fix(shared): keep valid escapes decoded when percent-decoding malformed input - #133
Conversation
…ed input `safeDecodeURIComponent` returned the whole input unchanged as soon as `decodeURIComponent` threw, so a single bad escape left every other escape encoded, e.g. `filename*=utf-8''%E4%B8%AD%ZZ.txt` parsed as `%E4%B8%AD%ZZ.txt`. Fall back to decoding each run of `%XX` escapes on its own and keep the runs that still fail as-is, like Hono's `tryDecode`. The filename above now parses as `中%ZZ.txt`. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014w3GkqAVAgbiC1X6JcXr4e
…ode errors `safeDecodeURIComponent` tried the whole string, then on failure decoded each `%XX` run in its own try/catch. Every run that is not valid UTF-8 threw a `URIError`, so a crafted 16 KB Content-Disposition header such as `'%FF-'.repeat(4096)` took ~13 ms to parse (vs ~2 µs before the fallback). Match only RFC 3629 UTF-8 sequences spelled as escapes and decode each match. `decodeURIComponent` never throws on a match, so both try/catches go away: that input now takes ~12 µs and a malformed filename ~0.3 µs (was ~3 µs), well-formed input stays within ~1 µs. Valid sequences inside a run that is invalid as a whole are now decoded too: `%E4%B8%AD%FF` gives `中%FF` instead of staying encoded. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014w3GkqAVAgbiC1X6JcXr4e
…-8 regex Go back to catching `decodeURIComponent` errors per `%XX` run instead of matching an RFC 3629 grammar, which is harder to read and maintain. The whole-string attempt before the per-run fallback is dropped too: a valid UTF-8 sequence cannot span a non-escape character, so decoding run by run gives the same output, leaving one decode path and one try/catch. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014w3GkqAVAgbiC1X6JcXr4e
Most values are valid, so one native `decodeURIComponent` call is about twice as fast as scanning and decoding run by run. Only malformed input pays for the throw and the per-run fallback, which keeps the valid runs decoded. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014w3GkqAVAgbiC1X6JcXr4e
@standard-server/aws-lambda
@standard-server/core
@standard-server/fastify
@standard-server/fetch
@standard-server/node
@standard-server/peer
@standard-server/shared
commit: |
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
Merging this PR will not alter performance
Comparing Footnotes
|
There was a problem hiding this comment.
✅ No new issues found.
Reviewed changes — the partial-decode rework of safeDecodeURIComponent and its test coverage.
- Partial-decode fallback (
packages/shared/src/uri.ts) — when the whole-valuedecodeURIComponentthrows, maximal contiguous runs of%XXescapes are decoded independently, and a run that still throws is kept verbatim; the whole-value fast path keeps valid input byte-identical to before. - Unit coverage (
packages/shared/src/uri.test.ts) — headline case changed from "unchanged" to the decoded result, plus mixed valid/invalid, invalid-UTF-8, poisoned-run, and no-double-decode (%2541%ZZ→%41%ZZ) cases. - Content-Disposition coverage (
packages/core/src/utils.test.ts) — twofilename*cases pin the new behavior at the only downstream caller.
Verified locally: pnpm exec vitest run packages/shared packages/core passes (197 tests), eslint is clean on the changed files, and the new uri.test.ts case fails against the base uri.ts (so the regression test is discriminating). The run-level (non-greedy) decode choice — an invalid byte poisons its whole hex run — is conservative and intentional, and partial decoding exposes no bytes (%00, %2f, CRLF) that a fully valid encoding could not already produce.
DeepSeek Flash (default — pick a model for stronger reviews) | 𝕏

getFilenameFromContentDispositionnow decodes the valid escapes in a malformedfilename*. Before,safeDecodeURIComponentreturned the whole input unchanged as soon asdecodeURIComponentthrew. So one bad escape left every other escape encoded:filename*=utf-8''%E4%B8%AD%ZZ.txtcame back as%E4%B8%AD%ZZ.txtinstead of中%ZZ.txt.Fixes
%XXescapes is decoded separately. Runs that still fail are kept as-is. This is the same fallback as Hono'stryDecode.%E4%B8%AD%FFis unchanged.safeDecodeURIComponentno longer returns malformed input unchanged. Its only caller in this repo isgetFilenameFromContentDisposition. Code outside the repo that uses@standard-server/sharedand relies on the old behaviour will now get partly decoded output.Performance
decodeURIComponentcall, so there is no extra cost.%E4%B8%AD%ZZ.txttakes about 3.8 µs. A crafted 16 KBContent-Dispositionheader made of thousands of invalid runs takes about 13 ms, compared with about 2 µs before.Testing
safeDecodeURIComponenttests cover:%, a truncated sequence and a non-hex escape;%2541%ZZgives%41%ZZ.getFilenameFromContentDispositiontests cover a malformedfilename*with and without a charset prefix.main.vitest run(1311 passed),eslintand the repo-wide type check are clean.🤖 Generated with Claude Code
https://claude.ai/code/session_014w3GkqAVAgbiC1X6JcXr4e