fix(node): encode string chunks back to bytes instead of corrupting them - #131
Conversation
A request stream yields string chunks when something upstream calls `setEncoding` on it, or when it is an object-mode stream. `toWebReadableStream` passed them to `new Uint8Array(value)`, which reads a string as a length: '123' became 123 zero bytes and most bodies vanished. The JSON and URL-encoded readers threw a TypeError on the same input, and file bodies ignored the stream's encoding. String chunks are now encoded back with the stream's own encoding (utf8 by default), which recovers the original bytes for base64, hex and latin1 as well.
@standard-server/aws-lambda
@standard-server/core
@standard-server/fastify
@standard-server/fetch
@standard-server/node
@standard-server/peer
@standard-server/shared
commit: |
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
Merging this PR will not alter performance
Comparing Footnotes
|
There was a problem hiding this comment.
ℹ️ No critical issues — a correct, well-tested fix. One adjacent body path (
form-data) still bypasses the new helper.
Reviewed changes
- New
readableChunkToByteshelper —packages/node/src/utils.ts:64re-encodes string chunks withBuffer.from(chunk, stream.readableEncoding ?? 'utf8')and passes byte chunks through. - Stream readers route through it —
toWebReadableStream,_streamToString, and_streamToFilenow convert chunks before use, fixing zero-filled/empty binary bodies and theTypeErrorthat hung JSON/url-encoded requests undersetEncodingor object mode. - Typing cleanup —
_streamToFilechunks retyped toUint8Array<ArrayBuffer>[]; the now-unusedBuffertype import was dropped frombody.ts. - Tests — object-mode string streams,
utf8/base64/hex/latin1encodings with bytes split across chunks, and JSON/file requests withsetEncoding.
ℹ️ form-data bodies still bypass the new helper
_streamToFormData (packages/node/src/body.ts:156) hands the raw request Readable straight to new Response(stream, …).formData(), so its chunks never pass through readableChunkToBytes. With a lossless non-utf8 encoding set upstream (base64/hex/latin1) multipart parsing throws Failed to parse body as FormData. no boundary found in multipart body. This is pre-existing, not introduced here, and the realistic setEncoding('utf8') case is unrecoverable for binary parts anyway — but for the lossless encodings the same helper would restore the bytes, so it's the one stream-reader left inconsistent with the fix.
Technical details
# form-data path bypasses readableChunkToBytes
## Affected sites
- `packages/node/src/body.ts:156-164` — `_streamToFormData` constructs `new Response(stream, …)` from the raw Node `Readable`. undici treats each string chunk as already-decoded text rather than re-encoding via the stream's `readableEncoding`, unlike the other three readers touched by this PR.
## Evidence
- Reproduced on Node 24: a multipart request stream with `req.setEncoding('base64')` (also `hex`/`latin1`) throws `TypeError: Failed to parse body as FormData. Caused by: TypeError: no boundary found in multipart body`. The same body parses with the encoding unset or `utf8`.
## Required outcome
- The `form-data` branch should feed the parser the same bytes an un-encoded stream would, for encodings where that is recoverable.
## Suggested approach (optional)
- Pipe the source through the existing `toWebReadableStream(stream)` (which now applies `readableChunkToBytes`) before constructing the `Response`, or apply the helper per chunk in a small transform.DeepSeek Flash (default — pick a model for stronger reviews) | 𝕏

Request bodies with string chunks are no longer corrupted. A Node stream yields strings instead of bytes when something upstream calls
setEncodingon the request, or when it is an object-mode stream.toWebReadableStreampassed those strings tonew Uint8Array(value), which reads a string as a length:'123'became 123 zero bytes and most bodies came through empty. String chunks are now encoded back to bytes with the stream's own encoding.Fixes
TypeError(and hang the request) on string chunks.base64, not the encoded text.@standard-server/node. AWS Lambda was never affected, since it doesn't read from a Node stream.API
readableChunkToBytes(stream, chunk), sinceutils.tsis re-exported from the package index.Testing
utf8/base64/hex/latin1encoding set (with a character split across chunks), and JSON and file requests withsetEncodingcalled.Fileencodes strings as utf8.