fix: stop asking for zstd that trafilatura cannot decompress - #18
Merged
Merged
Conversation
Pages served with `content-encoding: zstd` failed to extract at all, reporting a parse error against perfectly ordinary HTML. The cause is not the site and not bot detection: the response is HTTP 200 and identical whichever User-Agent asks for it. trafilatura 2.2.0 advertises `gzip,deflate,zstd`, and on receiving zstd calls `zstandard.decompress()` with no `max_output_size`. That raises `ZstdError: could not determine content size in frame header` for any frame omitting its decompressed size, which is what a server compressing on the fly sends. trafilatura catches the error, logs "invalid ZSTD file", and returns the compressed bytes, which decode into binary garbage. Asking only for encodings it can actually decode avoids the path entirely; gzip is universally served. On simonwillison.net this is the difference between a parse error and a 4,038-word extraction. Fixed upstream by a streaming `_decompress_zstd()` but unreleased as of 2.2.0, so this is a workaround with a removal condition, recorded in AGENTS.md: raise the trafilatura floor and delete the override once a release carries the fix. Verified against both regression pages from AGENTS.md plus the failing site, in all four output formats; ruff and basedpyright clean. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Pages served with
content-encoding: zstdfailed to extract at all,reporting a parse error against perfectly ordinary HTML:
What is actually going on
Not the site, and not bot detection. The response is HTTP 200 and byte-identical
whichever User-Agent asks for it — fetching that URL with trafilatura's own
trafilatura/2.2.0 (+https://github.com/adbar/trafilatura)User-Agent through aplain HTTP client returns the full 18,012 characters.
The body is a valid zstd stream that trafilatura fails to decompress:
accept-encoding: gzip,deflate,zstd.content-encoding: zstd.handle_compressed_file()callszstandard.decompress(filecontent)with nomax_output_size.ZstdError: could not determine content size in frame headerforany frame omitting its decompressed size — which is what a server compressing
a response on the fly sends.
"invalid ZSTD file"is logged, and the compressedbytes are returned unchanged, decoding into binary garbage.
Confirmed directly:
Installing
zstandarddoes not help — it makes the broken path morereachable, not less, and urllib3's
make_headers(accept_encoding=True)advertises zstd regardless of whether trafilatura can decode it.
The change
Ask only for encodings trafilatura can actually decode. gzip is universally
served, and this avoids the broken path entirely rather than working around it.
On simonwillison.net this is the difference between a parse error and a
4,038-word extraction.
Why a workaround rather than a fix
This is fixed upstream by a streaming
_decompress_zstd()whose own docstringnames the cause ("The one-shot API only accepts frames declaring their
decompressed size in the header, which servers compressing their responses on
the fly omit"). It is unreleased: 2.2.0 (31 Jul 2026) is the latest on PyPI
and has the bug. I verified trafilatura
masterfixes both failing pages.The removal condition is recorded in
AGENTS.mdunder a new Upstreamworkarounds section: raise the
trafilaturafloor inpyproject.tomlanddelete the override once a release carries the fix, with a one-line command to
check whether it has landed.
Verification
AGENTS.md(paulgraham.com/greatwork.html,gnu.org/philosophy/free-sw.en.html) plus the previously failing site-wwrapping, the-stdin path, and the installedreaderconsole scriptruff check .clean;basedpyright reader.pyzero errors and zero warningsNotes for review
# pyright: ignore[reportUnknownVariableType],matching the house style already used for the lxml typing gaps.
sandbox/directory, which[tool.pyright] include = ["**/*.py"]picks upeven though git excludes it. Untouched here.
🤖 Generated with Claude Code