Skip to content

fix: stop asking for zstd that trafilatura cannot decompress - #18

Merged
zyocum merged 1 commit into
masterfrom
fix/trafilatura-zstd
Sep 12, 2026
Merged

zyocum merged 1 commit into
masterfrom
fix/trafilatura-zstd

Conversation

@zyocum

@zyocum zyocum commented Sep 12, 2026

Copy link
Copy Markdown
Owner

Pages served with content-encoding: zstd failed to extract at all,
reporting a parse error against perfectly ordinary HTML:

$ reader https://simonwillison.net/2026/Sep/1/geojson/ -f json
[PARSE ERROR - failed to extract content] https://simonwillison.net/2026/Sep/1/geojson/

What is actually going on

Not the site, and not bot detection. The response is HTTP 200 and byte-identical
whichever User-Agent asks for it — fetching that URL with trafilatura's own
trafilatura/2.2.0 (+https://github.com/adbar/trafilatura) User-Agent through a
plain HTTP client returns the full 18,012 characters.

The body is a valid zstd stream that trafilatura fails to decompress:

  1. It advertises accept-encoding: gzip,deflate,zstd.
  2. The server obliges with content-encoding: zstd.
  3. handle_compressed_file() calls zstandard.decompress(filecontent) with no
    max_output_size.
  4. That raises ZstdError: could not determine content size in frame header for
    any frame omitting its decompressed size — which is what a server compressing
    a response on the fly sends.
  5. The error is caught, "invalid ZSTD file" is logged, and the compressed
    bytes are returned unchanged
    , decoding into binary garbage.

Confirmed directly:

>>> zstandard.decompress(data)
ZstdError: could not determine content size in frame header
>>> zstandard.decompress(data, max_output_size=10_000_000)[:15]
b'<!DOCTYPE html>'

Installing zstandard does not help — it makes the broken path more
reachable, not less, and urllib3's make_headers(accept_encoding=True)
advertises zstd regardless of whether trafilatura can decode it.

The change

Ask only for encodings trafilatura can actually decode. gzip is universally
served, and this avoids the broken path entirely rather than working around it.

DEFAULT_HEADERS["accept-encoding"] = "gzip, deflate"

On simonwillison.net this is the difference between a parse error and a
4,038-word extraction.

Why a workaround rather than a fix

This is fixed upstream by a streaming _decompress_zstd() whose own docstring
names the cause ("The one-shot API only accepts frames declaring their
decompressed size in the header, which servers compressing their responses on
the fly omit"). It is unreleased: 2.2.0 (31 Jul 2026) is the latest on PyPI
and has the bug. I verified trafilatura master fixes both failing pages.

The removal condition is recorded in AGENTS.md under a new Upstream
workarounds
section: raise the trafilatura floor in pyproject.toml and
delete the override once a release carries the fix, with a one-line command to
check whether it has landed.

Verification

  • Both regression pages from AGENTS.md (paulgraham.com/greatwork.html,
    gnu.org/philosophy/free-sw.en.html) plus the previously failing site
  • All four output formats, -w wrapping, the - stdin path, and the installed
    reader console script
  • ruff check . clean; basedpyright reader.py zero errors and zero warnings

Notes for review

  • The import needed a targeted # pyright: ignore[reportUnknownVariableType],
    matching the house style already used for the lxml typing gaps.
  • The ~200 other basedpyright findings are pre-existing and all in the untracked
    sandbox/ directory, which [tool.pyright] include = ["**/*.py"] picks up
    even though git excludes it. Untouched here.

🤖 Generated with Claude Code

Pages served with `content-encoding: zstd` failed to extract at all,
reporting a parse error against perfectly ordinary HTML.

The cause is not the site and not bot detection: the response is HTTP 200
and identical whichever User-Agent asks for it. trafilatura 2.2.0
advertises `gzip,deflate,zstd`, and on receiving zstd calls
`zstandard.decompress()` with no `max_output_size`. That raises
`ZstdError: could not determine content size in frame header` for any frame
omitting its decompressed size, which is what a server compressing on the
fly sends. trafilatura catches the error, logs "invalid ZSTD file", and
returns the compressed bytes, which decode into binary garbage.

Asking only for encodings it can actually decode avoids the path entirely;
gzip is universally served. On simonwillison.net this is the difference
between a parse error and a 4,038-word extraction.

Fixed upstream by a streaming `_decompress_zstd()` but unreleased as of
2.2.0, so this is a workaround with a removal condition, recorded in
AGENTS.md: raise the trafilatura floor and delete the override once a
release carries the fix.

Verified against both regression pages from AGENTS.md plus the failing
site, in all four output formats; ruff and basedpyright clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@zyocum
zyocum merged commit 873c0d4 into master Sep 12, 2026
1 check passed
@zyocum
zyocum deleted the fix/trafilatura-zstd branch September 12, 2026 19:18
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant