Format-preserving PDF translation. The goal is to replace BabelDOC: keep the layout, and never silently drop text.
This repository is milestone 1. It opens a PDF, interprets content streams down to individual glyphs, and checks that every glyph is still accounted for. A pluggable translator can mark those glyphs as translated. Rewriting the PDF comes later, so nothing here deletes a text-showing operator.
| Path | Role |
|---|---|
crates/rpt-core |
Interpreter, font decoding, glyph coverage, translator |
crates/rpt-ffi |
C ABI (rpt_open, rpt_extract, rpt_translate, rpt_save, rpt_free) |
crates/rpt-cli |
rpt extract and rpt translate |
crates/rpt-qa |
Corpus fidelity report (rpt-qa) |
crates/rpt-python |
Native Python module via PyO3 (not the C ABI) |
include/rapidpdftrans.h |
C header generated by cbindgen |
bindings/cpp |
Header-only RAII wrapper |
bindings/go |
cgo wrapper over the static library |
testdata/hello.pdf |
One-page fixture for the binding smoke tests |
Coordinates are PDF user space: origin at the bottom left, y increasing upward (pdf-user-space-origin-bottom-left-y-up). Page /Rotate is reported on each page and is not applied to glyph positions. q / Q save the graphics state and the text state (Tc Tw Tz TL Ts Tf Tr). They do not save the text matrix, matching the PDF specification.
Rust stable (1.88 or newer; the workspace pins stable):
cargo test --workspace
cargo build -p rpt-cli --release
./target/release/rpt extract file.pdf --jsonrpt extract writes per-glyph JSON, a plain-text reconstruction, and a coverage report to stdout. rpt translate translates that text and still does not write a PDF.
Python (needs python3-dev):
cargo build -p rpt-python --release
cp target/release/librapidpdftrans.so rapidpdftrans.so
python3 -c 'import rapidpdftrans; print(rapidpdftrans.extract("testdata/hello.pdf")[:200])'C ABI, used by Go and C++:
cargo build -p rpt-ffi --release
# header: include/rapidpdftrans.h
# static library: target/release/librapidpdftrans.a
g++ -std=c++17 -I include -I bindings/cpp bindings/cpp/smoke.cpp \
target/release/librapidpdftrans.a -ldl -lm -lpthread -lgcc_s -o smokeGo (from bindings/go, after the release static library exists):
go test ./...rpt_save translates the open document and writes a PDF. It returns -1, and does not create the file, when RPT_LLM_API_KEY is unset.
The default backend is an OpenAI-compatible chat endpoint:
- base URL
https://hub.mypapers.top/api/llm/v1(RPT_LLM_BASE_URLor thebase_urloption) - model
auto(RPT_LLM_MODELor themodeloption) - API key from
RPT_LLM_API_KEYonly
An api_key field in options JSON is ignored. Explicit options win over environment variables, which win over the built-in defaults. Tests talk to an in-process mock server. The live test runs only when RPT_LLM_API_KEY is set.
RPT_TRANSLATOR selects the backend: maclaw (default), openai, google / google-v2 (RPT_GOOGLE_API_KEY), google-v3 (a service account in RPT_GOOGLE_CREDENTIALS or a bearer token in RPT_GOOGLE_ACCESS_TOKEN), or google-unofficial. The unofficial client is opt-in only. It is never used as a silent fallback.
export RPT_LLM_API_KEY=...
./target/release/rpt translate paper.pdf --from en --to zh --output paper.zh.pdf \
--glossary transformer=TransformerThe default output is pure target-language text: the translation replaces the original in place. Bilingual output is opt-in. --mode side-by-side puts the original page on the left and the translation on the right. --mode alternating writes the original page, then the translated page. --mode overlay keeps the original operators and draws the translation on the same page. --bilingual selects side-by-side unless --layout alternating or --layout overlay is set. --layout alone does not change the default. The same choice is output_mode (default replace) and, only when bilingual is requested, bilingual_layout in options JSON for rpt_save, the C++ save, and Python save. A layout field by itself leaves the PDF as pure translation. --max-pages N limits extraction to the first N pages. The CJK face is RPT_CJK_FONT, or Droid Sans Fallback / WenQuanYi / Noto Sans CJK when one of those is installed. The embedded font is a glyf subset (Identity-H, ToUnicode). Horizontal advances come from hmtx; OpenType GSUB is not applied. A segment that does not fit, or that shares an operator with a formula or another kept glyph, stays original.
Before the call, URLs, email addresses, {brace} groups, numbers, in-text citations ([12], [1-3], (Smith et al., 2020), Smith (2019)), and glossary terms are replaced with ⟦N⟧ placeholders and restored afterwards. Glossary terms are substituted locally (longest match, ASCII word boundaries), so the model does not have to obey a glossary table. A dropped placeholder is retried once and then reported as an error.
The References / Bibliography section is not translated. Headings such as References, Bibliography, Works Cited, 参考文献, Literatur, and Références start the section; numbered and author-year entries do too. It continues across columns and pages and stops at the next section (Appendix, 附录, A. Proofs, …). Those glyphs are kept_original with reason references. skip_references defaults to true; set it to false in options JSON, or pass --translate-references on the CLI, to translate the section.
rpt translate --output then rewrites those glyphs. Without --output they stay translated_pending_rewrite.
Every extracted glyph starts as pending. A document is complete only when each glyph has exactly one final disposition:
rewritten— translated text has been written back into the PDFkept_original— the original drawing is kept, with a reasonnon_text— classified as non-text, with a reason
translated_pending_rewrite records a translation and stays unresolved until rewrite. Extraction alone therefore reports every glyph as unresolved. A test fails if an extraction path skips a glyph and shrinks that list.
Glyph records include the page, Unicode (which may be several characters, or empty when unmapped), character-code bytes, GID when known, font name and size, the text rendering matrix, an approximate bounding box, advance, colors, render mode, invisible and clip flags, and the source location (content stream, Form XObject, annotation appearance, or Type3 char proc, plus operator index and byte range). Later milestones use that location to delete exactly the original text-showing operator.
- Graphics state:
q/Q,cm, text state,Tm/Td/TD/T*/'/", andTJkerning - Widths from simple
/Widthsand CID/W//DW; vertical writing via/W2//DW2 - Form XObjects (matrix, resources, inheritance, cycles) and annotation
/APnormal appearances - Type3 char procs, including text drawn inside them
- Inline images, without aborting the page on a bad operator
- Invisible text (
Tr3 or 7) is recorded and flagged, not dropped - Rectangular clips flag glyphs that fall outside; other clips set
clip_uncertain - Unicode fallback: ToUnicode CMap, then predefined CJK CMaps (decoded with
encoding_rs, not Adobe's CID tables), then simple encodings and/Differencesvia the Adobe Glyph List, then glyph names from a Type 1FontFilecleartext encoding (dup 65 /A put, including PFB), then the embedded font cmap. Unmapped glyphs are kept and flagged.
corpus/manifest.json lists real papers and openly licensed books (direct URL, source, license, sha256, size, feature tags). python3 corpus/fetch.py downloads them into a gitignored cache. Those files are not committed. corpus/ci/ is the exception used by CI: nine CC BY 4.0 papers, each under 700 KB, with URLs and licenses in corpus/ci/manifest.json. testdata/hello.pdf is a one-page file generated by cargo run -p rpt-core --example hello_pdf, not a corpus document. rpt-qa extracts every cached file, checks glyph coverage, compares text with Poppler, and render-diffs an identity byte copy (SSIM on the page and on non-text blocks). A lopdf save is reported separately because it rewrites the file. CI runs the small ci: true subset. The full corpus is the nightly Corpus fidelity workflow. See corpus/README.md.
rpt-bench scores a translated PDF on dropped lines, placeholder and formula retention, overflow, style, non-text SSIM, and reference-operator byte identity. python3 corpus/bench/run.py records the identity ceiling (each CI PDF compared with itself) in corpus/benchmarks/. BabelDOC and PDFMathTranslate, when installed, should share corpus/bench/identity_server.py so the translator is the same. Metric definitions are in corpus/benchmarks/README.md.
- M1 (this tree). Content-stream interpreter, glyph coverage, pluggable translator, C / Python / Go / C++ skeletons, CLI.
- M2. Layout analysis, placeholder protection beyond the current shield, translator caching and glossary integration in the layout model.
- M3 (started). Re-layout with
hmtxadvances, CJK line breaking, font subsetting, deletion of the original text operators, bilingual mode, and a workingrpt_save. GSUB shaping and full paragraph layout are still open. - M4. Richer layout, OCR for pages that are not real text, render-diff checks against PDFium.
- For a non-Identity CJK CMap,
/Wis looked up with the character code itself, not an Adobe-GB1 / Japan1 CID./DWapplies when that code is absent. Identity-H/V uses the numeric code as the CID. - Standard 14 fonts without
/WidthsuseMissingWidth(default 0). AFM metrics are not vendored, so those advances are wrong until the PDF embeds widths. - Rectangle clips are exact enough to flag glyphs. Arbitrary paths only set
clip_uncertain. - A Type3 character records both the Type3 glyph and any text-showing operators inside its char proc.
- Glyph boxes are an em approximation (descent 0.2, ascent 0.8), not ink bounds.
- ZapfDingbats encoding is not included.
- PDFium is not a runtime dependency. A cross-check test was not added because libpdfium is not assumed to be installed.
- The in-memory translation cache is per call, keyed by the shielded source string.