Please report security issues privately via GitHub's private vulnerability reporting rather than opening a public issue.
Do not attach a document containing personal or confidential information. A structural description is almost always enough to reproduce a parsing bug:
doc-stitch yourfile.docx --dry-run --jsonEvery input this tool accepts is a file somebody else made. Both engines are written on that assumption.
The web app runs entirely in the browser and makes no network requests, so a malicious document is confined to the tab that opened it. The CLI and the engine packages have no such boundary — if you accept uploads and merge them server-side, the notes below are yours to think about.
packages/docx/test/security.test.ts locks in each of these.
| Attack | Behaviour |
|---|---|
| XXE — external entity pointing at a local file or URL | Entities are never resolved. The parser rejects the document rather than returning an empty value, so the attempt is visible instead of silent. |
| Billion laughs — recursive entity expansion | Rejected in milliseconds; no expansion is attempted. |
| Zip bomb — small archive, enormous contents | Total decompressed size is checked against a limit before inflating, so a bomb is refused rather than allocated. Default 256 MB, override with maxUncompressedBytes. |
| Zip slip — part names escaping the package root | Relationship targets are normalized; a target of ../../../../evil.png is imported under a name inside the package. Output part names can never begin with .., /, or a drive letter. |
| XSS via filename | The web UI builds DOM nodes and assigns textContent. There is no innerHTML, insertAdjacentHTML, or eval anywhere in the source. |
- PDF parsing is delegated. The PDF engine builds on
@cantoo/pdf-lib, and a malicious PDF's blast radius is whatever that library allows. There is no equivalent to the ZIP size guard, because PDF streams are decompressed lazily by the library rather than up front. - The zip guard trusts declared sizes. It reads each entry's stated uncompressed size from the archive. An archive that lies would be caught during inflation instead — later, and only after some memory has been committed.
- Merging is memory-resident. Both engines hold every input and the output in memory at once. Size your limits accordingly if you run this server-side.
--dry-runruns the real merge to collect accurate warnings. It writes nothing, but it does the same parsing work, so it is not a cheap way to pre-screen hostile input.
- Encrypted PDFs are refused, not opened. No attempt is made to work around a password.
- Nothing is uploaded, ever. The web app has no network code. If you find any, that is a bug worth reporting on its own.
Runtime dependencies are kept few and audited: fflate and @xmldom/xmldom for Word,
@cantoo/pdf-lib for PDF. npm audit --omit=dev should report zero vulnerabilities; CI fails if
it does not.