Skip to content

Read diagnostic codes off their text in linear time - #169

Closed
luisleo526 wants to merge 1 commit into
mainfrom
cg/diagcodes-linear-match
Closed

luisleo526 wants to merge 1 commit into
mainfrom
cg/diagcodes-linear-match

Conversation

@luisleo526

@luisleo526 luisleo526 commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

Summary

A follow-up to #166: diagnostic codes are now read off their text in linear time.

  • The problem. diagnostic_codes.classify fullmatched each catalog template as a regex with lazy (?P<x>.*?) arguments, compiled with DOTALL over the message, a NUL and the hint. Text a user's script spells (identifiers, expressions, request keys) reaches those arguments. With several arguments the regex backtracked over every split.
    • Example: PF-W1544 with each argument flooded with its own separator and a hint that does not match.
    • It took 0.55 s at 1.8 KB and 23 s at 3.4 KB, and kept growing.
    • Every glue envelope reads code, so the Pyodide worker reads it for every diagnostic.
  • The fix. A literal-split matcher:
    • Each literal segment goes to its leftmost place after the previous one, and the last to the text's end. That is exactly the split the lazy-regex fullmatch returns.
    • When no argument name repeats, that split succeeds whenever any does, because a leftmost place leaves the rest the most room. It is the only one tried, so the cost is linear.
    • 56 templates name an argument twice, such as {receiver} in the message and the hint (obj.field[1] holds the separator .). For those, later places are tried in the regex's order within a budget of 4,096 places. A crafted text that exhausts the budget gets PF-E0000 / PF-W0000.
    • The same crafted text, grown to 6.6 MB, now classifies in milliseconds. 1 MB texts take 0–6 ms against every template.
  • No API, catalog, pin or message change. The emitted C++ is untouched; the classifier is not on the transpile path.

Evidence

  • Same codes and args as the regex matcher:
    • 669/669 distinct real texts (corpus, fixtures and the test suite's captured diagnostics);
    • 7,512/7,512 generated samples, plain and with arguments holding the templates' own literals.
    • tests/test_diagnostic_codes.py::test_classification_is_the_lazy_regex_split pins this against a reference lazy-regex matcher. test_classification_is_linear_on_crafted_text bounds the time on flooded text.
  • Identity vs main 285ac03, one fresh process per file over 325 corpus and 263 fixture sources:
    • C++ (plus inputs, strategyParams, requests): 550/550 identical;
    • diagnostics: 1218/1218 identical;
    • glue envelope entries including code and args: 1218/1218 identical.
  • pr-gate: PASS (no-improve-no-regression) for engine f3bf9f07adfa and codegen 3ba9edf5f968, against the active baseline pineforge-parity-baseline-20261004-codegen-285ac035, recorded.
    • eventKey verdict-f3bf9f07adfa-3ba9edf5f968-04e60e8ca489, receipt 3f23bc32….
    • Hard lane: 1009 probes, 0 regressions, rank sum 4030 → 4030. Target lane: 6980 probes, score 0.
  • Remote CLOUD suite (rj-20261004t134808-4d1a31, engine f3bf9f07):
    • compile-corpus 314/314;
    • Pyodide gate PARITY OK over 277 fixtures (ok=264/264, err=13);
    • pytest 5677 passed, 41 skipped, 1 failed. The failure is test_array_history.py::test_many_bindings_are_walked_once_each, a 40 s wall-clock budget under full-suite xdist contention.
    • That test fails the same way on Stable diagnostic codes, an ICU message catalog, and first error in source order #166's head. An isolated CLOUD A/B, one VM with no xdist (rj-20261004t133235-e4c237), ran it at 13.4–14.3 s on 463a9af and 13.6–13.8 s on 521a6e9b.

Test plan

  • tests/test_diagnostic_codes.py: 600 tests pass locally, including the two new ones.
  • Identity, equivalence and pr-gate (above).
  • Remote pytest, compile-corpus and Pyodide gate on CLOUD (the timing test aside, above).

🤖 Generated with Claude Code

The classifier fullmatched each catalog template as a regex with lazy
(.*?) arguments over the message and hint. Text a script spells reaches
those arguments, and with several arguments the regex backtracked over
every split: 0.55 s for 1.8 KB and 23 s for 3.4 KB of one template's text.

Each literal now goes to its leftmost place after the previous one, the
last to the text's end. That is the split the lazy regex returns, and the
only one tried when no argument name repeats, since a leftmost place leaves
the rest the most room. Where a name repeats (a receiver in the message and
the hint), later places are tried in the regex's order within a budget of
4,096 places. The same crafted text, grown to 6.6 MB, now classifies in
milliseconds.

Codes and args are unchanged:
- 669 distinct real texts (corpus, fixtures, suite);
- 7,512 generated samples, plain and holding the templates' own literals;
- tests/test_diagnostic_codes.py now pins equality with a reference
  lazy-regex matcher, and a timing bound on flooded text.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@luisleo526

Copy link
Copy Markdown
Contributor Author

Superseded by #171, which carries this commit unchanged (3ba9edf) in one branch with a test-timing fix, the PineForge Source License 1.1 and the scoreboard README (#170), so the pr-gate measures them in one sweep. Closing in favour of #171.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant