Skip to content

Stable diagnostic codes, an ICU message catalog, and first error in source order - #166

Merged
luisleo526 merged 4 commits into
mainfrom
cg/diagnostic-codes
Oct 4, 2026
Merged

luisleo526 merged 4 commits into
mainfrom
cg/diagnostic-codes

Conversation

@luisleo526

@luisleo526 luisleo526 commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Lane CG-DIAGCODES, for the pineforge-app's 14-language diagnostics. Targets codegen 1.2.0: the public API grows, nothing changes in the emitted C++.

  • Codes and args on every diagnostic. Every Diagnostic now carries a stable code (PF-E1203 / PF-W0412) and raw named args. That covers transpile_full(...)["diagnostics"], CompileError.diagnostics (support checker, analyzer, codegen, lexer/parser/limits, libraries, request.security contexts) and both transpile_json envelopes.
  • How codes are assigned. No emitter call site changed. A lazy classifier (pineforge_codegen/diagnostic_codes.py) matches each diagnostic's English (message, hint) against the catalog's ICU templates. It runs on first access, after the transpile, so it never touches the C++. Literal text must match around the arguments, so a code is never guessed. A text no template renders gets PF-E0000 / PF-W0000, and the test suite refuses that.
  • Catalog. pineforge_codegen/diagnostics_catalog.json (628 codes, about 240 KB, one code per line) ships in the wheel and the Pyodide archive. It is exported as diagnostics_catalog() / render_diagnostic() and as @pineforge/codegen-pyodide/diagnostics_catalog.json, and attached to each GitHub release.
    • Each code has severity, area, an English ICU message template and an ICU hint template, plus a one-line explanation.
    • Each arg has a kind: identifier | type | keyword | number | vocab (with its closed values) | text. text marks open English the transpiler builds, such as a nested reason. It's an addition to the requested kinds.
  • Pin. tests/fixtures/diagnostic_codes_pin.json pins each code's meaning (severity, templates and arg names). Removing or repurposing a code fails CI, and adding one needs a catalog entry and a pin line.
    • scripts/gen_diagnostics_catalog.py extracts every template the source spells and --write adds codes. It reads emitter arguments, module tables, locals, callers' arguments, caught exceptions and helper return values.
  • Code ranges. The first digit is the area: 0 source (lexer, parser, limits), 1 support checker and requests, 2 analysis, 3 request.security, 4 libraries, 5 codegen, 6 array/matrix history, 7 ta.*. Inside a switch arm the support checker reports its refusals as warnings, so PF-W1nnn (nnn < 500) is the switch-arm twin of PF-E1nnn.
  • First error in source order (TOP's optional item). Since Emit checked strategy settings and contain C ABI exceptions #159, the settings metadata visited the input arguments (defval, options, minval, maxval, step) before the script body. An unknown name in a later input's minval was therefore raised ahead of an error on an earlier line. That error is now held until the body is generated, and the first error in source order is raised (tests/test_settings_error_order.py). No C++ changes.

Compatibility proof

  • Byte identity. Main ff67081 vs this head, one fresh process per file, transpile_full first:
    • 325 corpus + 263 fixture sources;
    • C++ (plus inputs, strategyParams and requests) identical 550/550;
    • diagnostics identical 1218/1218 (level, phase, location, message, hint);
    • glue envelope entries identical 1218/1218, code/args aside.
  • Coverage. No diagnostic in the corpus or fixtures falls back to PF-E0000 / PF-W0000. Across the pure-Python test suite's 669 distinct diagnostics, only 4 test-crafted texts (boom, …) are uncatalogued.
  • Pre-existing nondeterminism. _security_expr_hist_series_names sorts (sec_id, id(node)), so the order of clear_security lines follows memory addresses. Main's own C++ for tests/fixtures/tail_a_tv/taila_nested_bucket_hull_lag.pine differs between two harnesses (6a2434… vs ea70de…). This PR keeps codes lazy so the transpile's allocations stay main's. It does not fix the bug (a follow-up).

Test plan

  • tests/test_diagnostic_codes.py: catalog shape, pin, every spelled template coded, every corpus/fixture diagnostic renders back byte for byte, glue envelope.
  • tests/test_settings_error_order.py.
  • Rebased on main 463a9af (Emit immutable compiled execution capability receipts (capability extension v1) #167, capability receipts). Emit immutable compiled execution capability receipts (capability extension v1) #167 adds no diagnostic text; scripts/gen_diagnostics_catalog.py reports nothing uncoded.
  • Remote suite on CLOUD against engine f3bf9f07adfa (the active baseline):
    • compile-corpus 314/314 (rj-20261004t121634-9ddff2);
    • Pyodide gate PARITY OK over 277 fixtures (ok=264/264, err=13);
    • pytest 5675 passed, 41 skipped, 1 failed, in rj-20261004t121634-9ddff2 and again in rj-20261004t124132-4e0b90.
    • The failure is tests/test_array_history.py::test_many_bindings_are_walked_once_each, a 40 s wall-clock budget. It took 53–55 s on the CLOUD spot VMs and 38.9 s in the previous passing CLOUD run (pre-Emit immutable compiled execution capability receipts (capability extension v1) #167 head). On the Mac, old main 3cc1cf5, new main 463a9af and this head time equal (7.06–7.18 s). A CLOUD control run of plain main 463a9af is rj-20261004t130450-5f347d.
  • pr-gate PASS (no-improve-no-regression) for engine f3bf9f07adfa and codegen e81fe7ee2115, recorded (eventKey verdict-f3bf9f07adfa-e81fe7ee2115-80dc6517f2ac, receipt c6e96ed7…). Hard lane: 1009 probes, 0 regressions, rank sum 4030 → 4030. Target lane: 6980 probes, score 0, nothing changed.

🤖 Generated with Claude Code

luisleo526 and others added 4 commits October 4, 2026 20:14
Every Diagnostic carries a stable code (PF-E1203 / PF-W0412) and raw named
args, read off its English text by the ICU MessageFormat templates of
pineforge_codegen/diagnostics_catalog.json (628 codes: severity, area,
message and hint templates, argument kinds, a one-line explanation). The
message text and the emitted C++ are unchanged: the code is computed lazily,
after the transpile, so nothing in the pipeline reads it.

- diagnostics_catalog() / render_diagnostic() API; glue envelopes carry
  code and args; the release attaches the catalog.
- scripts/gen_diagnostics_catalog.py extracts every template the source
  spells and adds codes; tests/fixtures/diagnostic_codes_pin.json pins
  each code's meaning (never removed or repurposed).
- tests/test_diagnostic_codes.py: catalog shape, pin, every spelled
  template coded, and every diagnostic over the corpus and fixtures
  renders back to its message and hint.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…m twins in meaning

PF-W1077..PF-W1082 share their texts with PF-E1077..PF-E1082 but are emitted
as warnings wherever the variable diverges: their explanation drops the
switch-arm sentence.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…names

The settings metadata (1.1.0) visits every input's defval, options, minval,
maxval and step in the constructor, ahead of the body: an unknown name in a
later input's minval was raised before one on an earlier line. An error
there is now held until the body is generated, and the first in source
order is raised (tests/test_settings_error_order.py). No emitted C++
changes.

The catalog is rebuilt with clearer argument names (token, token_type,
node_type, timeframe, ...); every explanation and argument kind carried over.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@luisleo526
luisleo526 force-pushed the cg/diagnostic-codes branch from 8a588d4 to e81fe7e Compare October 4, 2026 12:36
@luisleo526
luisleo526 merged commit 521a6e9 into main Oct 4, 2026
9 checks passed
luisleo526 added a commit that referenced this pull request Oct 5, 2026
…he scoreboard README at pineforge-release 03b8dcc (#171)

* Read diagnostic codes off their text in linear time

The classifier fullmatched each catalog template as a regex with lazy
(.*?) arguments over the message and hint. Text a script spells reaches
those arguments, and with several arguments the regex backtracked over
every split: 0.55 s for 1.8 KB and 23 s for 3.4 KB of one template's text.

Each literal now goes to its leftmost place after the previous one, the
last to the text's end. That is the split the lazy regex returns, and the
only one tried when no argument name repeats, since a leftmost place leaves
the rest the most room. Where a name repeats (a receiver in the message and
the hint), later places are tried in the regex's order within a budget of
4,096 places. The same crafted text, grown to 6.6 MB, now classifies in
milliseconds.

Codes and args are unchanged:
- 669 distinct real texts (corpus, fixtures, suite);
- 7,512 generated samples, plain and holding the templates' own literals;
- tests/test_diagnostic_codes.py now pins equality with a reference
  lazy-regex matcher, and a timing bound on flooded text.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Keep two tests off a loaded host's clock

test_array_history.py::test_many_bindings_are_walked_once_each had a 40 s
wall-clock budget. It takes 13-14 s alone on our shared Linux test hosts, yet
every full-suite xdist run there since #166 failed it at 50-60 s. One such run,
timing three transpiles, read 58, 57 and 78 s of wall clock for 27, 27 and 33 s
of CPU at load 45-74: the wall clock counted the worker's wait for a CPU. The
test now budgets 60 s of CPU time, 1.8 times the most measured under that load.
The regression it guards, a walk of the rest of the script per binding, costs
about eight times today's time on any host.

test_every_spelled_template_has_a_code loaded scripts/gen_diagnostics_catalog.py
into the test process, and the generator's parse of every module of the package
then stayed in that xdist worker for the rest of the suite. The test now runs the
generator's own check (its default mode, the same comparison) in a subprocess,
and shows the templates without a code when it fails. That move alone left the
timing test at 50.3 s in the next full-suite run.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* License: the PineForge Source License 1.1; distributing the software bundled with a product is Commercial Use

The Distribution License said "Distributing copies under this section is not
Commercial Use", while Commercial Use (3) made embedding the software or its
Output into a product or service made available to others Commercial Use. A
vendor shipping the unmodified transpiler with or inside its product could
argue that it only distributed copies and owed nothing.

1.1 closes that, and changes nothing else:
- Distribution License: unless it is a permitted purpose, distributing the
  software, changed or not, embedded in or bundled with a product or service
  made available to others is Commercial Use under (3), which the license to
  distribute does not cover. Every other distribution of copies stays free and
  is not Commercial Use, as before (sharing copies, forks, mirrors, package
  registries).
- Commercial Use (3): embedding the software or its Output, changed or not,
  or distributing either bundled with such a product or service.
- Personal Trading, noncommercial purposes and organizations, personal uses,
  the Notices duty and Output are unchanged.

Surfaces: the LICENSE header, README (name and the Commercial Use bullet),
LEGAL.md (name, and one line: 1.1 replaces 1.0 for the code on main and in
future releases, and what changed), the pyproject.toml comment, AGENTS.md and
CLAUDE.md. Releases up to and including 1.1.0 keep the license text they
shipped with.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* docs: render the active scoreboard from the facts tokens at pineforge-release 03b8dcc

Re-render the marked README values against the hub's canonical facts at 03b8dcc
(the squash of release hub #27): the active scoreboard is
pineforge-parity-baseline-20261004-engine-6b77f061 (engine 6b77f061, codegen
285ac03), 7,970 excellent / 19 strong of 7,989 graded, none below strong.
Release scoreboards keep their own numbers. Only values inside pf markers change.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@luisleo526
luisleo526 deleted the cg/diagnostic-codes branch October 5, 2026 02:39
luisleo526 added a commit that referenced this pull request Oct 5, 2026
* docs: changelog and version scoping for 1.2.0

1.2.0 pairs with engine v1.2.0, whose pineforge.h defines
PF_CAPABILITIES_API_VERSION and declares the two capability functions
(engine #332) that the C++ of codegen 1.2.0 defines (#167). The changelog
section takes in the Unreleased notes and covers #166, #167, #168 and
#171: the capability receipt and what the live runner does with it, the
diagnostic codes and catalog, the first error in source order, the
linear-time classification as a security fix, the PineForge Source
License 1.1 and the pairing and migration steps. The Python and JSON
contract is additive; the engine's report harness is unchanged since
v1.1.0, so no report key changes.

The README, AGENTS.md / CLAUDE.md, CONTRIBUTING.md,
docs/PUBLIC_CONTRACT.md and docs/pine-cap-activation.md name 1.2.0 where
they named 1.1.0 as the current release or pair, the pairing table gets
a 1.2.0 row, the license badge reads 1.1, and the npm README lists the
catalog and LICENSE the package now carries. The release scoreboard
sentence and the pf markers are left to the release lane.

Measured for 1.2.0: the engine's 325 corpus sources and the 277 gate
fixtures transpile, with the capability block removed, to 1.1.0's C++
byte for byte, each in a fresh process (13 refusals, same messages), and
each in under 0.1 s on CPython 3.14.6 on an Apple M4 Max. The tutorial
run of the quick-start C++ against engine main 44eab7b1 (Linux x86_64,
GCC 13) prints 9 trades and +738.20, and 13 trades with the old sizing
declared, as 1.1.0 with engine v1.1.0 does.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* docs: release 1.2.0's grades from the facts tokens

README: the release sentence names 1.2.0 and renders releases[1.2.0]
(baseline pineforge-parity-baseline-20261005-engine-52292db9, 7,970
excellent / 19 strong of 7,989), and the scoreboard of main renders the
same baseline, from pineforge-release facts/facts.json as exported for the
1.2.0 release (sha256 c2f438f0). CHANGELOG: the 1.2.0 documentation entry
states the release's grades, as 1.1.0's does.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant