Skip to content

Linear-time diagnostic codes, the PineForge Source License 1.1, and the scoreboard README at pineforge-release 03b8dcc - #171

Merged
luisleo526 merged 4 commits into
mainfrom
gap/20261005
Oct 5, 2026
Merged

luisleo526 merged 4 commits into
mainfrom
gap/20261005

Conversation

@luisleo526

Copy link
Copy Markdown
Contributor

Summary

This PR puts three pending codegen changes in one branch, so the pr-gate measures them in one sweep. It supersedes #169 and #170. Each change is its own commit (the first has two), in this order, on main 285ac03:

  1. Diagnostic codes read off their text in linear time (was Read diagnostic codes off their text in linear time #169)
    • 3ba9edf is Read diagnostic codes off their text in linear time #169's head, unchanged.
      • diagnostic_codes.classify fullmatched each catalog template as a regex with lazy arguments. Text a user's script spells reaches those arguments, and the regex backtracked over every split: a crafted PF-W1544 text took 0.6 s at 1.8 KB and 30 s at 3.4 KB.
      • It is now a literal-split matcher. Each literal goes to its leftmost place, which is the split the lazy regex returns.
      • For the 56 templates that name an argument twice, later places are tried in the regex's order, within a budget of 4,096 places. Text that exhausts the budget gets PF-E0000 / PF-W0000.
      • No API, catalog, pin or message change. The emitted C++ is untouched.
    • 69d368f changes tests only: two tests no longer depend on a loaded host's clock.
      • test_array_history.py::test_many_bindings_are_walked_once_each budgeted 40 s of wall clock. It takes 13–14 s alone on our shared Linux test hosts, yet every full-suite xdist run there since Stable diagnostic codes, an ICU message catalog, and first error in source order #166 failed it, at 50–60 s.
      • One such run timed three transpiles: 58, 57 and 78 s of wall clock for 27, 27 and 33 s of CPU, at load 45–74.
      • The test now budgets 60 s of CPU time (time.process_time()), 1.8 times the most measured under that load.
      • The regression it guards costs about eight times today's time: about 110 s on hosts like these, which fails. On the workstation where this run takes 7 s, the regression took 57 s, which now passes. A follow-up lowers the budget to 45 s so a fast workstation fails it too.
      • test_diagnostic_codes.py::test_every_spelled_template_has_a_code runs the catalog generator's own check in a subprocess. Loaded into the test process, the generator's parse of the whole package stayed in that xdist worker for the rest of the suite. That move alone left the timing test at 50.3 s.
  2. The PineForge Source License 1.1 (e42f9c7)
    • Distributing the software, changed or not, embedded in or bundled with a product or service made available to others is Commercial Use, unless it is for a permitted purpose.
    • 1.0's Distribution License said "Distributing copies under this section is not Commercial Use". A vendor shipping the unmodified transpiler inside a product could read that as owing nothing.
    • Every other distribution of copies stays free and is not Commercial Use. Personal Trading, noncommercial use, the Notices duty and Output are unchanged.
    • Files: the LICENSE header and text, README (the license name and the Commercial Use bullet), LEGAL.md, the pyproject.toml comment, AGENTS.md and CLAUDE.md. Releases up to and including 1.1.0 keep the license text they shipped with.
  3. The scoreboard README at pineforge-release 03b8dcc (was docs: render the active scoreboard from the facts tokens at pineforge-release 03b8dcc #170; 6904c30)
    • Re-renders the README's marked scoreboard values from the release hub's canonical facts at 03b8dcc8c6f0d60531309bf0abbb3fcad749ec55.
    • That scoreboard is pineforge-parity-baseline-20261004-engine-6b77f061: 7,970 excellent and 19 strong of 7,989 graded, none below strong. Only values inside pf markers change.

3ba9edf keeps #169's SHA. e42f9c7 and 6904c30 carry the reviewed commits' patches unchanged (git range-diff reports them identical).

Evidence

  • pr-gate: PASS (no-improve-no-regression) for engine a80c623e0d2f and codegen 6904c3064dd2, against the active baseline pineforge-parity-baseline-20261004-engine-a80c623e. Recorded.

    • eventKey verdict-a80c623e0d2f-6904c3064dd2-d97b8579f85f, receipt e2b703f8….
    • Hard lane: 1009 probes, 0 regressions, rank sum 4032 → 4032, no moves.
    • Target lane: 6980 probes, score 0, no entrant or leaver.
    • Coverage: 7,989 of 7,989 probes measured.
  • Full suite on a remote Linux test host (this head, engine a80c623e):

    • pytest: 5,678 passed, 41 skipped, 0 failed. The skips are environmental: reference commits missing from the shipped history, no clang++, empty parameter sets.
    • test_many_bindings_are_walked_once_each passed within its CPU budget. Its wall clock there was 60.4 s, which the old 40 s wall-clock budget would have failed.
    • compile-corpus: 314/314.
    • Pyodide gate: PARITY OK over 277 fixtures (ok 264/264, err 13).
    • The same run on the previous head, before the CPU budget, failed only that test (50.3 s of wall clock).
  • Linear matcher on this branch:

    • Same code and args as main's regex matcher on 669/669 real texts and 7,512/7,512 generated samples.
    • Crafted PF-W1544 text from 1.8 KB to 6.6 MB classifies in 5 ms or less; main takes 0.6 s at 1.8 KB and 30 s at 3.4 KB.
    • The worst template that names an argument twice, flooded about 2,000 times, takes 0.026 s.
  • Local (macOS, engine a80c623e headers, syntax-only compiles):

    • tests/test_diagnostic_codes.py and tests/test_array_history.py: 749 passed, 0 skipped. This includes test_classification_is_the_lazy_regex_split and test_classification_is_linear_on_crafted_text.
    • tests/test_compile_corpus.py: 314/314 passed, 0 skipped.
    • test_many_bindings_are_walked_once_each with the CPU budget: passes in 7.9 s.
    • Fail check for the subprocess test: with PF-E0001 dropped from a copy of the catalog, it fails and names the template at pineforge_codegen/lexer.py:333.
    • ruff check: the changed Python files are clean, and the repository's 170 existing findings are the same as on main.
    • lab facts check against the facts pinned to 03b8dcc: PASS, 15 markers, 0 drift.

🤖 Generated with Claude Code

luisleo526 and others added 4 commits October 4, 2026 21:46
The classifier fullmatched each catalog template as a regex with lazy
(.*?) arguments over the message and hint. Text a script spells reaches
those arguments, and with several arguments the regex backtracked over
every split: 0.55 s for 1.8 KB and 23 s for 3.4 KB of one template's text.

Each literal now goes to its leftmost place after the previous one, the
last to the text's end. That is the split the lazy regex returns, and the
only one tried when no argument name repeats, since a leftmost place leaves
the rest the most room. Where a name repeats (a receiver in the message and
the hint), later places are tried in the regex's order within a budget of
4,096 places. The same crafted text, grown to 6.6 MB, now classifies in
milliseconds.

Codes and args are unchanged:
- 669 distinct real texts (corpus, fixtures, suite);
- 7,512 generated samples, plain and holding the templates' own literals;
- tests/test_diagnostic_codes.py now pins equality with a reference
  lazy-regex matcher, and a timing bound on flooded text.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
test_array_history.py::test_many_bindings_are_walked_once_each had a 40 s
wall-clock budget. It takes 13-14 s alone on our shared Linux test hosts, yet
every full-suite xdist run there since #166 failed it at 50-60 s. One such run,
timing three transpiles, read 58, 57 and 78 s of wall clock for 27, 27 and 33 s
of CPU at load 45-74: the wall clock counted the worker's wait for a CPU. The
test now budgets 60 s of CPU time, 1.8 times the most measured under that load.
The regression it guards, a walk of the rest of the script per binding, costs
about eight times today's time on any host.

test_every_spelled_template_has_a_code loaded scripts/gen_diagnostics_catalog.py
into the test process, and the generator's parse of every module of the package
then stayed in that xdist worker for the rest of the suite. The test now runs the
generator's own check (its default mode, the same comparison) in a subprocess,
and shows the templates without a code when it fails. That move alone left the
timing test at 50.3 s in the next full-suite run.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…bundled with a product is Commercial Use

The Distribution License said "Distributing copies under this section is not
Commercial Use", while Commercial Use (3) made embedding the software or its
Output into a product or service made available to others Commercial Use. A
vendor shipping the unmodified transpiler with or inside its product could
argue that it only distributed copies and owed nothing.

1.1 closes that, and changes nothing else:
- Distribution License: unless it is a permitted purpose, distributing the
  software, changed or not, embedded in or bundled with a product or service
  made available to others is Commercial Use under (3), which the license to
  distribute does not cover. Every other distribution of copies stays free and
  is not Commercial Use, as before (sharing copies, forks, mirrors, package
  registries).
- Commercial Use (3): embedding the software or its Output, changed or not,
  or distributing either bundled with such a product or service.
- Personal Trading, noncommercial purposes and organizations, personal uses,
  the Notices duty and Output are unchanged.

Surfaces: the LICENSE header, README (name and the Commercial Use bullet),
LEGAL.md (name, and one line: 1.1 replaces 1.0 for the code on main and in
future releases, and what changed), the pyproject.toml comment, AGENTS.md and
CLAUDE.md. Releases up to and including 1.1.0 keep the license text they
shipped with.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…-release 03b8dcc

Re-render the marked README values against the hub's canonical facts at 03b8dcc
(the squash of release hub #27): the active scoreboard is
pineforge-parity-baseline-20261004-engine-6b77f061 (engine 6b77f061, codegen
285ac03), 7,970 excellent / 19 strong of 7,989 graded, none below strong.
Release scoreboards keep their own numbers. Only values inside pf markers change.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
@luisleo526
luisleo526 merged commit 48e7a13 into main Oct 5, 2026
9 checks passed
@luisleo526
luisleo526 deleted the gap/20261005 branch October 5, 2026 00:37
luisleo526 added a commit that referenced this pull request Oct 5, 2026
* docs: changelog and version scoping for 1.2.0

1.2.0 pairs with engine v1.2.0, whose pineforge.h defines
PF_CAPABILITIES_API_VERSION and declares the two capability functions
(engine #332) that the C++ of codegen 1.2.0 defines (#167). The changelog
section takes in the Unreleased notes and covers #166, #167, #168 and
#171: the capability receipt and what the live runner does with it, the
diagnostic codes and catalog, the first error in source order, the
linear-time classification as a security fix, the PineForge Source
License 1.1 and the pairing and migration steps. The Python and JSON
contract is additive; the engine's report harness is unchanged since
v1.1.0, so no report key changes.

The README, AGENTS.md / CLAUDE.md, CONTRIBUTING.md,
docs/PUBLIC_CONTRACT.md and docs/pine-cap-activation.md name 1.2.0 where
they named 1.1.0 as the current release or pair, the pairing table gets
a 1.2.0 row, the license badge reads 1.1, and the npm README lists the
catalog and LICENSE the package now carries. The release scoreboard
sentence and the pf markers are left to the release lane.

Measured for 1.2.0: the engine's 325 corpus sources and the 277 gate
fixtures transpile, with the capability block removed, to 1.1.0's C++
byte for byte, each in a fresh process (13 refusals, same messages), and
each in under 0.1 s on CPython 3.14.6 on an Apple M4 Max. The tutorial
run of the quick-start C++ against engine main 44eab7b1 (Linux x86_64,
GCC 13) prints 9 trades and +738.20, and 13 trades with the old sizing
declared, as 1.1.0 with engine v1.1.0 does.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* docs: release 1.2.0's grades from the facts tokens

README: the release sentence names 1.2.0 and renders releases[1.2.0]
(baseline pineforge-parity-baseline-20261005-engine-52292db9, 7,970
excellent / 19 strong of 7,989), and the scoreboard of main renders the
same baseline, from pineforge-release facts/facts.json as exported for the
1.2.0 release (sha256 c2f438f0). CHANGELOG: the 1.2.0 documentation entry
states the release's grades, as 1.1.0's does.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant