[Deps] Require laya 0.3.9, lock 0.3.20, and drop the local weight-init skip - #37
cacheline999 wants to merge 5 commits into
Conversation
laya 0.3.9 builds the encoder under transformers' no_init_weights inside laya.load and loads the checkpoint with strict=True, so the without_weight_init() wrapper from ThinkFlowLab#22 and its test stubs are no longer needed. The lock moves laya from 0.3.5 to 0.3.9, the first release with the skip; later releases change MPS precision and are left for their own bump. The system_one config-error test pins the version lookup, so it passes with the laya extra installed as well as on the core install. Closes ThinkFlowLab#28
uv lock --upgrade-package laya, as ThinkFlowLab#28 asks, pinned to 0.3.20, the release system1-omni's worker runs. Only the laya entry of the lock changes. From 0.3.10 laya runs a request of five or more questions in fp16 on MPS. Against fp32 that moved probabilities by up to 0.05 on the multilingual checkpoint and flipped 2 of 180 decisions, and on an M1 Pro it was slower for short states. LayaModel.from_env now raises the agent's mps_amp_min_rows so such requests stay in fp32, unless LAYA_MPS_AMP_MIN_ROWS is set. With that, answers match main on CPU and MPS for both cached checkpoints. Also tests the 'laya unknown' branch, which the pinned version lookup had stopped covering.
Laya reads a value it cannot parse as its default of 5, which runs requests of five or more questions in fp16 on MPS. from_env saw the variable set and left that in place, so a mistyped value meant to keep fp32 did the opposite. It is now a configuration error, raised before the checkpoint loads.
|
@cacheline999 could you add a Demo / evidence section following the self-review guidance? The CPU/MPS measurements already answer the right questions; the missing piece is a reproducible artifact behind the tables. Please link the comparison script and sanitized per-run outputs for the same ticket inputs on baseline/head: model loading, one- and six-question answers, and the default MPS setting versus LAYA_MPS_AMP_MIN_ROWS=5. Include the exact source/checkpoint revisions and invocation, and retain the two changed multilingual decisions in the override results. Keep the existing cold-load/background-load notes and untested configurations explicit. Existing logs or a terminal trace are enough; no video or broader hardware sweep is needed. Please use safe sample data and redact credentials/private paths from anything shared. |
9bce57b to
ccd7903
Compare
docs/results/laya-upgrade: run.sh checks out two commits, installs each one's lock, runs --model laya's backend over the 30 ticket-router tickets (one choice, one noul and six questions) on both checkpoints and devices, adds the LAYA_MPS_AMP_MIN_ROWS=5 and no-skip controls, and prints the comparison. The README gives the steps; records/ holds the 28 outputs of the recorded run of 222e656 against c197fe6.
ccd7903 to
82f7f21
Compare
|
Added, see the Demo / evidence section. |
Why
Closes #28. #22 wrapped
laya.loadin transformers'no_init_weightsto skip the throwaway random weight init. Laya does this itself from 0.3.9, so the wrapper and its test stubs are no longer needed.Moving the lock past 0.3.9 brings one more change. From 0.3.10 Laya answers a request of five or more questions in fp16 on MPS. In s1a only the browser front asks that many at once, on pages that offer all four operations, so one episode would mix fp32 and fp16 steps. On the 30 ticket-router tickets with six questions, fp16 moved
multilingualprobabilities by up to 0.051 and changed two of 180 decisions, both close to even (records/head-multilingual-mps-override5.jsonbelow). Keeping fp32 can cost speed: the one M5 report (NandhaKishorM/laya#109) has fp16 5 ms faster at five questions.How
LayaModel.from_envno longer wrapslaya.load, since laya skips the init itself. After loading, it raises the agent'smps_amp_min_rowsso every MPS request stays in fp32, the precision main runs today, unlessLAYA_MPS_AMP_MIN_ROWS(Laya's own variable) is set; a value that is not a whole number is a configuration error, because Laya would read it as its default of 5. A laya without the attribute (0.3.9) is left as it is.Open first:
s1a/decision_models/laya.py(from_env), thentests/test_decision_models_laya.py.What
layaextra islaya>=0.3.9.uv lock --upgrade-package layamoves the lock from 0.3.5 to 0.3.20, the release system1-omni's Laya worker runs; only the laya entry changes. 0.3.26 is the latest today; say if you would rather lock that.without_weight_init()and itstransformers.initializationtest stubs are gone.--model layakeeps MPS requests in fp32 unlessLAYA_MPS_AMP_MIN_ROWSis set. Nothing changes on CPU or CUDA. Callers change nothing; settingLAYA_MPS_AMP_MIN_ROWS=5restores Laya's default.system_onetest now pins the version, so it passes with the laya extra installed too).docs/configuration.mdand.env.exampledescribe the variable.docs/results/laya-upgrade/: the comparison below, as a README with the steps,run.shand the recorded outputs.Verification
Run by me on
82f7f21, M1 Pro, Python 3.14, with the laya extra (0.3.20) unless noted.uv run ruff format --check . && uv run ruff check . && uv run ty check: all pass.uv run pytest -qandscripts/smoke.sh: 492 passed, 41 skipped;smoke: ok(core install). In a fresh environment whose packages are not byte-compiled yet, Python 3.14 printsSyntaxWarnings from pysbd and openjiuwen on stderr, which fails smoke's one-line check fordecideandtest_s1a_home_from_the_dotenv_file_places_the_logs_and_the_runtime_root. Main fails the same way there;uv sync --compile-bytecodeavoids it on both.CHANGELOG.mdand the docs say what the code does now.CI (core install, no laya extra) passes on
82f7f21.Demo / evidence
Task.
--model laya's backend,LayaModel.from_env(), answers the 30 tickets ins1a/agents/_data/ticket_router_eval.jsonl: one choice, one noul and one six-question request each. Success is every choice, probability, confidence and noul value equal between main and this branch.Actions.
docs/results/laya-upgrade/README.mdhas the steps andrun.shruns them: it checks out222e656(main, laya 0.3.5) andc197fe6(this branch's code, laya 0.3.20; the commits after it only adddocs/results/laya-upgrade), installs each lock, and runs both checkpoints on CPU and MPS, three fresh processes per side, alternating. Two controls: this branch withLAYA_MPS_AMP_MIN_ROWS=5on MPS, and main without its weight-init skip on CPU.Result. An actual run by me, 2026-10-05, M1 Pro, macOS 26.1, checkpoint revision
7b928d8; the 28 outputs are inrecords/, each naming its commit, versions, device and Laya variables.python compare.py records/printed:identicalbetween main and this branch, 3 + 3 runs each;LAYA_MPS_AMP_MIN_ROWS=5:multilingualchangedt_6cb98186dde0(pick, logistics to returns) andt_f5251390fb88(cancel, 0.4959 to 0.5074), largest probability change 0.051;english0.010, no decision changed;english) and 42.5 s (multilingual), against 6 to 10 s on either side with it.Limits. Load and latency columns compare the two sides of a row only: the machine had other work running (1-min load about 6 to 10), and the first
englishCPU load of each side read the model files from disk. Not covered: thetyped-decisionscheckpoint, a local checkpoint path, transformers 4.x, CUDA, and Apple chips other than the M1 Pro.