Repository navigation
qwen3_5: retain multimodal Graph regressions in workspace tests - #106
Open
Levius-Fubuki wants to merge 4 commits into
Open
Levius-Fubuki wants to merge 4 commits into
Levius-Fubuki wants to merge 4 commits into
Conversation
4 tasks done
Levius-Fubuki
force-pushed
the
codex/cua-native-mm-graph-tests
branch
from
October 7, 2026 03:21
975201c to
2dd88db
Compare
4 tasks done
Levius-Fubuki
marked this pull request as ready for review
October 8, 2026 01:11
4 tasks done
This was referenced Oct 8, 2026
1 of 10 tasks
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Purpose
Retain native multimodal language CUDA Graph regressions under root
tests/, addressing #103’s maintenance review. Three opt-in GPU cases and the capture-failure runner cover ordered cache keys, current IDs/images/positions, multimodal interleave, scratch growth, FIFO eviction and eager fallback.Depends on #103. The main-based diff includes its production implementation; this PR’s independent addition is the regression suite and four lines of private test-module wiring.
Submission scope
Core implementation and directly related tests are retained. Ordinary documentation, experiment reports/media and generated run outputs remain outside this submission. Test fixtures, build inputs, deployment data and third-party licensing remain as needed.
Validation
After cleaning workspace debug artifacts before each branch, fresh local fmt, Clippy/all-targets with warnings denied and 90 Rust tests pass (16 opt-in tests ignored). Independent GitHub CI passes on this exact head, including release workspace build and strict Docs. Restored test bodies/fixtures match the saved originals. Production execution is unchanged. GPU tests were not rerun; the server remains shut down.
Head
968a9be2763d0cf79a8935b89c493f562ded92e2; 3 changed files and 532 changed source/test/build lines, excluding licenses, notices, locks and static data.Self-review
Re-reviewed complete test-only commit and parent integration against current contributor/architecture rules. No generated run output is committed. Full agent-assisted self-review is complete under the author’s explicit Ready instruction. Retain the #103 dependency and rebase after that parent merges.
2026-10-08 verification follow-up
Validated source head:
2dd88db234b99a64061dbfd0094a4ab6b0b4e5cb. Upstream main snapshot:4a79980d; GitHub reports no merge conflict. Fresh exact-head fmt, strict workspace Clippy, workspace tests and release build pass: 90 passed, 0 failed, 16 ignored. Ignored tests are counted separately from explicit GPU execution. Exact commands and logs are in the archive. Exact test head2dd88db2: 3/3 release GPU cases plus both unconditional failure diagnostics pass on RTX 4090 with tracing unset. No assertion or tolerance changed. This closes the current-head GPU execution gap. Merge #103 first, then rebase to leave the dedicated test increment.Raw verification evidence, SHA-256
5df66bf3974d534875946168fba4d7fe543d08edbd658c3756f17d6f0c8f5773. Includes exact source manifests, environment/library/binary hashes, complete logs, synthetic reference/HTTP results and initial setup failures. CUDA execution is serial on one RTX 4090. No weights, credentials or private inputs are included. Contributor/maintainer review and dependency merges remain separate.Full self-review completed / Ready — 2026-10-08
Current head:
2dd88db234b99a64061dbfd0094a4ab6b0b4e5cb; merge base:99865743d27316fbe81362dc0f0e6de6fda86284. Full cumulative diff: 532 authored-code lines / 579 total diff lines. The author explicitly requested completed self-review and Ready conversion. The completed checklist records this agent-assisted review and its verification; it does not represent maintainer approval.Root integration review plus three independent component reviews covered all changed source, tests, helpers, manifests and documentation: CPU contract/checkpoint/tokenizer/calibration; serial worker/admission/GPU head and error retirement; ordered bounded packing and response reconstruction; Graph lifetime/cache/health/options; ABI7 CUDA continuations/fixed GEMM/request-local prefix state; root regressions, recipe reproduction, fixture consumers and evidence/claim boundaries. No functional blocker remains. For #112/#113 above 3,000 authored lines, this is full component/integration self-review, with the existing split rationale and review order retained.
Corrected stale unsupported-Graph and no-worker-prefix-consumer documentation where applicable. Removed unused CPU fixture run/duplicate fields on the Decider stack, preserving all 255 label IDs and all 12 used fixture values exactly; the original removed data remain in evidence outside committed source. Both pinned CPU contract opt-ins pass on every repaired Decider branch; repaired docs pass strict MkDocs. All five initial review heads passed fresh fmt, strict workspace Clippy, workspace tests, release build and strict docs; latest-head CI Rust/benchmark/docs also passes. Ignored hardware/model tests are separate from executed tests.
Runtime/CUDA/GPU regression code is unchanged by this cleanup. Earlier personally generated GPU/reference/HTTP campaigns retain their original source revisions; independent reviewers rehashed both source bindings, replayed all stored numerical gates and fixed/shared exact controls, and recomputed published timings.
ready-source-binding.jsonbinds these runtime checks to the latest heads and separately proves the used CPU oracle values are unchanged. No new CUDA/HTTP/performance campaign was run after server shutdown; the server remains off. Existing hardware, scope and performance limitations remain disclosed.Complete self-review reports and fresh checks; SHA-256
1edbd4a6e13576d4cc23638b7ec1e35c1248d6a712b8a6406e1f269ca685cb9d. Includes all four review reports, closure summary, original reviewed full diffs, new CPU/docs logs, latest-head CI, source/fixture binding and archived removed fields. Merge dependencies remain #103 before #106, and #94 → #110 → #111 → #112 → #113, including #97/#98/#99 prerequisites for #113; Ready conversion does not merge them.