Skip to content

qwen3_5: retain multimodal Graph regressions in workspace tests - #106

Open
Levius-Fubuki wants to merge 4 commits into
ThinkFlowLab:mainfrom
Levius-Fubuki:codex/cua-native-mm-graph-tests
Open

Levius-Fubuki wants to merge 4 commits into
ThinkFlowLab:mainfrom
Levius-Fubuki:codex/cua-native-mm-graph-tests

Conversation

@Levius-Fubuki

@Levius-Fubuki Levius-Fubuki commented Oct 6, 2026 •

Copy link
Copy Markdown
Collaborator

Purpose

Retain native multimodal language CUDA Graph regressions under root tests/, addressing #103’s maintenance review. Three opt-in GPU cases and the capture-failure runner cover ordered cache keys, current IDs/images/positions, multimodal interleave, scratch growth, FIFO eviction and eager fallback.

Depends on #103. The main-based diff includes its production implementation; this PR’s independent addition is the regression suite and four lines of private test-module wiring.

Submission scope

Core implementation and directly related tests are retained. Ordinary documentation, experiment reports/media and generated run outputs remain outside this submission. Test fixtures, build inputs, deployment data and third-party licensing remain as needed.

Validation

After cleaning workspace debug artifacts before each branch, fresh local fmt, Clippy/all-targets with warnings denied and 90 Rust tests pass (16 opt-in tests ignored). Independent GitHub CI passes on this exact head, including release workspace build and strict Docs. Restored test bodies/fixtures match the saved originals. Production execution is unchanged. GPU tests were not rerun; the server remains shut down.

Head 968a9be2763d0cf79a8935b89c493f562ded92e2; 3 changed files and 532 changed source/test/build lines, excluding licenses, notices, locks and static data.

Self-review

Re-reviewed complete test-only commit and parent integration against current contributor/architecture rules. No generated run output is committed. Full agent-assisted self-review is complete under the author’s explicit Ready instruction. Retain the #103 dependency and rebase after that parent merges.

  • I have reviewed the full diff and addressed the issues I found.
  • I have checked that the change follows the project's architecture and stays focused on the stated purpose.
  • I have run the checks appropriate to this change and reported commands, results, and anything I could not verify above.
  • I have checked that the PR description, documentation, and any accuracy or performance claims match the implementation and available evidence.

2026-10-08 verification follow-up

Validated source head: 2dd88db234b99a64061dbfd0094a4ab6b0b4e5cb. Upstream main snapshot: 4a79980d; GitHub reports no merge conflict. Fresh exact-head fmt, strict workspace Clippy, workspace tests and release build pass: 90 passed, 0 failed, 16 ignored. Ignored tests are counted separately from explicit GPU execution. Exact commands and logs are in the archive. Exact test head 2dd88db2: 3/3 release GPU cases plus both unconditional failure diagnostics pass on RTX 4090 with tracing unset. No assertion or tolerance changed. This closes the current-head GPU execution gap. Merge #103 first, then rebase to leave the dedicated test increment.

Raw verification evidence, SHA-256 5df66bf3974d534875946168fba4d7fe543d08edbd658c3756f17d6f0c8f5773. Includes exact source manifests, environment/library/binary hashes, complete logs, synthetic reference/HTTP results and initial setup failures. CUDA execution is serial on one RTX 4090. No weights, credentials or private inputs are included. Contributor/maintainer review and dependency merges remain separate.

Full self-review completed / Ready — 2026-10-08

Current head: 2dd88db234b99a64061dbfd0094a4ab6b0b4e5cb; merge base: 99865743d27316fbe81362dc0f0e6de6fda86284. Full cumulative diff: 532 authored-code lines / 579 total diff lines. The author explicitly requested completed self-review and Ready conversion. The completed checklist records this agent-assisted review and its verification; it does not represent maintainer approval.

Root integration review plus three independent component reviews covered all changed source, tests, helpers, manifests and documentation: CPU contract/checkpoint/tokenizer/calibration; serial worker/admission/GPU head and error retirement; ordered bounded packing and response reconstruction; Graph lifetime/cache/health/options; ABI7 CUDA continuations/fixed GEMM/request-local prefix state; root regressions, recipe reproduction, fixture consumers and evidence/claim boundaries. No functional blocker remains. For #112/#113 above 3,000 authored lines, this is full component/integration self-review, with the existing split rationale and review order retained.

Corrected stale unsupported-Graph and no-worker-prefix-consumer documentation where applicable. Removed unused CPU fixture run/duplicate fields on the Decider stack, preserving all 255 label IDs and all 12 used fixture values exactly; the original removed data remain in evidence outside committed source. Both pinned CPU contract opt-ins pass on every repaired Decider branch; repaired docs pass strict MkDocs. All five initial review heads passed fresh fmt, strict workspace Clippy, workspace tests, release build and strict docs; latest-head CI Rust/benchmark/docs also passes. Ignored hardware/model tests are separate from executed tests.

Runtime/CUDA/GPU regression code is unchanged by this cleanup. Earlier personally generated GPU/reference/HTTP campaigns retain their original source revisions; independent reviewers rehashed both source bindings, replayed all stored numerical gates and fixed/shared exact controls, and recomputed published timings. ready-source-binding.json binds these runtime checks to the latest heads and separately proves the used CPU oracle values are unchanged. No new CUDA/HTTP/performance campaign was run after server shutdown; the server remains off. Existing hardware, scope and performance limitations remain disclosed.

Complete self-review reports and fresh checks; SHA-256 1edbd4a6e13576d4cc23638b7ec1e35c1248d6a712b8a6406e1f269ca685cb9d. Includes all four review reports, closure summary, original reviewed full diffs, new CPU/docs logs, latest-head CI, source/fixture binding and archived removed fields. Merge dependencies remain #103 before #106, and #94 → #110 → #111 → #112 → #113, including #97/#98/#99 prerequisites for #113; Ready conversion does not merge them.

@Levius-Fubuki
Levius-Fubuki force-pushed the codex/cua-native-mm-graph-tests branch from 975201c to 2dd88db Compare October 7, 2026 03:21
@Levius-Fubuki
Levius-Fubuki marked this pull request as ready for review October 8, 2026 01:11
Copilot AI balanced review requested due to automatic review settings October 8, 2026 01:11

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@Levius-Fubuki Levius-Fubuki changed the title qwen3_5: retain multimodal Graph regressions in workspace tests cua_s1: retain native multimodal language Graph core Oct 8, 2026
@Levius-Fubuki Levius-Fubuki changed the title cua_s1: retain native multimodal language Graph core qwen3_5: retain multimodal Graph regressions in workspace tests Oct 8, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants