Skip to content

cuda: declare the qwen3_5 backend - #77

Open
xiaoyu-xyz wants to merge 6 commits into
ThinkFlowLab:mainfrom
xiaoyu-xyz:cuda-qwen3_5-manifest
Open

xiaoyu-xyz wants to merge 6 commits into
ThinkFlowLab:mainfrom
xiaoyu-xyz:cuda-qwen3_5-manifest

Conversation

@xiaoyu-xyz

@xiaoyu-xyz xiaoyu-xyz commented Oct 4, 2026 •

Copy link
Copy Markdown

Purpose

#19 added the kernels under src/backends/cuda/qwen3_5/ without declaring them, so the directory has kernel sources and no manifest. #25 defines the manifest and its checker, and against main the checker reports exactly that:

[error] qwen3_5: has kernel sources (attention.cu, common.cuh, elementwise.cu)
        but no qwen3_5.backend.json manifest

Rebased onto current main.

The change has three parts:

  1. The manifest. Values come from the merged files rather than from me: build.sh writes libqwen3_5_cuda.so and defaults arch to CUDA_COMPUTE_CAP-or-89 in a single -gencode flag, and the directory README says the tensor-core kernels need sm_80 or newer.
  2. The ABI version. The first commit said abi_version: 1; ops.h defines CS1_ABI_VERSION 4. Nothing caught that until cuda: add the backend contract and its checker #25's checker learned to read the header macro, which is what surfaced it.
  3. The CI step. ci.yml gains a cuda-contract job: the checker's own tests, the repository's CUDA tests under tests/cuda, and check_contract.py --repo-root .. No CUDA toolkit, no GPU.

status is experimental, not validated. #19 validated the kernels, but a validated manifest asserts a tolerance and a reference entrypoint, and those are twu3202's to state rather than mine. Promoting it once a tolerance is recorded is a one-line change.

Merge order

This depends on #25. The checker exits 1 on a tree with kernels and no declaration, which is main's state without this manifest, so the job cannot land before it. The checker's own tests and check_contract.py also arrive in #25, so those two steps skip until it lands — the job comments say so, and the notice names #25 rather than reporting a bare missing file. The tests/cuda step is main's own tree and runs in either order.

Test Result

Both merge orders, run from the commit rather than from a working tree:

The manifest itself: it parses, all nine declared sources exist, and the directory holds no .cu, .cuh or .h file the manifest omits — that last part by hand, since the checker does not compare a directory against sources. The ABI check cross-checks ops.h's CS1_ABI_VERSION 4 against abi_version: 4.

Deliberately not included: numerics and reference. Both are about a parity claim this does not make.

xiaoyu-xyz added a commit to xiaoyu-xyz/system1-omni that referenced this pull request Oct 5, 2026
The job failed on its own branch, in five seconds:

    ImportError: Start directory is not importable: 'src/backends/cuda/tests'

The checker and its tests live in ThinkFlowLab#25, so on ThinkFlowLab#77's tree there is nothing to run.
The guard I wrote checked for `check_contract.py` alone and ran *after* the test
step, so the first step failed before the guard was reached and the message
explaining the dependency never printed.

Now a first step decides whether the prerequisites are present and the three
steps run only if they are; otherwise the job reports a notice naming ThinkFlowLab#25 and
passes. That is what makes either merge order work: green on this branch before
ThinkFlowLab#25 lands, and running by itself once it does.

Verified both ways: with ThinkFlowLab#25's files absent all three steps skip, and with them
present 66 checker tests, 66 kernel tests and the contract check all pass.
@Levius-Fubuki

Copy link
Copy Markdown
Collaborator

Checked d1a0a12 with #25: 66 checker tests pass; current-main manifests yield zero errors and one disclosed Laya ABI warning. One CI nit: the new kernel-test loop scans beside kernels and misses root tests/cuda (8 CPU tests). Please use the repository test layout. No CUDA compilation or GPU parity was rerun.

ThinkFlowLab#19 added the kernels under src/backends/cuda/qwen3_5/ but left them
undeclared, so the directory has kernel sources and no manifest. ThinkFlowLab#25 defines the
manifest and its checker, and running that checker against main reports exactly
that:

    [error] qwen3_5: has kernel sources (attention.cu, common.cuh,
            elementwise.cu) but no qwen3_5.backend.json manifest

This adds the declaration, with the values the merged files already state:

- sources: the nine kernel and header files, from the directory listing.
- build.script and build.output: what build.sh actually writes,
  libqwen3_5_cuda.so.
- build.default_arch 89 and build.architectures [89]: build.sh defaults `arch` to
  CUDA_COMPUTE_CAP-or-89 and passes it to a single -gencode flag.
- build.min_capability 80: the README says "Tensor-core kernels need sm_80 or
  newer".

status is `experimental`, not `validated`. ThinkFlowLab#19 validated the kernels, but this
manifest is written by someone other than their author and a `validated` manifest
asserts a tolerance and a reference entrypoint, which are twu3202's to state. A
declaration that the backend exists and how it builds is useful on its own and
does not put words in anyone's mouth; promoting it is a one-line change once the
tolerance is recorded.

Checked with the checker from ThinkFlowLab#25: `1 backend(s), 0 error(s), 0 warning(s)`, and
with discover_buildable.py from #2, which reports it as compile-checkable.
The manifest said abi_version 1. ThinkFlowLab#19's ops.h defines CS1_ABI_VERSION 4, so the
declaration disagreed with the library it describes.

Nothing caught it when this was written, because the checker did not read
headers. ThinkFlowLab#25 now does: it reads the `#define <PREFIX>_ABI_VERSION N` out of the
declared sources and requires the manifest to match, which is what surfaced this.

    abi_version is 1 but ops.h defines CS1_ABI_VERSION 4

also strengthens the description in the same PR.
Nothing ran `src/backends/cuda/tests` or `check_contract.py`: the only Python
discovery in this workflow is `tests/benchmarks`.

Two things make this land here rather than with the checker. The step needs
`check_contract.py`, which arrives in ThinkFlowLab#25. And the checker exits 1 on a tree that
has kernels but no declaration for them, which is main's state until the manifest
in this branch lands — so wiring it up first would redden main. Both are stated
in the job's comment and in its failure message, so a wrong merge order is
self-explanatory rather than a bare "file not found".

No CUDA toolkit and no GPU: the checker and its tests read files as text.
The first version ran `discover -s src/backends/cuda/tests`, which finds the
contract checker's tests and nothing else. Kernel tests sit beside their kernels
-- `src/backends/cuda/scoring/test_scoring.py` is 17 of them -- so a backend that
merges later would have had no test coverage at all while appearing to be
covered.

The step now loops over `src/backends/cuda/*/` and discovers in any directory
holding `test_*.py`. `-t` has to point at the directory itself rather than the
repository root: the kernel directories are not packages, so `-t .` fails with
"Start directory is not importable".

Simulated on a tree carrying both: 17 from `scoring/`, 66 from `tests/`.
The job failed on its own branch, in five seconds:

    ImportError: Start directory is not importable: 'src/backends/cuda/tests'

The checker and its tests live in ThinkFlowLab#25, so on ThinkFlowLab#77's tree there is nothing to run.
The guard I wrote checked for `check_contract.py` alone and ran *after* the test
step, so the first step failed before the guard was reached and the message
explaining the dependency never printed.

Now a first step decides whether the prerequisites are present and the three
steps run only if they are; otherwise the job reports a notice naming ThinkFlowLab#25 and
passes. That is what makes either merge order work: green on this branch before
ThinkFlowLab#25 lands, and running by itself once it does.

Verified both ways: with ThinkFlowLab#25's files absent all three steps skip, and with them
present 66 checker tests, 66 kernel tests and the contract check all pass.
The loop scanned src/backends/cuda/*/ for test files, which matched the
checker's own tests/ directory -- running those 66 tests a second time --
and missed tests/cuda. CONTRIBUTING.md keeps test bodies under the root
tests/ tree, so use the same invocation as the benchmarks job.
@xiaoyu-xyz
xiaoyu-xyz force-pushed the cuda-qwen3_5-manifest branch from d1a0a12 to b1a886a Compare October 5, 2026 13:02
@xiaoyu-xyz

Copy link
Copy Markdown
Author

Fixed in b1a886a, rebased onto current main. The loop scanned src/backends/cuda/*/, so it matched the checker's own tests/ directory — running those 66 tests a second time — and never reached tests/cuda. It now uses the repository layout, the same invocation as the benchmarks job: python -m unittest discover -s tests/cuda -p 'test_*.py' -v.

tests/cuda is already in main, so that step runs in either merge order; only the checker steps stay behind the present gate. Verified from the commit in both orders, and CI on b1a886a ran the 8 tests.

Related, in #25 (184de43): a flat backend was taken to own the whole src/backends/cuda/ directory, so laya.backend.json exempted qwen3_5/ from the undeclared-backend check. Current main now reports that error, which is the state this PR clears.

@xiaoyu-xyz
xiaoyu-xyz marked this pull request as ready for review October 5, 2026 13:22

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants