Skip to content

Add CUDA resource ownership - #47

Open
linear3735 wants to merge 4 commits into
ThinkFlowLab:mainfrom
linear3735:codex/laya-runtime
Open

linear3735 wants to merge 4 commits into
ThinkFlowLab:mainfrom
linear3735:codex/laya-runtime

Conversation

@linear3735

@linear3735 linear3735 commented Sep 30, 2026 •

Copy link
Copy Markdown

Purpose

Add Rust ownership of a CUDA device, stream and buffers. Check copy bounds and keep resources alive until pending work completes. CPU builds do not link CUDA.

Part of #14. This is the resource layer for Laya; model execution comes separately. Core and configuration changes: 271 lines.

Tests now live under the repository-root tests/ directory. Cargo target names and test coverage are unchanged.

Test Plan

cargo fmt --all --check
cargo clippy --workspace --locked --all-targets -- -D warnings
cargo test --workspace --locked
cargo build --workspace --release --locked

On H800, check allocation, a 4 KiB copy roundtrip and buffer lifetime after dropping the context handle.

System1-Omni Version / Commit: a4787213f86f82c0986df2d6599a2fb8ab4e5de8.

Test Result

Merged main and resolved the workspace conflict by retaining both Laya and CUDA members. Runtime source is unchanged.

Formatting, strict Clippy and release build passed locally. The workspace passed 31 CPU tests; four GPU tests and two checkpoint-dependent CPU tests were skipped. All seven benchmark harness tests passed. The existing H800 resource test passed against the same runtime source; GPU tests were not rerun for this merge.

Rust CI, Docs build and benchmark harness tests passed for this commit.

Self-review

Before marking this PR ready for review or requesting maintainer review, complete
the self-review checklist.
Keep the PR in draft while this work is incomplete.
For agent assistance, use the optional precheck-pr skill.

  • I have reviewed the full diff and addressed the issues I found.
  • I have checked that the change follows the project's architecture and stays focused on the stated purpose.
  • I have run the checks appropriate to this change and reported commands, results, and anything I could not verify above.
  • I have checked that the PR description, documentation, and any accuracy or performance claims match the implementation and available evidence.

@hsliuustc0106 hsliuustc0106 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Independent local review — CUDA resource ownership

Verdict: approved, with one coordination note.

Verified locally at head a4787213: the CPU-fixture suite runs 6/6 — copy bounds, error-path synchronization, partial-resource release on failed allocations/uploads, buffers keeping the library and stream alive after Cuda is dropped, device restore on drop, and ABI-gated rejection of incompatible libraries — with the real-GPU round-trip properly #[ignore]d behind LAYA_CUDA_LIBRARY/LAYA_CUDA_DEVICE. Clippy -D warnings clean; [[test]] registered at repo-root tests/backends/cuda/ per CONTRIBUTING; CPU builds not linking CUDA matches the repo's dynamic-loading discipline. Copy bounds and keep-alive-until-pending-completes match the architecture contract's backend/lifetime rules.

Coordination note: this introduces a Rust bindings crate under src/backends/cuda alongside qwen3_5's model-crate loading — consistent with the target design's "Rust bindings/dispatch with CUDA C++ kernels", but when #25's backend-manifest contract lands, this directory needs a manifest and the two binding patterns should be reconciled (see my #25 review).

Approval per the repo review process; reflects head a4787213 only and asserts no GPU execution.

@hsliuustc0106 hsliuustc0106 mentioned this pull request Oct 5, 2026
2 of 4 tasks
@hsliuustc0106

Copy link
Copy Markdown
Contributor

fix conflicts

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants