fix(qwen3_8): use cached checkpoint offline in E2E test - #1563
zhenshanx-nv wants to merge 1 commit into
Conversation
Signed-off-by: Zhenshan Xie <zhenshanx@nvidia.com>
|
Navigate logical layers of code changes, visualize relationships, and explore their blast radius. No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Repository: NVIDIA/TensorRT-Model-Connect/.coderabbit.yaml Review profile: CHILL Plan: Enterprise Run ID: 📒 Files selected for processing (1)
Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 11 remain after this review. 📝 SummarySummary
Architecture impact
WalkthroughThe Qwen3-8 end-to-end test helper now requests checkpoint snapshots with ChangesQwen3-8 checkpoint resolution
Priority: ⬇️ Low Estimated code review effort: 1 (Trivial) | ~3 minutes Change: Bug fix Merge Risk: ⚪ Minimal · up to Qwen3.8 offline tests now request the locally staged checkpoint, matching the repository-controlled CI setup. No concrete merge blocker is established. 🚥 Pre-merge checks | ✅ 8 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (8 passed)
Comment |
|
This is an automated Internal CI result; no review from an individual maintainer is requested. Open the public Source Actions run from the automated status link above. |
Background
qwen3_8's premerge E2E test fails under Community GPU CI's offline mode.
_checkpoint()infamilies/qwen3_8/tests/test_e2e.pycallssnapshot_download()withoutlocal_files_only=True. Once the family container switches toHF_HUB_OFFLINE=1(after the trusted host coordinator has already staged the checkpoint locally), the call still tries to list the repo tree via the Hugging Face Hub API first, which raiseshuggingface_hub.errors.OfflineModeIsEnabled— even though the checkpoint is already present in the local cache. Every other family's equivalent checkpoint helper (e.g.families/bert/tests/test_e2e.py,families/sam3/tests/test_e2e.py) already passeslocal_files_only=True;qwen3_8was the only one missing it.Exit Criteria
HF_HUB_OFFLINE=1, with no network call.Implementation
local_files_only=Trueto thesnapshot_download()call in_checkpoint()(families/qwen3_8/tests/test_e2e.py), matching the existing pattern already used byfamilies/bert/tests/test_e2e.pyandfamilies/sam3/tests/test_e2e.py.Change categories
Validation
Commands and Results
I don't have a local dev environment with the pinned dependencies/CUDA toolchain in this session to run
pytest/ruffdirectly, so validation here is based on direct inspection of live Community GPU CI runs rather than a local test invocation:ci/developer-dispatched Dev Community CI run: traceback showedhuggingface_hub.errors.OfflineModeIsEnabledraised fromfamilies/qwen3_8/tests/test_e2e.py:110inside_checkpoint().ci/developeras fix(qwen3_8): use cached checkpoint offline in E2E test #1407 and matches the long-standing working pattern already used bybert/sam3onmain.mainspecifically, since family test files are sourced from each PR's own merge commit againstmain, not fromci/developer— this PR is what makes that verification possible.Hardware, Environment, and Revisions
e5475571707d1fb4e7f9f5805fcb82794f308c5Qwen/Qwen3.8-27Bat the revision pinned infamilies/qwen3_8/tests/manifestsNot Run / Remaining Gaps
main; recommend validating via a fresh Community GPU CI dispatch afterward.Contributor Self-Review
Notes For Future Readers
ci/developeras fix(qwen3_8): use cached checkpoint offline in E2E test #1407. Family-owned test files (families/*/tests/test_e2e.py) are checked out from each PR's ownmain-based merge commit during Community GPU CI, not fromci/developer— so a fix like this only takes effect for real PRs once it lands onmaindirectly.qwenfamily hit the identical underlying bug, butmainalready carries a different fix for it (local_files_only=constants.HF_HUB_OFFLINE, from feat(build): add optional native Edge-LLM SDK #1305) — that one does not need porting here.Risk level
Risk rationale: single-parameter addition to a test-only helper function, matching an existing pattern already used elsewhere in the codebase; no production runtime or API change.