fix(doctor): budget boot stages on guest work, not host steal - #269
Merged
Merged
Conversation
The parser's eleven tests were spread over three places in the agent's 1280-line tests.rs, which sat at its size ratchet. They move to boot_timing/tests.rs and the ratchet drops to 1115. Adds a test that a line carrying capsem-init's new steal_ms field, sane or hostile, still yields its stage: the field is for the in-guest doctor and BootStage does not carry it.
The guest only accounts steal time when KVM advertises it in CPUID 0x40000001. setup_cpuid passes KVM_GET_SUPPORTED_CPUID through and only rewrites topology leaves; this pins that the paravirt leaves survive, so the boot budget's steal subtraction keeps a source on x86_64.
Release run 36351667646 failed test_boot_stages_within_budget on profile_root_seed at 510ms; the stage copies sixteen small files and measured 40ms on the previous run. On a shared nested-virt runner the vCPU was descheduled for most of it. capsem-init's boot_mark now records steal_ms per stage: the largest per-vCPU delta of the /proc/stat steal column (USER_HZ ticks), so vCPUs waiting in parallel cannot add up to more than one vCPU's wall time. One awk reads /proc/uptime and /proc/stat per mark (busybox-compatible; the kernel line goes through the same mark). assess_boot_timing budgets duration_ms minus steal_ms, capped at the duration. A missing field means 0, and anything that is not a finite non-negative integer discounts nothing, so a hostile or corrupt line can only make the budget stricter. Slow stages are reported raw, with their steal. The field stays in the guest: BootStage and the IPC schema hash are unchanged, and the agent ignores it.
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #269 +/- ##
=========================================
- Coverage 65.4% 65.3% -0.2%
=========================================
Files 1450 1452 +2
Lines 127617 127936 +319
Branches 91731 91872 +141
=========================================
+ Hits 83554 83604 +50
- Misses 39120 39336 +216
- Partials 4943 4996 +53
Flags with carried forward coverage won't be shown. Click here to find out more.
🚀 New features to boost your workflow:
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Release run 36351667646 failed test_boot_stages_within_budget on
profile_root_seed at 510ms; the stage copies sixteen small files and
measured 40ms on the previous run. On a shared nested-virt runner the
vCPU was descheduled for most of it.
capsem-init's boot_mark now records steal_ms per stage: the largest
per-vCPU delta of the /proc/stat steal column (USER_HZ ticks), so vCPUs
waiting in parallel cannot add up to more than one vCPU's wall time.
One awk reads /proc/uptime and /proc/stat per mark (busybox-compatible;
the kernel line goes through the same mark).
assess_boot_timing budgets duration_ms minus steal_ms, capped at the
duration. A missing field means 0, and anything that is not a finite
non-negative integer discounts nothing, so a hostile or corrupt line can
only make the budget stricter. Slow stages are reported raw, with their
steal. The field stays in the guest: BootStage and the IPC schema hash
are unchanged, and the agent ignores it.
Blocked the 0.6.4 stable release: run 36351667646, x86_64 KVM on a shared 4-vCPU nested runner with the ironbank suite running 4 VMs, failed only
test_boot_stages_within_budget:profile_root_seed510 ms > 500 ms. The same code measured it at 40 ms in run 36290484745; the stage copies 16 files (100K). The release process refuses an unchanged retry, so this is the fix forward.Tests: Python doctor/initrd suites 454 passed; citadel 1255; capsem-agent musl tests; clippy agent (musl+host) and capsem-core; capsem-proto and guest_report tests.
🤖 Generated with Claude Code