Soak test: long-run daemon health tripwires (#101) - #103
Merged
Merged
Conversation
One compressed daemon lifetime: 3,000 scheduler ticks (~10 simulated days) via the fake clock against real SQLite — one watch with the full stage-2 kit (digest/stale/flaky, never completing), a conclusion flip every 500 ticks and a new commit every 1,500 so every watch kind and the statemachine transitions fire repeatedly. Budgets are generous tripwires, not benchmarks; widening one requires a finding, not a fix: - steady-state traced-memory growth < 1 MB; traced peak < 10 MB - last-200-tick mean latency ≤ 3x first-200 mean (+5ms) — no drift with accumulated history - check-observed retention ≤ OBSERVATION_RETENTION after the whole run - SQLite < 10 MB for the simulated fortnight - digests ≈ simulated days; stale and flaky both fire; the watch ends ACTIVE and schedulable Measured ~7s with tracemalloc active (~3s untraced); the suite grows accordingly. 708 pass; demo 8/8.
7 tasks done
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Implements #101 — the first guard for slow failure modes now that the daemon really runs for days. One offline test compresses a ~10-simulated-day daemon lifetime (3,000
RunnerDaemon.tick()calls, fake clock, real SQLite) with the full stage-2 watch kit, periodic conclusion flips and new commits, and asserts the budgets from the issue: bounded steady-state memory (<1 MB growth, <10 MB peak, tracemalloc after warmup), no per-tick latency drift (last-200 mean ≤ 3× first-200 mean), observation retention ≤ 20 after the whole run, DB < 10 MB, and functional sanity (digests ≈ days, stale + flaky fire, task stays ACTIVE).Budgets are tripwires — widening any requires a finding, not a fix. Measured ~7s with tracing active (~3s untraced); suite now 708 passing, demo 8/8.