Skip to content

usage.get and simlock stats: the figures from the event history, on a worker and a gateway #347

Description

@V3RON

Part of #329.

Scope

After this PR, simlock stats prints the usage figures for a window, as a table and as --json, on a worker and on a gateway, from the event history and nothing else. The programmatic client gets usage.get. ADR 0016 §1, §2, §5, §6, §7 and §8 fix how the figures are computed.

Technical spec

Modules touched

  • src/core/usage/ (new, pure, platform-agnostic) — computeUsage(events, window, options) takes event envelopes in time order, including the carried steps below, and returns the figures. options.fleet selects the gateway rule (ADR 0016 §6); options.labels is a map from requester id to label the caller fills.
  • Which window a fact belongs to: a request belongs to the window its lease.requested falls in; its grant, rejection and end are joined from events up to to, so an end after to is unseen and the lease is open at to with no held or turnaround sample. Provisioning, boot and incident counts go by the device event's timestamp. Utilisation and queue depth come from the steps; where no step precedes the window, the series points before the first step in it are null and excluded from peak and mean. requests counts lease.requested in the window (on a worker, only those without fleetRequestId; those with it count under probes, ADR 0021 §5); rejected counts the rejections of those requests, plus rejections refused before admission (no lease.requested), which belong to the window their lease.rejected falls in; so granted + rejected may exceed requests and docs/CLI.md says so. A request from before from that is granted or rejected inside the window is in no count.
  • Incidents: quarantined counts device.quarantined; crashRecovered counts device.recovered; quarantineRecovered counts device.quarantine-recovered; lost counts device.recovery-failed, device.quarantine-abandoned and device.quarantine-stranded.
  • Fleet rule (ADR 0016 §6 as amended by ADR 0021 §5): a gateway's own events are those with no payload.workerId. A fleet request's outcome is the gateway's own request.granted or lease.rejected for its request id; with neither, a later own daemon.started of the gateway ends it as rejected daemon-restarted, else it is open. Its wait runs from the gateway's lease.requested to that outcome, by the gateway's clock. The grant source and held time come from the relayed lease.granted whose workerId is the request.granted's worker and whose leaseId is its workerLeaseId, and that lease's relayed end, by the worker's clock; with no relayed grant the source is unknown and there is no held sample; turnaround is wait plus held. A grant counts for its request.granted's worker, a worker-failed rejection for its own worker. Relayed lease.declined count as events per worker and per platform (from requestSpec) under declined, in the window their timestamp falls in. Relayed lease.granted, lease.rejected, lease.declined, request.dispatched and daemon.started never decide an outcome. This PR does not reword ADR 0016 §6; ADR 0021 replaced its join. Per worker on a gateway (ADR 0021 §5): a worker's entry has its device facts, the fleet requests its request.granted names and its worker-failed rejections (with their waits and turnarounds), and its declined events; a fleet request that ended any other way, or is open, counts in totals and per platform under no worker.
  • Series: the bucket width is the smallest of 1 minute, 5 minutes, 15 minutes, 1 hour, 6 hours and 1 day that keeps the series at or under 200 points.
  • src/bus/event-file.ts and EventHistory.replay — a carry option: for each named event (capacity.changed, queue.changed), the latest envelope at or before sinceTs is kept in the result, one per workerId where present. The reader scans every line already, so this adds no I/O.
  • src/contract/errors.ts — HISTORY_NOT_KEPT joins the error table as a domain error with cliExitCode: 12, httpStatus: 422 and details { oldestTs }, so the CLI and the HTTP route take their behaviour from the table.
  • src/contract/operations.ts — usage.get, role admin, effect read, input { from, to } as epoch milliseconds with from < to and to - from bounded (at most 90 days). src/contract/schemas.ts for the output shape. src/contract/operations.test.ts role matrix.
  • src/daemon/dispatcher.ts — worker handler: reads the history for [from, to] through EventHistory (file plus ring, as events.replay with sinceTs, with carry for the two step events), fills labels from the token store, calls computeUsage with fleet: false, and answers HISTORY_NOT_KEPT when the oldest held timestamp is later than to. It rounds from and to down to bucketMs, memoises the last answer by (rounded window, newest event id), and answers a call inside the same bucket while no event arrived from the memo without reading the history; the answer's window is the rounded one.
  • src/gateway/dispatcher.ts — the same from the gateway's merged history with fleet: true, stripping its own gw:<instance>: prefix before the label lookup, with the same memo. The memo exists once, shared by both handlers.
  • src/daemon/server.ts socket switch; src/simlock-client/client.ts and types.ts — usage(window).
  • src/cli/index.ts — simlock stats [--since <duration> | --from <ISO> [--to <ISO>]] [--json]; default window the last 24 hours; --since and --from together is a usage error. Human output: a header with the window and coversFrom, the "partial" note when set, then the totals, the per-platform rows, the per-worker rows, the per-requester rows.
  • docs/CLI.md, docs/CLIENT.md, docs/HTTP-API.md error code list if it is shared.

Contract and event changes

  • New operation usage.get. Output, one object:
    • window { from, to }, coversFrom, partial: boolean, bucketMs.
    • totals: requests, granted, bySource { warm, booted, provisioned }, rejected { total, byReason }, declined (lease.declined events in the window by their timestamp: the worker's own on a worker, relayed ones on a gateway; not requests, ADR 0021 §5), probes (on a worker: lease.requested carrying fleetRequestId; requests counts only those without it; grants, source and held time count both; a probe gives no wait or turnaround sample; in requesters a probe counts under its gw: requester in granted and heldTotalMs only), bySource gains unknown on a gateway, wait { p50, p95, max, count }, held { p50, p95, max, count }, turnaround { p50, p95, max, count }, provisioning { p50, p95, max, count }, boot { p50, p95, max, count }, utilisation { slots { peak, mean, max }, ram? { peakBytes, meanBytes, limitBytes } }, queue { peakDepth, meanDepth }, incidents { quarantined, crashRecovered, quarantineRecovered, lost }, failures { byEvent }: a count per failure event name (device.purge-failed, device.recovery-failed, component.install-failed, and any other *-failed event) in the window by the event's timestamp.
    • platforms: the same figures keyed by ios and android.
    • workers: one entry per worker id with the same figures and a label; on a worker, one entry for itself.
    • requesters: { id, label?, requests, granted, rejected, heldTotalMs }[], by requests descending.
    • series: { at, slotsUsed, slotsMax, ramUsedBytes?, queueDepth, waiting }[], one point per bucket, waiting the number of requests waiting at the bucket's end.
    • Durations and percentiles in milliseconds; a percentile over an empty set is null.
  • New error code HISTORY_NOT_KEPT in the contract's error table (domain, cliExitCode: 12, httpStatus: 422) with oldestTs in its details.
  • Protocol range unchanged: the gateway does not call workers for this.

Rules in play

  • architecture.md rules 1 and 10: the figures module has no platform or transport imports; the fleet dedupe rule and the plan-source map exist once.
  • safety.md rule 10: from and to are wire input, bounded before use.
  • testing.md rule 1: each percentile test states the input set and the expected number.

Tests

  • computeUsage counts a request granted, held and released as one request, one grant, and one sample in wait, held and turnaround with the expected milliseconds
  • computeUsage reports p50, p95 and max of a set of twenty known waits
  • computeUsage counts a rejected request under its reason and not under grants
  • computeUsage attributes grants to warm, booted and provisioned from lease.granted.source
  • computeUsage reads provisioning and boot durations from device.provisioned and device.ready
  • computeUsage derives peak and time-weighted mean slot utilisation from capacity.changed steps, including the step in force before the window starts
  • computeUsage reports RAM utilisation when capacity.changed carries ramBudget and omits it otherwise
  • computeUsage derives peak and mean queue depth from queue.changed
  • computeUsage counts quarantined, crashRecovered, quarantineRecovered and lost from the named device events and nothing else
  • computeUsage counts a request whose lease.requested is before from and whose grant is inside the window under neither requests nor grants
  • computeUsage gives no held or turnaround sample for a lease still open at to, and counts its request and grant
  • computeUsage counts a rejection with no preceding lease.requested under rejected and not under requests
  • computeUsage counts a request from before from that is rejected inside the window under neither requests nor rejections
  • computeUsage reports null series points and excludes them from peak and mean before the first step when no step precedes the window
  • computeUsage groups by platform and by workerId
  • computeUsage with fleet: true counts a fleet request under the worker its request.granted or worker-failed rejection names, with its wait and turnaround, and a timeout rejection under no worker
  • computeUsage on a worker gives a probe no turnaround sample and lists its gw: requester with its grant and held time and no request
  • computeUsage with fleet: true counts a fleet request once, from the gateway's own events, and attributes its grant to the worker of its request.granted
  • computeUsage with fleet: true ignores relayed lease.requested, lease.queued, queue.changed, lease.rejected and lease.declined as request facts
  • computeUsage with fleet: true does not count a relayed lease.granted carrying a request's fleetRequestId as its outcome when the gateway rejected it
  • computeUsage with fleet: true does not close an open fleet request at a relayed worker daemon.started
  • computeUsage with fleet: true settles a request as rejected from the gateway's own lease.rejected, and a worker-failed one under its worker
  • computeUsage with fleet: true measures a fleet wait from the gateway's lease.requested to its request.granted, ignoring a relayed lease.granted stamped earlier
  • computeUsage with fleet: true takes the grant source and held time from the relayed lease.granted whose workerId and leaseId match the request.granted's worker and workerLeaseId, and counts source unknown with no held sample when it is missing
  • computeUsage with fleet: true counts each relayed lease.declined under its worker's and platform's declined, two declines of one request as two
  • computeUsage with fleet: true ends a request with no outcome at the gateway's next daemon.started as rejected daemon-restarted, and leaves one with no later start open
  • computeUsage on a worker counts a lease.requested carrying fleetRequestId under probes, not requests, its grant under granted, its lease.declined under declined by the decline's timestamp, and gives it no wait sample
  • computeUsage sets partial and coversFrom when the oldest event is inside the window
  • computeUsage picks 1-minute buckets for one hour, 15-minute buckets for one day and 1-hour buckets for seven days, and never more than 200 points for a 90-day window
  • readEventFile with carry returns the latest capacity.changed and queue.changed at or before sinceTs, one per worker id, and none when there is none
  • the worker handler answers a second call inside the same bucket from the memo without reading the history, and reads again after an event arrives or the bucket moves
  • computeUsage lists requesters by request count with the label the caller supplied
  • the worker handler answers HISTORY_NOT_KEPT with the oldest held timestamp when the window ends before it
  • the gateway handler strips its own gw: prefix before the label lookup and leaves another prefix alone
  • usage.get is an admin operation an agent token cannot call
  • simlock stats rejects --since with --from, and --from later than --to, as usage errors
  • simlock stats --json prints the operation's output unchanged
  • (e2e) after a scripted run of leases against the fake driver, simlock stats --since 1h --json reports the counts a test derives independently from simlock events --since 1h
  • (e2e) simlock stats --since 1h returns the same figures before and after simlock daemon stop and simlock daemon start
  • (e2e) on a gateway with two workers, simlock stats has fleet totals and one row per worker, and each worker's own simlock stats covers itself only

Done when

  • After a scripted run against the fake driver with waits, a --no-wait rejection and a timeout, simlock stats --since 1h prints request count, grants by source, held time and turnaround, wait p50, p95 and max, provisioning and boot durations, peak utilisation, queue depth, rejections by reason and incidents; --json prints the same figures; and the counts match what a reader counts by hand from simlock events --since 1h.
  • The same window prints the same figures after simlock daemon stop and simlock daemon start.
  • On a gateway with two workers, the output has fleet totals and one row per worker; each worker's own output has one row, itself.
  • A request that waited shows in the wait figures; a request that timed out and one cancelled are counted under rejections, not grants.
  • The per-requester rows show each requester's leases in the window, with the token label beside the id for a lease taken over HTTP.
  • After a --no-wait refusal, simlock stats --since 1h shows requests and rejected.byReason.no-wait each one higher than before it and granted unchanged.
  • simlock stats --from <a day ago> --to <two days ago> is a usage error; --since 30d on a history whose oldest event is a minute old prints the figures with "Figures cover from " above them; a window that ends before the oldest held event prints the HISTORY_NOT_KEPT message and exits non-zero.
  • docs/CLI.md describes simlock stats, its flags, the partial note and the error.

Out of scope

  • The HTTP route and the CSV export (next task).
  • The console view.

Depends on

Approval

  • Approved for delivery

Written by an agent.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    task:draftScope written; technical spec, approval, or deps missing.

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions