diff --git a/change_log.md b/change_log.md index 08be1de..38475e7 100644 --- a/change_log.md +++ b/change_log.md @@ -7,6 +7,7 @@ * Checker enforces `docs/reference/NUMERIC_CONTRACT.md`: mixed int/float `+`, `-`, `*` are valid and typed float, `/` always types float, and int/float comparisons are valid (HFREC-063). Void expressions in value position (`let`, `set`, call/print/list arguments, `append`/`map_set` values) and statically provable builtin argument mismatches (`str_split`, `str_join`, `regex_match`, `list_len`) are rejected before any backend runs (HFREC-009, HFREC-010). `parse_json` bodies must be variable names outside `web_app` (HFREC-004). Generated JavaScript keeps typed parameter names (HFREC-049). Go and JavaScript `to_int` truncate toward zero like the VM; `web_app` rejects integer literals and JSON integers outside +/-(2^53-1) instead of rounding. Wasm integer `/` fails closed. Artifact format version 2 (sorted function table, deterministic bytes) still reads version 1 artifacts; version 1 readers cannot read version 2 artifacts (see `docs/reference/artifact_compatibility.md`). New corpus: `tests/parity_stabilization/`. ### Changed +* Review C5: optional runner-written `--receipt` v0 JSON on both bytecode runners records artifact hash, grant, limits, instruction usage, redacted capability decisions, and exit outcome. Journal: [C5 execution receipt](docs/journals/2026-10-06_c5_execution_receipt.md). * Review C4b: bytecode cumulative allocation accounting enforces MaxMemoryBytes (64 MiB, `--max-memory-bytes`); fetch body and exec output caps default to 10 MiB (`--max-fetch-bytes`, `--max-exec-output-bytes`). Optional `--deadline` cancels fetch/exec/sleep/model calls and checks the instruction loop. Limits raise structured `LIMIT_EXCEEDED`; AST interpreter unchanged, no C5. Journal: [C4b resource limits](docs/journals/2026-10-06_c4b_resource_limits.md). * Review C4a: bytecode `CALL` now enforces `MaxCallDepth`, with default raised 128→1000 (shared with SPAWN_AGENT nesting), and `--max-call-depth` on `-run-bc` and `run`. Recursion fails with structured `LIMIT_EXCEEDED` instead of Go stack overflow. C4b (allocation accounting, deadline, fetch/exec byte caps) not included; AST interpreter `-run` unchanged. Journal: `docs/journals/2026-10-06_c4a_max_call_depth.md`. * Review C1/C2: `bytecode.RequiredCapabilities`, `inspect --caps` and `-required-caps` report compact JSON without execution. `check|build --profile governed` rejects excluded constructs, outside-root imports and Go/JS targets before output. Default behavior and HFBC encoding are unchanged. Journal: `docs/journals/2026-10-05_c1_c2_required_caps_governed_profile.md`. diff --git a/docs/ai-native-language-review/00-current-state.md b/docs/ai-native-language-review/00-current-state.md index 43a77fd..50bd2a2 100644 --- a/docs/ai-native-language-review/00-current-state.md +++ b/docs/ai-native-language-review/00-current-state.md @@ -56,7 +56,7 @@ Other paths that are still in the binary: | Deterministic artifacts | EXISTS (v2 sorted function table) | `artifact_determinism_test.go` | | Model-facing contract | PROTOTYPE, library only. Strict JSON candidate transport (128 nodes, 64 KiB), verifier, direct lowering, bounded `replace_node` repair with hash preconditions. It is not wired to the CLI. | `internal/hfir/model_adapter.go`, `docs/hfir_model_adapter_status.md` | | Build manifest and CAS | PROTOTYPE, library only. `hfir-manifest/v1` has graph, node, and module hashes, dependency maps, evidence hash, and artifact hash. It has no compiler version, policy, or grant. Disk and memory CAS. Incremental compiler. | `internal/hfir/manifest.go`, `storage.go`, `incremental.go` | -| Execution receipt (what ran, with which grant, what effects happened) | ABSENT | none | +| Execution receipt (what ran, with which grant, what effects happened) | PARTIAL (C5, 2026-10-06): runner-written bytecode v0; no signing/attestation or interpreter receipt | `internal/vm/receipt_test.go`, `receipt_cli_test.go`; C5 journal | | Static "what does this artifact need" report | ABSENT from the CLI. The data exists: per-opcode capabilities plus store URIs. | `opcode.go` | | Signed or attested artifacts | ABSENT | none | diff --git a/docs/ai-native-language-review/02-product-thesis.md b/docs/ai-native-language-review/02-product-thesis.md index acc7e47..d3e4e14 100644 --- a/docs/ai-native-language-review/02-product-thesis.md +++ b/docs/ai-native-language-review/02-product-thesis.md @@ -45,7 +45,8 @@ and VM, all written in about 10 weeks. 3. **"Verified" means less than it sounds.** The production gate blocks on two diagnostic codes. There is no stack or type validation of artifacts, - no execution receipt, and no signed provenance. + at review no execution receipt, and no signed provenance. C5 now supplies + optional bytecode receipts; signing/attestation remains open. 4. **No demonstrated user.** The best candidate job, the Release Authority / Action Executor pattern, can be built today with a JSON action schema and a 300-line Go executor. Nobody has shown that programmability beats diff --git a/docs/ai-native-language-review/05-competitive-prior-art.md b/docs/ai-native-language-review/05-competitive-prior-art.md index 7c3fe76..506ce63 100644 --- a/docs/ai-native-language-review/05-competitive-prior-art.md +++ b/docs/ai-native-language-review/05-competitive-prior-art.md @@ -10,7 +10,7 @@ HowlFrame describe `d496d972` plus the #57 fix in this PR. | System | Determinism | Termination / cost bound | Capability model | Static checks before run | Audit / provenance | Model-generation fit | Maturity | | --- | --- | --- | --- | --- | --- | --- | --- | -| **HowlFrame (today)** | ◐ (clock, sleep, model calls, and stdin are ambient) | ◐ instruction count only. No memory, wall-clock, or output bound. | ◐ 5 coarse classes, runner-owned, deny by default | ◐ checker + construct registry + 2 blocking HFIR codes | ◐ deterministic artifact + hash. Manifest in library only. No receipt. | ◐ small grammar. JSON graph transport prototype with bounded repair. | ○ ~10 weeks, one author | +| **HowlFrame (today)** | ◐ (clock, sleep, model calls, and stdin are ambient) | ◐ instruction count only. No memory, wall-clock, or output bound. | ◐ 5 coarse classes, runner-owned, deny by default | ◐ checker + construct registry + 2 blocking HFIR codes | ◐ deterministic artifact + hash. Manifest in library only. C5 adds optional runner-written bytecode receipts; signing/attestation still open. | ◐ small grammar. JSON graph transport prototype with bounded repair. | ○ ~10 weeks, one author | | **WASM + WASI 0.2 / component model (wasmtime)** | ◐ (deterministic core; host imports decide) | ● fuel and epoch interruption, `ResourceLimiter` for memory and tables | ● capability-based imports. Preopened dirs. WIT-typed interfaces. | ● module validation (type-checked stack machine) | ◐ module hashes. Signing via external tooling. | ○ models don't write Wasm. They write a source language that compiles to it. | ● | | **Starlark (Bazel, starlark-go)** | ● hermetic by design. No I/O unless the host injects it. Frozen values. | ◐ `SetMaxExecutionSteps` in starlark-go. Bazel forbids recursion. | ◐ host-injected builtins are the only effects | ◐ parse/resolve. Dynamic types. | ○ | ● Python-like, so models write it well | ● | | **CEL** | ● | ● non-Turing-complete. Static cost estimation plus runtime cost limit. | ● host provides all data. No effects. | ● type-checked expressions | ○ | ● for predicates. ○ for workflows. | ● (Kubernetes admission, IAM conditions) | diff --git a/docs/ai-native-language-review/06-reconciliation-matrix.md b/docs/ai-native-language-review/06-reconciliation-matrix.md index 8bf4ee6..4573c26 100644 --- a/docs/ai-native-language-review/06-reconciliation-matrix.md +++ b/docs/ai-native-language-review/06-reconciliation-matrix.md @@ -49,7 +49,7 @@ reviewer, checked against `d496d97` plus this PR's fix. Statuses: | Model synthesis evidence | **PARTIAL** (tiny) | 11 stored candidates. HFIR 10/10 against `.howl` 3/10, from one author (`hfir_model_adapter_status.md:35-43`). Benchmarks: v1 has 3 tasks; v2 has a harness and no results. | D, E, F | | Token-density advantage over Python | **CONFLICTS** | v1 CSV: HowlFrame used fewer tokens than Python in 1 of 3 tasks | C | | Required-capabilities inspection | **ABSENT** (CLI) | Effect inference exists in the HFIR verifier, but nothing is exposed for artifacts | F, executor | -| Execution receipt / audit record | **ABSENT** | The CLI throws away the bounded evidence object (`vm.go:1629-1644`) | A, F, executor | +| Execution receipt / audit record | **PARTIAL (C5, 2026-10-06)** | Optional runner-written bytecode v0 receipt; no signing/attestation, interpreter receipt, or sealed evidence binding | A, F, executor; C5 VM/CLI tests and journal | | CAS manifest and incremental compile | **PROTOTYPE** | `internal/hfir/manifest.go`, `storage.go`, `incremental.go`. Library only. No issuer, compiler version, or grant binding. | A, D | | Governed profile (construct allow-list for model-authored code) | **ABSENT** | None. See 11, C2. | D, F, executor | diff --git a/docs/ai-native-language-review/07-security-and-authority-findings.md b/docs/ai-native-language-review/07-security-and-authority-findings.md index 812329d..6ea4776 100644 --- a/docs/ai-native-language-review/07-security-and-authority-findings.md +++ b/docs/ai-native-language-review/07-security-and-authority-findings.md @@ -24,7 +24,7 @@ their own code. | S11 | **Generated Go/JS gate only 6 effects.** Model calls, SQL, listeners, goroutines, and JS `spawn` run ungated. A generated-Go panic writes `crash.json` unconditionally. | Medium (documented) | OPEN. Keep Go/JS out of the governed profile. | README states it. Codex B confirmed. | | S12 | **Artifact validation is structural only.** Magic, version, SHA-256, opcode, jump, and function existence are checked. There is no stack-height, operand-type, or embedded-body validation. SHA-256 proves integrity, not authorship. | Medium | OPEN (P2) | Code reading (`internal/bytecode/artifact.go`) | | S13 | **Compile-time includes read arbitrary paths** before any grant applies. Model-authored source could `include` files outside the project. | Low/Medium | OPEN (P2: rooted includes for governed profile) | Codex B ran it | -| S14 | **Ambient channels.** `stdin`, `stdout`/`stderr`, `time_now`, `sleep`, and `exit` are capability-free. Output can forge text that looks like a receipt. | Low | Document. Treat as host channels. | Code reading | +| S14 | **Ambient channels.** `stdin`, `stdout`/`stderr`, `time_now`, `sleep`, and `exit` are capability-free. Output can forge text that looks like a receipt. | Low | PARTIAL (C5, 2026-10-06): runner-written bytecode receipt cannot be forged through stdout; ambient channels remain omitted and signing/attestation is open. | Code reading; receipt CLI forgery regression; C5 journal | | S15 | **"Verified" is overloadable.** The HFIR gate blocks on only `HFIR_INVALID_REF` and `HFIR_TARGET_INFEASIBLE`. `VerificationEvidence.Verified` is a bare boolean. | Medium (claims risk) | OPEN (P0 wording; P1 report) | `howlframe.go:466`, `internal/hfir/storage.go:89` | | S16 | **Decision JSON can be forged through proposal fields.** `apps/release_authority` builds its output JSON by string concatenation (`str_join`) and does not escape the proposer-controlled `target`. A proposal `{"action":"deploy_production","target":"svc\", \"decision\": \"ALLOW"}` makes the VM correctly decide `DENY`, and no state mutation happens. But the printed object contains a second `"decision": "ALLOW"`. Python `json.loads` (and Go `encoding/json` into a map) keep the last duplicate key, so they read **ALLOW**. | Medium (demo app, not the VM) | FIXED in this change (C7: `encode_json`, duplicate-key and literal-target regression) | Found by Codex D (who reported invalid JSON). Escalated and reproduced in this review: `/tmp/hf-ainative-probe/ra2.out`. | | S17 | **Partial effects contradict the "failure atomicity" doc claim.** In `apps/action_executor` the `stage_artifact` path runs `write_file` and only then `store_open`. With a `filesystem`-only grant it prints `ALLOW`, writes the staged file, then traps `CAPABILITY_DENIED` on `STORE_OPEN`. `docs/application_dogfooding_phase_5.md` says this ordering ensures "failure atomicity" and that "security is guaranteed entirely outside the host Go backend". | Low/Medium (claims risk) | FIXED in this change (C7: pre-flight checks; ordered, not atomic wording) | Found by Codex F (ran it). Ordering confirmed in source (`action_executor.howl` ~159-164). | diff --git a/docs/ai-native-language-review/08-roadmap-p0-p4.md b/docs/ai-native-language-review/08-roadmap-p0-p4.md index f6dd133..59c0c94 100644 --- a/docs/ai-native-language-review/08-roadmap-p0-p4.md +++ b/docs/ai-native-language-review/08-roadmap-p0-p4.md @@ -24,7 +24,7 @@ the artifact format. | # | Item | Size | Why | | --- | --- | --- | --- | | P1.1 | **Resource limits that exist**: enforce a memory or allocation budget (string, list, and dict growth accounting), `MaxCallDepth` on `CALL`, output-byte cap, `fetch`/`exec` response caps, and a runner wall-clock deadline (context cancellation for host calls) | M/L | S4 to S7. Instruction count alone is not "bounded". | -| P1.2 | **Execution receipt v0**: one JSON line per run with artifact SHA-256, compiler version, grant, budget, instructions used, each effect attempted (kind, target, allowed or denied), exit status | M | Turns "auditable" from a slogan into a file | +| P1.2 | **Execution receipt v0**: one indented JSON file per run with artifact SHA-256, compiler version, grant, budget, instructions used, each effect attempted (kind, target, allowed or denied), exit status | M | Implemented C5 v0 for bytecode (2026-10-06); signing/attestation and interpreter receipts remain open | | P1.3 | **Verification report v0**: replace the bare `Verified` boolean with named checks, versions, and pass/fail (checker, construct registry, HFIR codes, artifact validation) | S/M | S15 | | P1.4 | **Decide `and`/`or` semantics** (short-circuit everywhere is the conventional choice) and add a differential test across interpreter, VM, and Go | S/M | S10 changes effects between backends | | P1.5 | **Run the first falsifiable experiment** (09) | M | Decides whether P2+ is worth doing | diff --git a/docs/ai-native-language-review/11-implementation-candidates.md b/docs/ai-native-language-review/11-implementation-candidates.md index eb5448d..d8186dc 100644 --- a/docs/ai-native-language-review/11-implementation-candidates.md +++ b/docs/ai-native-language-review/11-implementation-candidates.md @@ -89,7 +89,7 @@ C4a+C4b bytecode runner limits are implemented. AST interpreter memory, deadline, and recursion limits remain open; print output remains uncapped. No HFBC or production compiler change; #90 remains Partial. -## C5. Execution receipt v0 (P1.2): M +## C5. Execution receipt v0 (P1.2): M — v0 implemented - `-run-bc --receipt out.json`. The receipt holds `{schema:"howlframe.receipt/v0", artifact_sha256, compiler_version, grant, limits, instructions_used, @@ -99,6 +99,14 @@ No HFBC or production compiler change; #90 remains Partial. value. - Written by the runner, never by program code, so `print` cannot forge it. +**Status (C5, 2026-10-06):** implemented on legacy `-run-bc` and +`run --target bytecode/bc`, including in-process source compilation. Receipts +are host-written before process exit on normal and failed execution, with +bounded redacted allowed/denied decisions and shared child recording. No +signing/attestation, sealed ExecutionEvidence binding, AST interpreter receipt, +or harness in v0; no HFBC/compiler flip. #90 remains Partial. See +[C5 journal](../journals/2026-10-06_c5_execution_receipt.md). + ## C6. `and`/`or` semantics (P1.4): S/M, needs a decision - Choose short-circuit (Go and JS already do this, and it matches what diff --git a/docs/ai-native-language-review/FINAL_REPORT.md b/docs/ai-native-language-review/FINAL_REPORT.md index af8e87e..d75faf0 100644 --- a/docs/ai-native-language-review/FINAL_REPORT.md +++ b/docs/ai-native-language-review/FINAL_REPORT.md @@ -136,7 +136,7 @@ Not for the thesis. authority-preserving, hash-preconditioned local repair; 2. effect-requirement inference before running; 3. "intent is not authority" built into the runner; - 4. (planned) runner-written receipts. + 4. runner-written bytecode receipts (C5 v0; signing/attestation deferred). - **CaMeL-style provenance tracking** is the most relevant research to borrow. See 05. @@ -250,7 +250,7 @@ See the table below and 08. | P0 | Honest wording: what "verified", "bounded", and "atomic" mean; mark extreme paradigms superseded; fix the stale roadmap | S | Planned (docs) | | P0 | Demo-app output integrity (`encode_json`) and pre-flight or "ordered, not atomic" | S | Done in this change (C7: JSON integrity + pre-flight, ordered effects) | | **P1** | Memory/allocation, call-depth, deadline, and byte limits | M/L | Planned (C4) | -| P1 | Execution receipt v0, written by the runner | M | Planned (C5) | +| P1 | Execution receipt v0, written by the runner | M | Implemented C5 v0 for bytecode; signing/attestation and interpreter receipts deferred | | P1 | Verification report instead of the bare boolean | S/M | Planned | | P1 | `and`/`or` semantics decision plus a differential test | S/M | Planned (C6) | | P1 | **Run the first falsifiable experiment** | M | Designed (09) | diff --git a/docs/cli.md b/docs/cli.md index 9a679bf..17ed11b 100644 --- a/docs/cli.md +++ b/docs/cli.md @@ -30,6 +30,39 @@ Executes a compiled HowlFrame bytecode artifact. * `--max-exec-output-bytes` : Positive combined stdout/stderr ceiling for exec (default `10485760`, 10 MiB). * `--deadline` : Optional positive duration such as `5s` or `100ms`; absent means no wall-clock deadline. Explicit `0s`, negative durations, and invalid durations are rejected. Cancels bytecode fetch, exec, sleep, and model requests and checks between instructions. +* `--receipt ` : Optional runner-written `howlframe.receipt/v0` JSON file, + supported by `run --target bytecode`/`bc` for artifacts and `.howl` source, + and by legacy `-run-bc --receipt `. Interpreter targets + reject this option. Absent or empty means no receipt and unchanged output. + +The receipt records the exact artifact file SHA-256; for source compiled in +process it hashes canonical `bytecode.WriteArtifact` bytes. HFBC carries a +format version but no compiler identity, so `compiler_version` is the runner's +HowlFrame Version. It includes sorted unique grants, configured limits, executed +instructions (including synchronous SPAWN_AGENT work), bounded capability +decisions, and exit status/error. JSON uses fixed field order, two-space indent, +and a trailing newline; grant/effects are always arrays. No fields are omitted. +A nonzero exit/return has status `error` with `error: null`; runtime traps have +code 1 and a structured error. Zero deadline_ms means none (positive durations are +rounded up to whole milliseconds). + +Targets exclude URL paths/query/fragment/userinfo, env values, command arguments, +DSNs, SQL, store keys/values, response bodies, and model prompts. Filesystem +paths are recorded as given; resource names may still be sensitive. Summaries +are bounded to 256 bytes. Host failure messages are replaced by a code-only +message, except safe limit/capability messages; stderr diagnostics are unchanged. +Ambient print/stderr/sleep/time/read_line/exit are omitted. The first 10,000 +capability decisions are retained; `effects_truncated` reports overflow. Children +share the recorder; detached SPAWN effects after finalization are dropped. + +The runner writes after execution and before exit using a same-directory temp +file (0600 permissions) and rename. Program stdout/stderr are unchanged and +cannot forge the receipt through print. If writing fails, a diagnostic goes to +stderr: successful execution becomes exit 1; an existing failure code is kept. +Parent directories must exist. Receipts have no signature or attestation and +are independent of sealed ExecutionEvidence. See the +[C5 journal](journals/2026-10-06_c5_execution_receipt.md). + These limits also apply to legacy `-run-bc`. The new flags apply to `run` targets `bytecode`/`bc`; AST interpreter memory and deadline behavior is unchanged. Exceeded ceilings produce structured `LIMIT_EXCEEDED`. Blocking `read_line` and HTTP serving are not cancelled by the deadline; print output remains uncapped. diff --git a/docs/journals/2026-10-06_c5_execution_receipt.md b/docs/journals/2026-10-06_c5_execution_receipt.md new file mode 100644 index 0000000..e1e401b --- /dev/null +++ b/docs/journals/2026-10-06_c5_execution_receipt.md @@ -0,0 +1,158 @@ +# 2026-10-06: Runner-written execution receipt v0 (C5) + +## Findings and change + +- C5 supplies an optional durable execution record for the bytecode runner. + Review S14 found that program stdout could imitate a receipt. `--receipt` + now selects a separate host-written file; printing fake receipt JSON leaves + stdout unchanged and does not affect the real record. +- Both legacy `-run-bc` and `run --target bytecode/bc` support the option, + including `.howl` compiled in process. Empty/absent selects the existing + execution path. Interpreter targets (including legacy `-run`) reject a + nonempty receipt option. No new opcode exposes the recorder to programs. +- `RunBytecodeWithReceipt` reuses the evidence runner's recovery boundary but + returns code and receipt without terminating the process. Runtime diagnostics + remain identical. The CLI persists the receipt before `os.Exit`, including + capability/limit traps, other structured runtime failures, exit N, and return. + Parse, compile, and artifact validation errors occur before execution and do + not produce execution receipts. + +## Schema + +JSON uses explicit struct field order, two-space indentation, and a trailing +newline. No fields are omitted; empty grant/effects are `[]`, never null. + +| Field | v0 meaning | +| --- | --- | +| schema | `howlframe.receipt/v0` | +| artifact_sha256 | Hex SHA-256 of exact artifact file bytes; source uses canonical `bytecode.WriteArtifact` bytes generated before execution | +| compiler_version | Runner's HowlFrame Version; HFBC has format/program version but no recorded compiler identity | +| grant | Sorted, deduplicated runner capability allow-list | +| limits | max_instructions, max_call_depth, max_memory_bytes, max_fetch_body_bytes, max_exec_output_bytes, deadline_ms | +| instructions_used | Existing VM executed count, including synchronous SPAWN_AGENT work even on failures | +| effects | Ordered `{op, capability, target_summary, decision}` for each capability decision, allowed or denied | +| effects_truncated | True if the 10,000-entry effect cap is reached and another entry is attempted | +| exit | `{status, code, error}`; code is the intended program process exit code | +| exit.error | null or `{code, opcode, instruction, message}` | + +Zero deadline_ms means no deadline; positive durations are rounded up to whole +milliseconds so a configured submillisecond deadline does not appear absent. +A runtime trap has status error and code 1. Explicit nonzero exit/return has +status error with null error; zero exit/return and normal completion have status +ok. Host error messages can contain secrets, so receipt messages are code-only +(`runtime failure: CODE`) except safe LIMIT_EXCEEDED/CAPABILITY_DENIED messages. +Full existing diagnostics remain on stderr. The artifact checksum in HFBC covers +its payload, not its exact file envelope, so it is not reused for this digest. +Source serialization uses the existing deterministic v2 writer; no wire change. + +## Redaction policy + +Summaries are computed non-destructively at instruction decision time. Every +summary is bounded to 256 bytes, cutting at a UTF-8 boundary. Operand values are +never dumped by default. + +| Operation | Summary | Excluded | +| --- | --- | --- | +| FETCH | scheme://host[:port]; `` when unparseable or missing scheme/host | userinfo, path, query, fragment | +| ENV | Variable name | Value | +| READ_FILE, WRITE_FILE, MKDIR | Path as given, without cleaning | Contents | +| EXEC | Basename of program operand | Arguments | +| SPAWN, SPAWN_AGENT | Empty (VM tasks have no external argv[0]) | Task/body text | +| DB_CONNECT | Driver | DSN | +| SQL_QUERY | DB handle name | Query | +| STORE_OPEN | Store URI as given (memory://name or file://path) | Keys/values | +| Other STORE_* | Store handle name | Keys/values | +| HTTP_SERVER_START / HTTP_ROUTE / HTTP_SERVER_SERVE | Listen address / route pattern / saved listen address | Requests and bodies | +| RES, RES_JSON, HTTP_RES_HEADER, HTTP_REQ_METHOD | Empty | Bodies/headers | +| Model operations and lazy CALL | Empty | Prompts, inputs, model output | +| Default | Empty | All operands | + +Filesystem paths, store names, routes, and env names can themselves be sensitive; +the receipt is a resource-level record, not a blanket removal of personal data. + +## Decisions and limits + +- Central registry checks and explicit capability/store checks use the same + recorder. Each executed instruction has a separate decision context; repeated + checks of the same capability deduplicate, while database plus filesystem + store decisions remain separate. Denials record before the runtime panic. +- PRINT, STDERR, SLEEP, TIME_NOW, READ_LINE, EXIT and other capability-free + operations are ambient host channels and omitted in v0. +- Recorder pointer is shared by SPAWN goroutines, HTTP route children, and + SPAWN_AGENT. A mutex serializes effect append and finalization. Detached + SPAWN effects after finalization are dropped; those children retain historical + separate instruction accounting. HTTP child instruction counts likewise retain + existing behavior. Concurrent effects have scheduler-dependent order; byte + determinism applies to deterministic runs without concurrency/timing. +- First 10,000 decisions are retained and truncation is explicit. Summary + computation, decision allocation, and synchronization are absent when the + recorder pointer is nil. Child errors keep existing isolation semantics: a denied child effect need not + fail the parent run. Existing evidence sealing remains unchanged; the + receipt is an independent host reporting object. +- Persistence uses a same-directory temporary file, 0600 permissions, close, + then rename. Parent directories must exist. This is atomic replacement on + supported filesystems, not an fsync durability guarantee. Temp files are + removed on failures. If writing fails, stderr gets a host diagnostic: exit 1 + if execution succeeded, otherwise preserve the program's failure code. + `exit.code` describes execution rather than a subsequent persistence failure. + +## Verification + +- VM tests cover empty/pure programs, byte determinism, allowed/denied URL + redaction (real httptest and in-memory transport), environment/exec secrets, + limits with instruction usage, spawn-agent effects, per-capability store + decisions, exit/return and structured failures, cap/dedup/finalization, and + denied summary policy including invalid URLs and length bounds, every registry + gate and lazy CALL, listener-free HTTP children, and concurrent recording. +- CLI tests build the real binary and production artifacts, test both runners + and source compilation, verify the exact/canonical hash, fake stdout receipt, + capability/limit/runtime failures, exit/return codes, identical streams with + and without the option, interpreter rejection, no output file when absent or + empty, 0600 permissions, and write-failure exit behavior. +- Commands use `GOCACHE=/tmp/howlframe-c5-go-cache` because the default cache + under `/home/box/.cache` is read-only in this sandbox. Go is 1.24.4. +- Required: `gofmt -l .`, `go vet ./...`, + `go test ./internal/vm/ -count=1`, `go test ./... -count=1`. +- Supplemental hermetic receipt tests: + `go test ./internal/vm/ -run 'TestReceipt(Empty|FetchRedactionInMemory|Environment|Denied|SpawnAgent|Store|Exit|Cap|Concurrent|Registry|HTTPChild|Deadline)' -count=1`. +- C4 compatibility and CLI receipts: + `go test . -run 'TestReceipt|TestRunBytecodeResourceFlags|TestRunBytecodeMaxCallDepthFlag|TestRunBytecodeMaxInstructionsFlag' -count=1`. + +- `gofmt -l .` printed nothing; `git diff --check` and `go vet ./...` + passed. Hermetic receipt tests passed, also with `-race`. C4 CLI limits + and new receipt CLI tests passed. C4 VM compatibility checks passed: + `go test ./internal/vm -run 'ResourceDefaults|Memory|AllocationOpcode|FetchLimitsInMemory|Exec(OutputCap|LargeOutputAndExit|WithoutDeadline)|Deadline(SleepAndLoop|Model|HTTPResponseBodies)|CallDepth|SpawnAgent|Instruction|DefaultExecutionPolicy|EffectGateSource' -count=1`. +- Full VM validation stops at `TestInterpretFetchDeniedBeforeRequest`: + httptest listen on `[::1]:0` returns `socket: operation not permitted`. + The new real-listener `TestReceiptFetchRedaction` independently fails for + the same reason; its in-memory companion passes without weakening it. +- Full `go test ./... -count=1` was attempted and interrupted after the + existing app cleanup hangs were confirmed. A subsequent full run used + `-timeout=60s` to bound existing cleanup hangs; no tests were skipped or + weakened in source. The root package (including C4/C5 CLI tests) passes. + Sandbox blockers observed in the full suite: + +| Package / test | Sandbox evidence | +| --- | --- | +| apps/status_api TestStatusAPI | Listener on :8090 denied; cleanup waits on an already-consumed process-exit channel and times out | +| apps/task_api TestTaskAPI | Listener on :8091 denied; same cleanup-channel timeout | +| examples TestHTTPServerServeHFBC | listen tcp :8080: socket: operation not permitted | +| internal/backend/gogen TestGoFetchRequiresGrantBeforeRequest | httptest loopback listen denied | +| internal/backend/javascript TestJSExecRequiresGrantBeforeSpawn | spawnSync touch EPERM | +| internal/backend/javascript TestJSExecGrantedPrintsCapturedOutput | spawnSync printf EPERM | +| internal/backend/javascript TestJSFetchRequiresGrantBeforeRequest | httptest loopback listen denied | +| internal/vm TestInterpretFetchDeniedBeforeRequest | httptest loopback listen denied | +| tools/difftest TestLoweredHFIRABIConformance | listen tcp 127.0.0.1:47653: socket: operation not permitted | + +The two app artifact runners were also executed directly with their grants: +status_api and task_api each report structured IO_ERROR at HTTP_SERVER_SERVE +with the denied listener syscall. Full validation remains incomplete until +rerun in an environment permitting these syscalls; this journal does not claim +an unrestricted full-suite pass. + +## Deferrals + +No signing/attestation of receipts in v0. Receipts are not bound into sealed +ExecutionEvidence. AST interpreter receipts and a consumer experiment/harness +are deferred. No HFBC change, opcode registry change, DOM work, C6 work, or +production `-compile-bc`/HFIR flip. Improvement #90 remains Partial. diff --git a/howlframe.go b/howlframe.go index 5acfcaf..42b9eb7 100644 --- a/howlframe.go +++ b/howlframe.go @@ -2,6 +2,7 @@ package main import ( "bytes" + "crypto/sha256" "encoding/gob" "encoding/json" "errors" @@ -56,6 +57,7 @@ func main() { compileHfirBc := flag.Bool("compile-hfir-bc", false, "EXPERIMENTAL: compile directly from semantic HFIR to bytecode JSON") compileWasm := flag.Bool("compile-wasm", false, "compile typed SSA/CFG to WebAssembly Text") requiredCaps := flag.Bool("required-caps", false, "report required capabilities as compact JSON without execution") + receiptPath := flag.String("receipt", "", "write runner-owned bytecode execution receipt JSON to this path") runBc := flag.Bool("run-bc", false, "run bytecode from JSON file") allowCaps := flag.String("allow-caps", "", "comma-separated capabilities to allow when running bytecode with -run-bc (network,filesystem,process,environment,database); instructions requiring an unlisted capability are denied") maxInstructions := flag.Int("max-instructions", vm.DefaultLimits.MaxInstructions, "positive finite instruction ceiling for -run-bc (default 100000; zero and negative values are invalid)") @@ -68,6 +70,10 @@ func main() { maskPlan := flag.Bool("mask-plan", false, "print the deterministic constrained-decoding mask plan and exit") optimizationPlan := flag.Bool("optimization-plan", false, "print the deterministic compile-time optimization plan and exit") flag.Parse() + if *receiptPath != "" && !*runBc { + fmt.Fprintln(os.Stderr, "--receipt requires -run-bc; interpreter receipts are unsupported") + os.Exit(1) + } if *runBc { if _, err := resourcePolicy(*maxMemory, *maxFetch, *maxExec, *deadline); err != nil { ast.ReportError(err.Error(), 0, 0) @@ -154,7 +160,7 @@ func main() { } executionPolicy.Limits.MaxInstructions = *maxInstructions executionPolicy.Limits.MaxCallDepth = *maxCallDepth - os.Exit(vm.RunBytecodeWithPolicy(prog, programArgs, executionPolicy, parseAllowedCaps(*allowCaps), os.Stdin, os.Stdout, os.Stderr)) + os.Exit(runBytecodeArtifact(prog, content, programArgs, executionPolicy, parseAllowedCaps(*allowCaps), *receiptPath)) } lx := lexer.NewLexer(string(content)) @@ -901,6 +907,7 @@ func buildSource() { func runArtifact() { runFlags := flag.NewFlagSet("run", flag.ExitOnError) + receiptPath := runFlags.String("receipt", "", "write runner-owned bytecode execution receipt JSON to this path") target := runFlags.String("target", "bytecode", "execution target: bytecode (or bc), interpreter (or run)") allowCaps := runFlags.String("allow-caps", "", "comma-separated capabilities to allow (network,filesystem,process,environment,database)") maxInst := runFlags.Int("max-instructions", vm.DefaultLimits.MaxInstructions, "finite instruction limit") @@ -953,6 +960,10 @@ func runArtifact() { targetVal := strings.ToLower(strings.TrimSpace(*target)) switch targetVal { case "interpreter", "run": + if *receiptPath != "" { + fmt.Fprintln(os.Stderr, "--receipt requires bytecode target; interpreter receipts are unsupported") + os.Exit(1) + } lx := lexer.NewLexer(string(content)) p := parser.NewParser(lx, filepath.Base(inputFile)) root := p.ParseExpression() @@ -995,6 +1006,14 @@ func runArtifact() { hfirModule := filepath.Base(inputFile) runHFIRGate(root, hfirModule, hfirTargetBytecode) prog = bytecode.CompileToBytecode(root) + if *receiptPath != "" { + var canonical bytes.Buffer + if err := bytecode.WriteArtifact(&canonical, prog); err != nil { + fmt.Fprintf(os.Stderr, "Cannot serialize receipt artifact: %v\n", err) + os.Exit(1) + } + content = canonical.Bytes() + } } else { var err error prog, err = bytecode.ReadArtifact(bytes.NewReader(content)) @@ -1003,7 +1022,7 @@ func runArtifact() { os.Exit(1) } } - os.Exit(vm.RunBytecodeWithPolicy(prog, programArgs, executionPolicy, parseAllowedCaps(*allowCaps), os.Stdin, os.Stdout, os.Stderr)) + os.Exit(runBytecodeArtifact(prog, content, programArgs, executionPolicy, parseAllowedCaps(*allowCaps), *receiptPath)) default: fmt.Fprintf(os.Stderr, "Unknown target %q: valid run targets are bytecode, interpreter\n", *target) @@ -1062,3 +1081,40 @@ func resourcePolicy(memory, fetch, output int, deadline string) (vm.ExecutionPol } return policy, nil } + +// runBytecodeArtifact keeps the legacy path unchanged when reporting is absent. +func runBytecodeArtifact(prog *bytecode.BCProgram, artifact []byte, args []string, policy vm.ExecutionPolicy, caps []capability.Capability, receiptPath string) int { + if receiptPath == "" { + return vm.RunBytecodeWithPolicy(prog, args, policy, caps, os.Stdin, os.Stdout, os.Stderr) + } + digest := sha256.Sum256(artifact) + code, receipt := vm.RunBytecodeWithReceipt(prog, args, policy, caps, os.Stdin, os.Stdout, os.Stderr, fmt.Sprintf("%x", digest), Version) + if err := writeExecutionReceipt(receiptPath, receipt); err != nil { + fmt.Fprintf(os.Stderr, "Cannot write execution receipt: %v\n", err) + if code == 0 { + code = 1 + } + } + return code +} + +func writeExecutionReceipt(path string, receipt vm.ExecutionReceipt) error { + data, err := receipt.JSON() + if err != nil { + return err + } + f, err := os.CreateTemp(filepath.Dir(path), ".howlframe-receipt-*") + if err != nil { + return err + } + defer os.Remove(f.Name()) + // CreateTemp uses 0600; receipts may contain sensitive resource names. + if _, err := f.Write(data); err != nil { + f.Close() + return err + } + if err := f.Close(); err != nil { + return err + } + return os.Rename(f.Name(), path) +} diff --git a/internal/vm/receipt.go b/internal/vm/receipt.go new file mode 100644 index 0000000..a8e9da5 --- /dev/null +++ b/internal/vm/receipt.go @@ -0,0 +1,207 @@ +package vm + +import ( + "encoding/json" + "io" + "net/url" + "path/filepath" + "sort" + "sync" + "time" + "unicode/utf8" + + "github.com/howlcipher/howlframe/internal/bytecode" + "github.com/howlcipher/howlframe/internal/capability" +) + +// ExecutionReceipt is host-owned reporting, independent of sealed ExecutionEvidence. +// ArtifactSHA256 and CompilerVersion are supplied by the trusted artifact runner. +type ExecutionReceipt struct { + Schema string `json:"schema"` + ArtifactSHA256 string `json:"artifact_sha256"` + CompilerVersion string `json:"compiler_version"` + Grant []string `json:"grant"` + Limits ReceiptLimits `json:"limits"` + InstructionsUsed int `json:"instructions_used"` + Effects []ReceiptEffect `json:"effects"` + EffectsTruncated bool `json:"effects_truncated"` + Exit ReceiptExit `json:"exit"` +} +type ReceiptLimits struct { + MaxInstructions int `json:"max_instructions"` + MaxCallDepth int `json:"max_call_depth"` + MaxMemoryBytes int `json:"max_memory_bytes"` + MaxFetchBodyBytes int `json:"max_fetch_body_bytes"` + MaxExecOutputBytes int `json:"max_exec_output_bytes"` + DeadlineMS int64 `json:"deadline_ms"` +} +type ReceiptEffect struct { + Op string `json:"op"` + Capability string `json:"capability"` + TargetSummary string `json:"target_summary"` + Decision string `json:"decision"` +} +type ReceiptExit struct { + Status string `json:"status"` + Code int `json:"code"` + Error *ReceiptError `json:"error"` +} +type ReceiptError struct { + Code string `json:"code"` + Opcode string `json:"opcode"` + Instruction int `json:"instruction"` + Message string `json:"message"` +} + +func (r ExecutionReceipt) JSON() ([]byte, error) { + b, err := json.MarshalIndent(r, "", " ") + if err != nil { + return nil, err + } + return append(b, '\n'), nil +} + +const maxReceiptEffects = 10000 + +type receiptRecorder struct { + mu sync.Mutex + effects []ReceiptEffect + truncated bool + finalized bool + instructions int +} +type effectDecision struct { + summary string + seen []capability.Capability +} + +func (r *receiptRecorder) finish(instructions int) { + r.mu.Lock() + defer r.mu.Unlock() + r.instructions = instructions + r.finalized = true +} +func (vm *BCVM) recordEffect(cap capability.Capability, op bytecode.Opcode, decision string) { + if vm.receipt == nil { + return + } + d := vm.receiptDecision + if d == nil { + return + } + for _, seen := range d.seen { + if seen == cap { + return + } + } + d.seen = append(d.seen, cap) + r := vm.receipt + r.mu.Lock() + defer r.mu.Unlock() + if r.finalized { + return + } + if len(r.effects) >= maxReceiptEffects { + r.truncated = true + return + } + r.effects = append(r.effects, ReceiptEffect{bytecode.Registry[op].Name, string(cap), d.summary, decision}) +} + +// RunBytecodeWithReceipt returns normally on every VM outcome. Diagnostics match +// RunBytecodeWithPolicy; the caller must persist the receipt before process exit. +func RunBytecodeWithReceipt(prog *bytecode.BCProgram, args []string, policy ExecutionPolicy, caps []capability.Capability, in io.Reader, out, errOut io.Writer, artifactSHA256, compilerVersion string) (int, ExecutionReceipt) { + recorder := &receiptRecorder{effects: []ReceiptEffect{}} + ev := runBytecodeEvidence(prog, args, policy, caps, in, out, errOut, 0, recorder) + grant := []string{} + for _, cap := range caps { + grant = append(grant, string(cap)) + } + sort.Strings(grant) + grant = compactGrant(grant) + l := policy.Limits + var deadlineMS int64 + if policy.Deadline > 0 { + // Round up so zero always denotes an absent deadline, without overflow. + deadlineMS = int64((policy.Deadline-1)/time.Millisecond) + 1 + } + r := ExecutionReceipt{Schema: "howlframe.receipt/v0", ArtifactSHA256: artifactSHA256, CompilerVersion: compilerVersion, Grant: grant, + Limits: ReceiptLimits{l.MaxInstructions, l.MaxCallDepth, l.MaxMemoryBytes, l.MaxFetchBodyBytes, l.MaxExecOutputBytes, deadlineMS}, + InstructionsUsed: recorder.instructions, Effects: recorder.effects, EffectsTruncated: recorder.truncated, Exit: ReceiptExit{Status: "ok", Code: ev.ExitCode}} + if ev.RuntimeFailure != nil { + writeRuntimeFailure(errOut, ev.RuntimeFailure) + f := ev.RuntimeFailure + // Host errors can embed URLs, SQL, arguments, or prompts. Keep those only in + // the existing stderr diagnostic, never in a receipt. + message := "runtime failure: " + f.Code + if f.Code == "LIMIT_EXCEEDED" || f.Code == "CAPABILITY_DENIED" { + message = f.Message + } + r.Exit = ReceiptExit{Status: "error", Code: 1, Error: &ReceiptError{f.Code, f.Opcode, f.Instruction, message}} + } else if ev.ExitCode != 0 { + r.Exit.Status = "error" + } + return r.Exit.Code, r +} +func compactGrant(grant []string) []string { + n := 0 + for _, cap := range grant { + if n == 0 || grant[n-1] != cap { + grant[n] = cap + n++ + } + } + return grant[:n] +} + +func (vm *BCVM) effectSummary(inst bytecode.BCInstruction, env *BcEnv) string { + peek := func(depth int) string { + i := len(vm.stack) - 1 - depth + if i < 0 || i >= len(vm.stack) { + return "" + } + s, _ := vm.stack[i].(string) + return s + } + summary := "" + switch inst.Op { + case bytecode.OpFetch: + u, err := url.Parse(peek(1)) + if err != nil || u.Scheme == "" || u.Host == "" { + summary = "" + } else { + summary = u.Scheme + "://" + u.Host + } + case bytecode.OpEnv: + summary = peek(0) + case bytecode.OpReadFile, bytecode.OpMkdir: + summary = peek(0) + case bytecode.OpWriteFile: + summary = peek(1) + case bytecode.OpExec: + if inst.IntOperand >= 0 && inst.IntOperand < int64(len(vm.stack)) { + summary = peek(int(inst.IntOperand)) + if summary != "" { + summary = filepath.Base(summary) + } + } + case bytecode.OpDbConnect: + summary = inst.StringOperand2 + case bytecode.OpSqlQuery, bytecode.OpStorePut, bytecode.OpStoreGet, bytecode.OpStoreDelete, bytecode.OpStoreKeys: + summary = inst.StringOperand + case bytecode.OpStoreOpen: + summary = inst.StringOperand2 + case bytecode.OpHttpServerStart, bytecode.OpHttpRoute: + summary = inst.StringOperand + case bytecode.OpHttpServerServe: + v, _ := env.get("__http_port") + summary, _ = v.(string) + } + if len(summary) > 256 { + summary = summary[:256] + for !utf8.ValidString(summary) { + summary = summary[:len(summary)-1] + } + } + return summary +} diff --git a/internal/vm/receipt_test.go b/internal/vm/receipt_test.go new file mode 100644 index 0000000..0d29c6f --- /dev/null +++ b/internal/vm/receipt_test.go @@ -0,0 +1,270 @@ +package vm + +import ( + "bytes" + "fmt" + "io" + "net/http" + "net/http/httptest" + "strings" + "sync" + "testing" + "time" + + "github.com/howlcipher/howlframe/internal/bytecode" + "github.com/howlcipher/howlframe/internal/capability" +) + +func receiptRun(t *testing.T, source string, policy ExecutionPolicy, caps ...capability.Capability) (ExecutionReceipt, string, string) { + t.Helper() + _, prog := parseAndCompile(t, source) + var out, stderr bytes.Buffer + code, r := RunBytecodeWithReceipt(prog, nil, policy, caps, strings.NewReader(""), &out, &stderr, "trusted-hash", "test-version") + if code != r.Exit.Code { + t.Fatalf("code %d != receipt %d", code, r.Exit.Code) + } + b, err := r.JSON() + if err != nil { + t.Fatal(err) + } + return r, string(b), out.String() +} +func TestReceiptEmptyAndDeterministic(t *testing.T) { + for _, source := range []string{`(cli_app)`, `(cli_app (print "hello"))`} { + r, a, _ := receiptRun(t, source, DefaultExecutionPolicy()) + _, b, _ := receiptRun(t, source, DefaultExecutionPolicy()) + if a != b || !strings.Contains(a, `"effects": []`) || !strings.Contains(a, `"grant": []`) || r.Exit.Status != "ok" || !strings.HasSuffix(a, "\n") { + t.Fatalf("receipt: %s / %s", a, b) + } + } +} +func TestReceiptFetchRedaction(t *testing.T) { + server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { io.WriteString(w, "ok") })) + defer server.Close() + testReceiptFetch(t, server.URL) +} +func TestReceiptFetchRedactionInMemory(t *testing.T) { + old := http.DefaultClient + http.DefaultClient = &http.Client{Transport: receiptRoundTripFunc(func(r *http.Request) (*http.Response, error) { + return &http.Response{StatusCode: 200, Body: io.NopCloser(strings.NewReader("ok")), Header: make(http.Header)}, nil + })} + defer func() { http.DefaultClient = old }() + testReceiptFetch(t, "http://127.0.0.1:12345") +} +func testReceiptFetch(t *testing.T, origin string) { + t.Helper() + address := strings.Replace(origin, "http://", "http://user:pw@", 1) + "/path?token=secret#frag" + r, raw, _ := receiptRun(t, fmt.Sprintf(`(cli_app (fetch %q "GET"))`, address), DefaultExecutionPolicy(), capability.Network) + if r.Exit.Code != 0 || len(r.Effects) != 1 || r.Effects[0] != (ReceiptEffect{"FETCH", "network", origin, "allowed"}) { + t.Fatalf("receipt: %s", raw) + } + for _, secret := range []string{"secret", "token", "pw", "/path", "user:"} { + if strings.Contains(raw, secret) { + t.Fatalf("leaked %q: %s", secret, raw) + } + } +} +func TestReceiptEnvironmentAndExec(t *testing.T) { + t.Setenv("HF_RECEIPT_TEST", "secret-value-123") + r, raw, _ := receiptRun(t, `(cli_app (env "HF_RECEIPT_TEST") (exec "/bin/echo" "secret-argument-456"))`, DefaultExecutionPolicy(), capability.Environment, capability.Process, capability.Environment) + if r.Exit.Code != 0 || len(r.Effects) != 2 || r.Effects[0].TargetSummary != "HF_RECEIPT_TEST" || r.Effects[1].TargetSummary != "echo" || len(r.Grant) != 2 { + t.Fatalf("receipt: %s", raw) + } + for _, secret := range []string{"secret-value-123", "secret-argument-456"} { + if strings.Contains(raw, secret) { + t.Fatalf("leaked %s", raw) + } + } +} +func TestReceiptDeniedAndLimit(t *testing.T) { + r, raw, _ := receiptRun(t, `(cli_app (fetch "http://user:pw@example.com/path?token=secret" "GET"))`, DefaultExecutionPolicy()) + if r.Exit.Code != 1 || r.Exit.Error == nil || r.Exit.Error.Code != "CAPABILITY_DENIED" || len(r.Effects) != 1 || r.Effects[0].Decision != "denied" || r.Effects[0].TargetSummary != "http://example.com" { + t.Fatal(raw) + } + p := DefaultExecutionPolicy() + p.Limits.MaxInstructions = 1 + r, raw, _ = receiptRun(t, `(cli_app (print 42))`, p) + if r.Exit.Error == nil || r.Exit.Error.Code != "LIMIT_EXCEEDED" || r.InstructionsUsed != 1 { + t.Fatal(raw) + } +} +func TestReceiptSpawnAgent(t *testing.T) { + t.Setenv("HF_RECEIPT_CHILD", "child-secret") + r, raw, _ := receiptRun(t, `(cli_app (spawn_agent "child" (task "work" (env "HF_RECEIPT_CHILD"))))`, DefaultExecutionPolicy(), capability.Process, capability.Environment) + if len(r.Effects) != 2 || r.Effects[0].Op != "SPAWN_AGENT" || r.Effects[1].Op != "ENV" || r.Effects[1].TargetSummary != "HF_RECEIPT_CHILD" || r.InstructionsUsed < 4 { + t.Fatal(raw) + } + if strings.Contains(raw, "child-secret") || strings.Contains(raw, "work") { + t.Fatal(raw) + } +} +func TestReceiptStoreDecisions(t *testing.T) { + file := t.TempDir() + "/store.json" + source := fmt.Sprintf(`(cli_app (store_open data %q) (store_put data "secret-key" (dict ("x" "secret-value"))) (store_get data "secret-key"))`, "file://"+file) + r, raw, _ := receiptRun(t, source, DefaultExecutionPolicy(), capability.Database, capability.Filesystem) + if r.Exit.Code != 0 || len(r.Effects) != 6 { + t.Fatal(raw) + } + for i := 0; i < len(r.Effects); i += 2 { + if r.Effects[i].Capability != "database" || r.Effects[i+1].Capability != "filesystem" { + t.Fatal(raw) + } + } + if strings.Contains(raw, "secret-key") || strings.Contains(raw, "secret-value") { + t.Fatal(raw) + } +} +func TestReceiptExitAndFailure(t *testing.T) { + for _, source := range []string{`(cli_app (exit 7))`, `(cli_app (return 7))`} { + r, raw, _ := receiptRun(t, source, DefaultExecutionPolicy()) + if r.Exit.Code != 7 || r.Exit.Status != "error" || r.Exit.Error != nil { + t.Fatal(raw) + } + } + r, raw, _ := receiptRun(t, `(cli_app (read_file "/missing/secret-file"))`, DefaultExecutionPolicy(), capability.Filesystem) + if r.Exit.Error == nil || r.Exit.Error.Code != "IO_ERROR" || strings.Contains(r.Exit.Error.Message, "secret-file") { + t.Fatal(raw) + } +} +func TestReceiptCapAndFinalization(t *testing.T) { + recorder := &receiptRecorder{effects: []ReceiptEffect{}} + vm := &BCVM{receipt: recorder} + for i := 0; i < maxReceiptEffects+1; i++ { + vm.receiptDecision = &effectDecision{summary: "name"} + vm.recordEffect(capability.Environment, bytecode.OpEnv, "allowed") + vm.recordEffect(capability.Environment, bytecode.OpEnv, "allowed") + } + if len(recorder.effects) != maxReceiptEffects || !recorder.truncated { + t.Fatal("cap or dedup failed") + } + recorder.finish(3) + vm.receiptDecision = &effectDecision{} + vm.recordEffect(capability.Network, bytecode.OpFetch, "denied") + if recorder.instructions != 3 || len(recorder.effects) != maxReceiptEffects { + t.Fatal("late effect recorded") + } +} + +type receiptRoundTripFunc func(*http.Request) (*http.Response, error) + +func (f receiptRoundTripFunc) RoundTrip(r *http.Request) (*http.Response, error) { return f(r) } + +func TestReceiptDeniedSummaryPolicy(t *testing.T) { + for _, tc := range []struct { + op bytecode.Opcode + operands []any + inst bytecode.BCInstruction + want string + }{ + {op: bytecode.OpFetch, operands: []any{"%%%", "GET"}, want: ""}, + {op: bytecode.OpWriteFile, operands: []any{"a/../file", "content-secret"}, want: "a/../file"}, + {op: bytecode.OpReadFile, operands: []any{"file"}, want: "file"}, + {op: bytecode.OpMkdir, operands: []any{"directory"}, want: "directory"}, + {op: bytecode.OpExec, operands: []any{"/bin/echo", "arg-secret"}, inst: bytecode.BCInstruction{IntOperand: 1}, want: "echo"}, + {op: bytecode.OpDbConnect, inst: bytecode.BCInstruction{StringOperand: "db", StringOperand2: "sqlite", StringOperand3: "dsn-secret"}, want: "sqlite"}, + {op: bytecode.OpSqlQuery, inst: bytecode.BCInstruction{StringOperand: "db", StringOperand2: "query-secret"}, want: "db"}, + {op: bytecode.OpHttpRoute, inst: bytecode.BCInstruction{StringOperand: "/route"}, want: "/route"}, + {op: bytecode.OpResJson, operands: []any{int64(200), "body-secret"}, want: ""}, + {op: bytecode.OpLlmGenerate, operands: []any{"prompt-secret"}, inst: bytecode.BCInstruction{StringOperand: "model"}, want: ""}, + {op: bytecode.OpEnv, operands: []any{strings.Repeat("x", 300)}, want: strings.Repeat("x", 256)}, + } { + t.Run(bytecode.Registry[tc.op].Name, func(t *testing.T) { + prog := &bytecode.BCProgram{} + for _, value := range tc.operands { + prog.Main = append(prog.Main, bytecode.BCInstruction{Op: bytecode.OpLoadConst, ValueOperand: value}) + } + inst := tc.inst + inst.Op = tc.op + prog.Main = append(prog.Main, inst) + var out, stderr bytes.Buffer + _, r := RunBytecodeWithReceipt(prog, nil, DefaultExecutionPolicy(), nil, strings.NewReader(""), &out, &stderr, "hash", "version") + if r.Exit.Error == nil || r.Exit.Error.Code != "CAPABILITY_DENIED" || len(r.Effects) != 1 || r.Effects[0].TargetSummary != tc.want { + t.Fatalf("receipt %+v", r) + } + raw, err := r.JSON() + if err != nil { + t.Fatal(err) + } + if bytes.Contains(raw, []byte("secret")) { + t.Fatal(string(raw)) + } + }) + } +} + +func TestReceiptConcurrentRecorder(t *testing.T) { + recorder := &receiptRecorder{effects: []ReceiptEffect{}} + var workers sync.WaitGroup + for i := 0; i < 8; i++ { + workers.Add(1) + go func() { + defer workers.Done() + child := &BCVM{receipt: recorder} + for j := 0; j < 100; j++ { + child.receiptDecision = &effectDecision{summary: "name"} + child.recordEffect(capability.Environment, bytecode.OpEnv, "allowed") + } + }() + } + workers.Wait() + recorder.finish(0) + if len(recorder.effects) != 800 { + t.Fatalf("lost effects: %d", len(recorder.effects)) + } +} + +func TestReceiptRegistryAndLazyDecisions(t *testing.T) { + for op, spec := range bytecode.Registry { + if spec.Capability == capability.None { + continue + } + t.Run(spec.Name, func(t *testing.T) { + prog := &bytecode.BCProgram{Main: []bytecode.BCInstruction{{Op: op}}} + var out, stderr bytes.Buffer + _, r := RunBytecodeWithReceipt(prog, nil, DefaultExecutionPolicy(), nil, strings.NewReader(""), &out, &stderr, "hash", "version") + if len(r.Effects) != 1 || r.Effects[0].Op != spec.Name || r.Effects[0].Decision != "denied" || r.Exit.Error == nil || r.Exit.Error.Code != "CAPABILITY_DENIED" { + t.Fatalf("receipt %+v", r) + } + }) + } + prog := &bytecode.BCProgram{Functions: map[string]*bytecode.BCFunction{"lazy": {Name: "lazy", LazySynthesize: true, Docstring: "prompt-secret"}}, Main: []bytecode.BCInstruction{{Op: bytecode.OpCall, StringOperand: "lazy"}}} + var out, stderr bytes.Buffer + _, r := RunBytecodeWithReceipt(prog, nil, DefaultExecutionPolicy(), nil, strings.NewReader(""), &out, &stderr, "hash", "version") + if len(r.Effects) != 1 || r.Effects[0].Op != "CALL" || r.Effects[0].Capability != "network" || r.Effects[0].Decision != "denied" || r.Effects[0].TargetSummary != "" { + t.Fatalf("lazy receipt %+v", r) + } +} + +func TestReceiptHTTPChild(t *testing.T) { + recorder := &receiptRecorder{effects: []ReceiptEffect{}} + prog := &bytecode.BCProgram{Main: []bytecode.BCInstruction{ + {Op: bytecode.OpHttpServerStart, StringOperand: ":12345"}, + {Op: bytecode.OpHttpRoute, StringOperand: "/test", StringOperand2: "req", IntOperand: 2}, + {Op: bytecode.OpLoadConst, ValueOperand: "HF_RECEIPT_HTTP"}, + {Op: bytecode.OpEnv}, + }} + env := NewBcEnv(nil) + machine := &BCVM{prog: prog, env: env, receipt: recorder, Limits: DefaultLimits, AllowedCaps: []capability.Capability{capability.Network, capability.Environment}, Out: io.Discard, ErrOut: io.Discard} + machine.run(prog.Main, env) + mux, _ := env.get("__http_mux") + mux.(*http.ServeMux).ServeHTTP(httptest.NewRecorder(), httptest.NewRequest("GET", "http://localhost/test", nil)) + recorder.finish(machine.executed) + if len(recorder.effects) != 3 || recorder.effects[2].Op != "ENV" || recorder.effects[2].TargetSummary != "HF_RECEIPT_HTTP" { + t.Fatalf("HTTP child effects: %+v", recorder.effects) + } +} + +func TestReceiptDeadlineUnits(t *testing.T) { + for _, deadline := range []time.Duration{0, time.Nanosecond, time.Millisecond, time.Millisecond + 1} { + policy := DefaultExecutionPolicy() + policy.Deadline = deadline + r, _, _ := receiptRun(t, `(cli_app)`, policy) + want := int64(0) + if deadline > 0 { + want = int64((deadline-1)/time.Millisecond) + 1 + } + if r.Limits.DeadlineMS != want { + t.Fatalf("deadline %s: %d want %d", deadline, r.Limits.DeadlineMS, want) + } + } +} diff --git a/internal/vm/vm.go b/internal/vm/vm.go index 4cb17f8..ed6d1d5 100644 --- a/internal/vm/vm.go +++ b/internal/vm/vm.go @@ -1388,6 +1388,9 @@ func SliceToAny(strs []string) []any { } type BCVM struct { + receipt *receiptRecorder + receiptDecision *effectDecision + ctx context.Context allocations *allocationBudget prog *bytecode.BCProgram @@ -1650,6 +1653,10 @@ func RunBytecodeWithPolicy(prog *bytecode.BCProgram, cliArgs []string, policy Ex // returns a bounded runner-owned trace for trusted consumers such as HFIR // failure localization. traceLimit is a runner policy, never program input. func RunBytecodeWithEvidence(prog *bytecode.BCProgram, cliArgs []string, policy ExecutionPolicy, allowedCaps []capability.Capability, in io.Reader, out io.Writer, errOut io.Writer, traceLimit int) (evidence bytecode.ExecutionEvidence) { + return runBytecodeEvidence(prog, cliArgs, policy, allowedCaps, in, out, errOut, traceLimit, nil) +} + +func runBytecodeEvidence(prog *bytecode.BCProgram, cliArgs []string, policy ExecutionPolicy, allowedCaps []capability.Capability, in io.Reader, out io.Writer, errOut io.Writer, traceLimit int, recorder *receiptRecorder) (evidence bytecode.ExecutionEvidence) { ctx := context.Background() if policy.Deadline > 0 { var cancel context.CancelFunc @@ -1657,7 +1664,7 @@ func RunBytecodeWithEvidence(prog *bytecode.BCProgram, cliArgs []string, policy defer cancel() } vm := &BCVM{ - ctx: ctx, allocations: &allocationBudget{}, + ctx: ctx, allocations: &allocationBudget{}, receipt: recorder, prog: prog, env: NewBcEnv(nil), insts: prog.Main, @@ -1709,6 +1716,9 @@ func RunBytecodeWithEvidence(prog *bytecode.BCProgram, cliArgs []string, policy evidence.RuntimeFailure = &bytecode.RuntimeFailure{Code: "VM_INTERNAL", Instruction: vm.ip, Message: fmt.Sprintf("%v", r)} } } + if recorder != nil { + recorder.finish(vm.executed) + } bytecode.SealExecutionEvidence(prog, &evidence) }() @@ -1716,6 +1726,13 @@ func RunBytecodeWithEvidence(prog *bytecode.BCProgram, cliArgs []string, policy return evidence } +func writeRuntimeFailure(errOut io.Writer, failure *bytecode.RuntimeFailure) { + if errOut == nil { + errOut = os.Stderr + } + fmt.Fprintln(errOut, mustVMErrorString(failure)) +} + func mustVMErrorString(failure *bytecode.RuntimeFailure) string { return (&VMError{Phase: "runtime", Code: failure.Code, Function: "main", Instruction: failure.Instruction, Opcode: failure.Opcode, NodeID: failure.NodeID, Message: failure.Message}).Error() } @@ -1771,9 +1788,11 @@ func (vm *BCVM) storeHandle(env *BcEnv, name string, op bytecode.Opcode) *bcMemo func (vm *BCVM) requireCapability(cap capability.Capability, op bytecode.Opcode) { for _, allowed := range vm.AllowedCaps { if allowed == cap { + vm.recordEffect(cap, op, "allowed") return } } + vm.recordEffect(cap, op, "denied") panic(NewRuntimeError("CAPABILITY_DENIED", "main", vm.ip, op, "capability denied: %s", cap)) } @@ -1963,6 +1982,9 @@ func (vm *BCVM) run(insts []bytecode.BCInstruction, env *BcEnv) any { panic(NewRuntimeError("LIMIT_EXCEEDED", "main", vm.ip, inst.Op, "instruction limit exceeded")) } vm.executed++ + if vm.receipt != nil { + vm.receiptDecision = &effectDecision{summary: vm.effectSummary(inst, env)} + } spec := bytecode.Registry[inst.Op] if spec.Capability != capability.None { @@ -2300,7 +2322,7 @@ func (vm *BCVM) run(insts []bytecode.BCInstruction, env *BcEnv) any { } go func(cEnv *BcEnv) { defer func() { recover() }() - childVM := &BCVM{ctx: vm.ctx, allocations: vm.allocations, prog: vm.prog, env: cEnv, stores: vm.stores, Limits: vm.Limits, AllowedCaps: vm.AllowedCaps, Out: vm.Out, ErrOut: vm.ErrOut} + childVM := &BCVM{ctx: vm.ctx, allocations: vm.allocations, receipt: vm.receipt, prog: vm.prog, env: cEnv, stores: vm.stores, Limits: vm.Limits, AllowedCaps: vm.AllowedCaps, Out: vm.Out, ErrOut: vm.ErrOut} childVM.run(bodyInsts, cEnv) }(capturedEnv) ip += bodyLen @@ -2377,7 +2399,7 @@ func (vm *BCVM) run(insts []bytecode.BCInstruction, env *BcEnv) any { reqEnv.vars["w"] = http.ResponseWriter(recorder) reqEnv.vars[reqVar] = r reqEnv.vars["req"] = r - childVM := &BCVM{ctx: vm.ctx, allocations: vm.allocations, prog: prog, env: reqEnv, stores: vm.stores, Limits: vm.Limits, AllowedCaps: vm.AllowedCaps, Out: vm.Out, ErrOut: vm.ErrOut} + childVM := &BCVM{ctx: vm.ctx, allocations: vm.allocations, receipt: vm.receipt, prog: prog, env: reqEnv, stores: vm.stores, Limits: vm.Limits, AllowedCaps: vm.AllowedCaps, Out: vm.Out, ErrOut: vm.ErrOut} func() { defer func() { recovered := recover() @@ -3170,7 +3192,7 @@ func (vm *BCVM) run(insts []bytecode.BCInstruction, env *BcEnv) any { } } // Synchronous children share the buffered reader to preserve stdin order. - childVM := &BCVM{ctx: vm.ctx, allocations: vm.allocations, prog: vm.prog, env: capturedEnv, args: vm.args, In: vm.In, lineReader: vm.lineReader, stores: vm.stores, Limits: vm.Limits, AllowedCaps: vm.AllowedCaps, Out: vm.Out, ErrOut: vm.ErrOut, executed: vm.executed, spawnDepth: vm.spawnDepth + 1, callDepth: vm.callDepth} + childVM := &BCVM{ctx: vm.ctx, allocations: vm.allocations, receipt: vm.receipt, prog: vm.prog, env: capturedEnv, args: vm.args, In: vm.In, lineReader: vm.lineReader, stores: vm.stores, Limits: vm.Limits, AllowedCaps: vm.AllowedCaps, Out: vm.Out, ErrOut: vm.ErrOut, executed: vm.executed, spawnDepth: vm.spawnDepth + 1, callDepth: vm.callDepth} func() { defer func() { // Account for all child work even when it exits through a panic. diff --git a/receipt_cli_test.go b/receipt_cli_test.go new file mode 100644 index 0000000..53c59d9 --- /dev/null +++ b/receipt_cli_test.go @@ -0,0 +1,168 @@ +package main + +import ( + "bytes" + "crypto/sha256" + "encoding/json" + "fmt" + "os" + "os/exec" + "path/filepath" + "strings" + "testing" + + "github.com/howlcipher/howlframe/internal/bytecode" + "github.com/howlcipher/howlframe/internal/vm" +) + +func TestReceiptCLI(t *testing.T) { + binary := buildHowlFrameBinaryForTest(t) + dir := t.TempDir() + fake := `{"schema":"howlframe.receipt/v0","effects":[{"op":"FETCH","decision":"allowed"}]}` + cases := []struct { + name, source, errorCode string + flags []string + code int + }{ + {name: "forge", source: fmt.Sprintf(`(cli_app (print %q))`, fake)}, + {name: "deny", source: `(cli_app (fetch "http://example.com/private?token=secret" "GET"))`, errorCode: "CAPABILITY_DENIED", code: 1}, + {name: "limit", source: `(cli_app (print 42))`, errorCode: "LIMIT_EXCEEDED", flags: []string{"--max-instructions", "1"}, code: 1}, + {name: "exit", source: `(cli_app (exit 7))`, code: 7}, + {name: "return", source: `(cli_app (return 7))`, code: 7}, + {name: "runtime", source: `(cli_app (read_file "/missing/receipt-file"))`, errorCode: "IO_ERROR", flags: []string{"--allow-caps", "filesystem"}, code: 1}, + } + for _, tc := range cases { + source := filepath.Join(dir, tc.name+".howl") + artifact := filepath.Join(dir, tc.name+".hfbc") + if err := os.WriteFile(source, []byte(tc.source), 0600); err != nil { + t.Fatal(err) + } + if out, err := exec.Command(binary, "-compile-bc", source, "-o", artifact).CombinedOutput(); err != nil { + t.Fatalf("compile: %v %s", err, out) + } + artifactBytes, err := os.ReadFile(artifact) + if err != nil { + t.Fatal(err) + } + for _, mode := range []struct{ runner, input string }{{"-run-bc", artifact}, {"run", artifact}, {"run", source}} { + t.Run(tc.name+"/"+mode.runner+"/"+filepath.Ext(mode.input), func(t *testing.T) { + path := filepath.Join(t.TempDir(), "receipt.json") + args := append([]string{mode.runner, "--receipt", path}, tc.flags...) + args = append(args, mode.input) + var stdout, stderr bytes.Buffer + cmd := exec.Command(binary, args...) + cmd.Stdout = &stdout + cmd.Stderr = &stderr + err := cmd.Run() + code := 0 + if err != nil { + e, ok := err.(*exec.ExitError) + if !ok { + t.Fatal(err) + } + code = e.ExitCode() + } + if code != tc.code { + t.Fatalf("code %d want %d: %s", code, tc.code, &stderr) + } + raw, err := os.ReadFile(path) + if err != nil { + t.Fatal(err) + } + var r vm.ExecutionReceipt + if err := json.Unmarshal(raw, &r); err != nil { + t.Fatal(err) + } + if r.Schema != "howlframe.receipt/v0" || r.CompilerVersion != Version || r.Exit.Code != code { + t.Fatalf("receipt %s", raw) + } + canonical := artifactBytes + if filepath.Ext(mode.input) == ".howl" { + prog, err := bytecode.ReadArtifact(bytes.NewReader(artifactBytes)) + if err != nil { + t.Fatal(err) + } + var b bytes.Buffer + if err := bytecode.WriteArtifact(&b, prog); err != nil { + t.Fatal(err) + } + canonical = b.Bytes() + } + if r.ArtifactSHA256 != fmt.Sprintf("%x", sha256.Sum256(canonical)) { + t.Fatalf("wrong artifact hash: %s", raw) + } + if tc.errorCode != "" && (r.Exit.Error == nil || r.Exit.Error.Code != tc.errorCode) { + t.Fatal(string(raw)) + } + if tc.name == "forge" && (stdout.String() != fake+"\n" || len(r.Effects) != 0 || !bytes.Contains(raw, []byte(`"effects": []`))) { + t.Fatalf("forged receipt %s stdout %s", raw, &stdout) + } + stat, err := os.Stat(path) + if err != nil { + t.Fatal(err) + } + if stat.Mode().Perm() != 0600 { + t.Fatalf("permissions: %v", stat.Mode()) + } + // The receipt option must not change program streams or process status. + plainArgs := append([]string{mode.runner}, tc.flags...) + plainArgs = append(plainArgs, mode.input) + var plainOut, plainErr bytes.Buffer + plain := exec.Command(binary, plainArgs...) + plain.Stdout = &plainOut + plain.Stderr = &plainErr + plain.Run() + if stdout.String() != plainOut.String() || stderr.String() != plainErr.String() { + t.Fatalf("streams changed: %q %q vs %q %q", &stdout, &stderr, &plainOut, &plainErr) + } + }) + } + } + for _, runner := range []string{"-run-bc", "run"} { + for _, flags := range [][]string{nil, {"--receipt", ""}} { + args := append([]string{runner}, flags...) + args = append(args, filepath.Join(dir, "forge.hfbc")) + cmd := exec.Command(binary, args...) + cmd.Dir = t.TempDir() + output, err := cmd.CombinedOutput() + if err != nil || string(output) != fake+"\n" { + t.Fatalf("no receipt: %v %s", err, output) + } + files, err := os.ReadDir(cmd.Dir) + if err != nil || len(files) != 0 { + t.Fatalf("unexpected output files: %v %v", files, err) + } + } + } + out, err := exec.Command(binary, "run", "--target", "interpreter", "--receipt", filepath.Join(dir, "unsupported.json"), filepath.Join(dir, "forge.howl")).CombinedOutput() + if err == nil || !strings.Contains(string(out), "--receipt requires bytecode target") { + t.Fatalf("interpreter: %v %s", err, out) + } + out, err = exec.Command(binary, "-run", "--receipt", filepath.Join(dir, "unsupported.json"), filepath.Join(dir, "forge.howl")).CombinedOutput() + if err == nil || !strings.Contains(string(out), "--receipt requires -run-bc") { + t.Fatalf("legacy interpreter: %v %s", err, out) + } + if _, err := os.Stat(filepath.Join(dir, "unsupported.json")); !os.IsNotExist(err) { + t.Fatalf("unexpected receipt: %v", err) + } +} + +func TestReceiptWriteFailureCLI(t *testing.T) { + binary := buildHowlFrameBinaryForTest(t) + dir := t.TempDir() + for _, code := range []int{0, 7} { + source := filepath.Join(dir, fmt.Sprintf("exit%d.howl", code)) + if err := os.WriteFile(source, []byte(fmt.Sprintf(`(cli_app (exit %d))`, code)), 0600); err != nil { + t.Fatal(err) + } + out, err := exec.Command(binary, "run", "--receipt", filepath.Join(dir, "absent", "receipt.json"), source).CombinedOutput() + e, ok := err.(*exec.ExitError) + want := code + if want == 0 { + want = 1 + } + if !ok || e.ExitCode() != want || !strings.Contains(string(out), "Cannot write execution receipt") { + t.Fatalf("write failure: %v %s", err, out) + } + } +}