Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions change_log.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,7 @@
* Checker enforces `docs/reference/NUMERIC_CONTRACT.md`: mixed int/float `+`, `-`, `*` are valid and typed float, `/` always types float, and int/float comparisons are valid (HFREC-063). Void expressions in value position (`let`, `set`, call/print/list arguments, `append`/`map_set` values) and statically provable builtin argument mismatches (`str_split`, `str_join`, `regex_match`, `list_len`) are rejected before any backend runs (HFREC-009, HFREC-010). `parse_json` bodies must be variable names outside `web_app` (HFREC-004). Generated JavaScript keeps typed parameter names (HFREC-049). Go and JavaScript `to_int` truncate toward zero like the VM; `web_app` rejects integer literals and JSON integers outside +/-(2^53-1) instead of rounding. Wasm integer `/` fails closed. Artifact format version 2 (sorted function table, deterministic bytes) still reads version 1 artifacts; version 1 readers cannot read version 2 artifacts (see `docs/reference/artifact_compatibility.md`). New corpus: `tests/parity_stabilization/`.

### Changed
* Review C5: optional runner-written `--receipt` v0 JSON on both bytecode runners records artifact hash, grant, limits, instruction usage, redacted capability decisions, and exit outcome. Journal: [C5 execution receipt](docs/journals/2026-10-06_c5_execution_receipt.md).
* Review C4b: bytecode cumulative allocation accounting enforces MaxMemoryBytes (64 MiB, `--max-memory-bytes`); fetch body and exec output caps default to 10 MiB (`--max-fetch-bytes`, `--max-exec-output-bytes`). Optional `--deadline` cancels fetch/exec/sleep/model calls and checks the instruction loop. Limits raise structured `LIMIT_EXCEEDED`; AST interpreter unchanged, no C5. Journal: [C4b resource limits](docs/journals/2026-10-06_c4b_resource_limits.md).
* Review C4a: bytecode `CALL` now enforces `MaxCallDepth`, with default raised 128→1000 (shared with SPAWN_AGENT nesting), and `--max-call-depth` on `-run-bc` and `run`. Recursion fails with structured `LIMIT_EXCEEDED` instead of Go stack overflow. C4b (allocation accounting, deadline, fetch/exec byte caps) not included; AST interpreter `-run` unchanged. Journal: `docs/journals/2026-10-06_c4a_max_call_depth.md`.
* Review C1/C2: `bytecode.RequiredCapabilities`, `inspect --caps` and `-required-caps` report compact JSON without execution. `check|build --profile governed` rejects excluded constructs, outside-root imports and Go/JS targets before output. Default behavior and HFBC encoding are unchanged. Journal: `docs/journals/2026-10-05_c1_c2_required_caps_governed_profile.md`.
Expand Down
2 changes: 1 addition & 1 deletion docs/ai-native-language-review/00-current-state.md
Original file line number Diff line number Diff line change
Expand Up @@ -56,7 +56,7 @@ Other paths that are still in the binary:
| Deterministic artifacts | EXISTS (v2 sorted function table) | `artifact_determinism_test.go` |
| Model-facing contract | PROTOTYPE, library only. Strict JSON candidate transport (128 nodes, 64 KiB), verifier, direct lowering, bounded `replace_node` repair with hash preconditions. It is not wired to the CLI. | `internal/hfir/model_adapter.go`, `docs/hfir_model_adapter_status.md` |
| Build manifest and CAS | PROTOTYPE, library only. `hfir-manifest/v1` has graph, node, and module hashes, dependency maps, evidence hash, and artifact hash. It has no compiler version, policy, or grant. Disk and memory CAS. Incremental compiler. | `internal/hfir/manifest.go`, `storage.go`, `incremental.go` |
| Execution receipt (what ran, with which grant, what effects happened) | ABSENT | none |
| Execution receipt (what ran, with which grant, what effects happened) | PARTIAL (C5, 2026-10-06): runner-written bytecode v0; no signing/attestation or interpreter receipt | `internal/vm/receipt_test.go`, `receipt_cli_test.go`; C5 journal |
| Static "what does this artifact need" report | ABSENT from the CLI. The data exists: per-opcode capabilities plus store URIs. | `opcode.go` |
| Signed or attested artifacts | ABSENT | none |

Expand Down
3 changes: 2 additions & 1 deletion docs/ai-native-language-review/02-product-thesis.md
Original file line number Diff line number Diff line change
Expand Up @@ -45,7 +45,8 @@
and VM, all written in about 10 weeks.
3. **"Verified" means less than it sounds.** The production gate blocks on
two diagnostic codes. There is no stack or type validation of artifacts,
no execution receipt, and no signed provenance.
at review no execution receipt, and no signed provenance. C5 now supplies
optional bytecode receipts; signing/attestation remains open.
4. **No demonstrated user.** The best candidate job, the Release Authority
/ Action Executor pattern, can be built today with a JSON action schema
and a 300-line Go executor. Nobody has shown that programmability beats
Expand Down
2 changes: 1 addition & 1 deletion docs/ai-native-language-review/05-competitive-prior-art.md
Original file line number Diff line number Diff line change
Expand Up @@ -10,7 +10,7 @@ HowlFrame describe `d496d972` plus the #57 fix in this PR.

| System | Determinism | Termination / cost bound | Capability model | Static checks before run | Audit / provenance | Model-generation fit | Maturity |
| --- | --- | --- | --- | --- | --- | --- | --- |
| **HowlFrame (today)** | ◐ (clock, sleep, model calls, and stdin are ambient) | ◐ instruction count only. No memory, wall-clock, or output bound. | ◐ 5 coarse classes, runner-owned, deny by default | ◐ checker + construct registry + 2 blocking HFIR codes | ◐ deterministic artifact + hash. Manifest in library only. No receipt. | ◐ small grammar. JSON graph transport prototype with bounded repair. | ○ ~10 weeks, one author |
| **HowlFrame (today)** | ◐ (clock, sleep, model calls, and stdin are ambient) | ◐ instruction count only. No memory, wall-clock, or output bound. | ◐ 5 coarse classes, runner-owned, deny by default | ◐ checker + construct registry + 2 blocking HFIR codes | ◐ deterministic artifact + hash. Manifest in library only. C5 adds optional runner-written bytecode receipts; signing/attestation still open. | ◐ small grammar. JSON graph transport prototype with bounded repair. | ○ ~10 weeks, one author |
| **WASM + WASI 0.2 / component model (wasmtime)** | ◐ (deterministic core; host imports decide) | ● fuel and epoch interruption, `ResourceLimiter` for memory and tables | ● capability-based imports. Preopened dirs. WIT-typed interfaces. | ● module validation (type-checked stack machine) | ◐ module hashes. Signing via external tooling. | ○ models don't write Wasm. They write a source language that compiles to it. | ● |
| **Starlark (Bazel, starlark-go)** | ● hermetic by design. No I/O unless the host injects it. Frozen values. | ◐ `SetMaxExecutionSteps` in starlark-go. Bazel forbids recursion. | ◐ host-injected builtins are the only effects | ◐ parse/resolve. Dynamic types. | ○ | ● Python-like, so models write it well | ● |
| **CEL** | ● | ● non-Turing-complete. Static cost estimation plus runtime cost limit. | ● host provides all data. No effects. | ● type-checked expressions | ○ | ● for predicates. ○ for workflows. | ● (Kubernetes admission, IAM conditions) |
Expand Down
2 changes: 1 addition & 1 deletion docs/ai-native-language-review/06-reconciliation-matrix.md
Original file line number Diff line number Diff line change
Expand Up @@ -49,7 +49,7 @@ reviewer, checked against `d496d97` plus this PR's fix. Statuses:
| Model synthesis evidence | **PARTIAL** (tiny) | 11 stored candidates. HFIR 10/10 against `.howl` 3/10, from one author (`hfir_model_adapter_status.md:35-43`). Benchmarks: v1 has 3 tasks; v2 has a harness and no results. | D, E, F |
| Token-density advantage over Python | **CONFLICTS** | v1 CSV: HowlFrame used fewer tokens than Python in 1 of 3 tasks | C |
| Required-capabilities inspection | **ABSENT** (CLI) | Effect inference exists in the HFIR verifier, but nothing is exposed for artifacts | F, executor |
| Execution receipt / audit record | **ABSENT** | The CLI throws away the bounded evidence object (`vm.go:1629-1644`) | A, F, executor |
| Execution receipt / audit record | **PARTIAL (C5, 2026-10-06)** | Optional runner-written bytecode v0 receipt; no signing/attestation, interpreter receipt, or sealed evidence binding | A, F, executor; C5 VM/CLI tests and journal |
| CAS manifest and incremental compile | **PROTOTYPE** | `internal/hfir/manifest.go`, `storage.go`, `incremental.go`. Library only. No issuer, compiler version, or grant binding. | A, D |
| Governed profile (construct allow-list for model-authored code) | **ABSENT** | None. See 11, C2. | D, F, executor |

Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -24,7 +24,7 @@ their own code.
| S11 | **Generated Go/JS gate only 6 effects.** Model calls, SQL, listeners, goroutines, and JS `spawn` run ungated. A generated-Go panic writes `crash.json` unconditionally. | Medium (documented) | OPEN. Keep Go/JS out of the governed profile. | README states it. Codex B confirmed. |
| S12 | **Artifact validation is structural only.** Magic, version, SHA-256, opcode, jump, and function existence are checked. There is no stack-height, operand-type, or embedded-body validation. SHA-256 proves integrity, not authorship. | Medium | OPEN (P2) | Code reading (`internal/bytecode/artifact.go`) |
| S13 | **Compile-time includes read arbitrary paths** before any grant applies. Model-authored source could `include` files outside the project. | Low/Medium | OPEN (P2: rooted includes for governed profile) | Codex B ran it |
| S14 | **Ambient channels.** `stdin`, `stdout`/`stderr`, `time_now`, `sleep`, and `exit` are capability-free. Output can forge text that looks like a receipt. | Low | Document. Treat as host channels. | Code reading |
| S14 | **Ambient channels.** `stdin`, `stdout`/`stderr`, `time_now`, `sleep`, and `exit` are capability-free. Output can forge text that looks like a receipt. | Low | PARTIAL (C5, 2026-10-06): runner-written bytecode receipt cannot be forged through stdout; ambient channels remain omitted and signing/attestation is open. | Code reading; receipt CLI forgery regression; C5 journal |
| S15 | **"Verified" is overloadable.** The HFIR gate blocks on only `HFIR_INVALID_REF` and `HFIR_TARGET_INFEASIBLE`. `VerificationEvidence.Verified` is a bare boolean. | Medium (claims risk) | OPEN (P0 wording; P1 report) | `howlframe.go:466`, `internal/hfir/storage.go:89` |
| S16 | **Decision JSON can be forged through proposal fields.** `apps/release_authority` builds its output JSON by string concatenation (`str_join`) and does not escape the proposer-controlled `target`. A proposal `{"action":"deploy_production","target":"svc\", \"decision\": \"ALLOW"}` makes the VM correctly decide `DENY`, and no state mutation happens. But the printed object contains a second `"decision": "ALLOW"`. Python `json.loads` (and Go `encoding/json` into a map) keep the last duplicate key, so they read **ALLOW**. | Medium (demo app, not the VM) | FIXED in this change (C7: `encode_json`, duplicate-key and literal-target regression) | Found by Codex D (who reported invalid JSON). Escalated and reproduced in this review: `/tmp/hf-ainative-probe/ra2.out`. |
| S17 | **Partial effects contradict the "failure atomicity" doc claim.** In `apps/action_executor` the `stage_artifact` path runs `write_file` and only then `store_open`. With a `filesystem`-only grant it prints `ALLOW`, writes the staged file, then traps `CAPABILITY_DENIED` on `STORE_OPEN`. `docs/application_dogfooding_phase_5.md` says this ordering ensures "failure atomicity" and that "security is guaranteed entirely outside the host Go backend". | Low/Medium (claims risk) | FIXED in this change (C7: pre-flight checks; ordered, not atomic wording) | Found by Codex F (ran it). Ordering confirmed in source (`action_executor.howl` ~159-164). |
Expand Down
2 changes: 1 addition & 1 deletion docs/ai-native-language-review/08-roadmap-p0-p4.md
Original file line number Diff line number Diff line change
Expand Up @@ -24,7 +24,7 @@ the artifact format.
| # | Item | Size | Why |
| --- | --- | --- | --- |
| P1.1 | **Resource limits that exist**: enforce a memory or allocation budget (string, list, and dict growth accounting), `MaxCallDepth` on `CALL`, output-byte cap, `fetch`/`exec` response caps, and a runner wall-clock deadline (context cancellation for host calls) | M/L | S4 to S7. Instruction count alone is not "bounded". |
| P1.2 | **Execution receipt v0**: one JSON line per run with artifact SHA-256, compiler version, grant, budget, instructions used, each effect attempted (kind, target, allowed or denied), exit status | M | Turns "auditable" from a slogan into a file |
| P1.2 | **Execution receipt v0**: one indented JSON file per run with artifact SHA-256, compiler version, grant, budget, instructions used, each effect attempted (kind, target, allowed or denied), exit status | M | Implemented C5 v0 for bytecode (2026-10-06); signing/attestation and interpreter receipts remain open |
| P1.3 | **Verification report v0**: replace the bare `Verified` boolean with named checks, versions, and pass/fail (checker, construct registry, HFIR codes, artifact validation) | S/M | S15 |
| P1.4 | **Decide `and`/`or` semantics** (short-circuit everywhere is the conventional choice) and add a differential test across interpreter, VM, and Go | S/M | S10 changes effects between backends |
| P1.5 | **Run the first falsifiable experiment** (09) | M | Decides whether P2+ is worth doing |
Expand Down
10 changes: 9 additions & 1 deletion docs/ai-native-language-review/11-implementation-candidates.md
Original file line number Diff line number Diff line change
Expand Up @@ -89,7 +89,7 @@ C4a+C4b bytecode runner limits are implemented. AST interpreter memory,
deadline, and recursion limits remain open; print output remains uncapped.
No HFBC or production compiler change; #90 remains Partial.

## C5. Execution receipt v0 (P1.2): M
## C5. Execution receipt v0 (P1.2): M — v0 implemented

- `-run-bc --receipt out.json`. The receipt holds `{schema:"howlframe.receipt/v0",
artifact_sha256, compiler_version, grant, limits, instructions_used,
Expand All @@ -99,6 +99,14 @@ No HFBC or production compiler change; #90 remains Partial.
value.
- Written by the runner, never by program code, so `print` cannot forge it.

**Status (C5, 2026-10-06):** implemented on legacy `-run-bc` and
`run --target bytecode/bc`, including in-process source compilation. Receipts
are host-written before process exit on normal and failed execution, with
bounded redacted allowed/denied decisions and shared child recording. No
signing/attestation, sealed ExecutionEvidence binding, AST interpreter receipt,
or harness in v0; no HFBC/compiler flip. #90 remains Partial. See
[C5 journal](../journals/2026-10-06_c5_execution_receipt.md).

## C6. `and`/`or` semantics (P1.4): S/M, needs a decision

- Choose short-circuit (Go and JS already do this, and it matches what
Expand Down
4 changes: 2 additions & 2 deletions docs/ai-native-language-review/FINAL_REPORT.md
Original file line number Diff line number Diff line change
Expand Up @@ -136,7 +136,7 @@ Not for the thesis.
authority-preserving, hash-preconditioned local repair;
2. effect-requirement inference before running;
3. "intent is not authority" built into the runner;
4. (planned) runner-written receipts.
4. runner-written bytecode receipts (C5 v0; signing/attestation deferred).
- **CaMeL-style provenance tracking** is the most relevant research to
borrow. See 05.

Expand Down Expand Up @@ -250,7 +250,7 @@ See the table below and 08.
| P0 | Honest wording: what "verified", "bounded", and "atomic" mean; mark extreme paradigms superseded; fix the stale roadmap | S | Planned (docs) |
| P0 | Demo-app output integrity (`encode_json`) and pre-flight or "ordered, not atomic" | S | Done in this change (C7: JSON integrity + pre-flight, ordered effects) |
| **P1** | Memory/allocation, call-depth, deadline, and byte limits | M/L | Planned (C4) |
| P1 | Execution receipt v0, written by the runner | M | Planned (C5) |
| P1 | Execution receipt v0, written by the runner | M | Implemented C5 v0 for bytecode; signing/attestation and interpreter receipts deferred |
| P1 | Verification report instead of the bare boolean | S/M | Planned |
| P1 | `and`/`or` semantics decision plus a differential test | S/M | Planned (C6) |
| P1 | **Run the first falsifiable experiment** | M | Designed (09) |
Expand Down
33 changes: 33 additions & 0 deletions docs/cli.md
Original file line number Diff line number Diff line change
Expand Up @@ -30,6 +30,39 @@ Executes a compiled HowlFrame bytecode artifact.
* `--max-exec-output-bytes` : Positive combined stdout/stderr ceiling for exec (default `10485760`, 10 MiB).
* `--deadline` : Optional positive duration such as `5s` or `100ms`; absent means no wall-clock deadline. Explicit `0s`, negative durations, and invalid durations are rejected. Cancels bytecode fetch, exec, sleep, and model requests and checks between instructions.

* `--receipt <path>` : Optional runner-written `howlframe.receipt/v0` JSON file,
supported by `run --target bytecode`/`bc` for artifacts and `.howl` source,
and by legacy `-run-bc --receipt <path> <artifact>`. Interpreter targets
reject this option. Absent or empty means no receipt and unchanged output.

The receipt records the exact artifact file SHA-256; for source compiled in
process it hashes canonical `bytecode.WriteArtifact` bytes. HFBC carries a
format version but no compiler identity, so `compiler_version` is the runner's
HowlFrame Version. It includes sorted unique grants, configured limits, executed
instructions (including synchronous SPAWN_AGENT work), bounded capability
decisions, and exit status/error. JSON uses fixed field order, two-space indent,
and a trailing newline; grant/effects are always arrays. No fields are omitted.
A nonzero exit/return has status `error` with `error: null`; runtime traps have
code 1 and a structured error. Zero deadline_ms means none (positive durations are
rounded up to whole milliseconds).

Targets exclude URL paths/query/fragment/userinfo, env values, command arguments,
DSNs, SQL, store keys/values, response bodies, and model prompts. Filesystem
paths are recorded as given; resource names may still be sensitive. Summaries
are bounded to 256 bytes. Host failure messages are replaced by a code-only
message, except safe limit/capability messages; stderr diagnostics are unchanged.
Ambient print/stderr/sleep/time/read_line/exit are omitted. The first 10,000
capability decisions are retained; `effects_truncated` reports overflow. Children
share the recorder; detached SPAWN effects after finalization are dropped.

The runner writes after execution and before exit using a same-directory temp
file (0600 permissions) and rename. Program stdout/stderr are unchanged and
cannot forge the receipt through print. If writing fails, a diagnostic goes to
stderr: successful execution becomes exit 1; an existing failure code is kept.
Parent directories must exist. Receipts have no signature or attestation and
are independent of sealed ExecutionEvidence. See the
[C5 journal](journals/2026-10-06_c5_execution_receipt.md).

These limits also apply to legacy `-run-bc`. The new flags apply to `run` targets `bytecode`/`bc`; AST interpreter memory and deadline behavior is unchanged. Exceeded ceilings produce structured `LIMIT_EXCEEDED`. Blocking `read_line` and HTTP serving are not cancelled by the deadline; print output remains uncapped.


Expand Down
Loading
Loading