You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
A backend that refused a policy exited -1 with a generic backend_error, so a caller could not
tell a refused policy from a crash, a launch failure or a timeout.
FailurePhase now derives the wire error code, and the executor envelope and the engine share
that one mapping:
phase
code
exit
Rejected
policy_validation
1
BackendUnavailable
backend_unavailable
-1
everything else
backend_error
-1
Every ScriptResponse failure site across Bubblewrap, LXC, Seatbelt, process container,
IsolationSession, Nanvix, Hyperlight, and Windows Sandbox is tagged with the phase that matches its cause:
rejected — the request itself is refusable; only a changed request can succeed.
unavailable — the host cannot run this backend, so a caller may fall back a tier.
error — a runtime failure (spawn, syscall, IO, crash, invariant).
Genuinely mixed cases stay error. Mislabelling infrastructure as policy_validation is the
failure this contract exists to prevent.
From review:
Rejected is policy-only. Windows Sandbox mapped SandboxUnavailable onto it, reporting a
disabled optional feature as policy_validation; it now uses BackendUnavailable.
emit_captured_stderr no longer suppresses a workload's own stderr. Windows Sandbox mirrors a
failing guest's stderr into error_message, so the de-duplication swallowed real output; it now
applies only to phases where no workload ran.
validate_legacy_host_lists refusals are tagged, matching the rest of the shared validator.
captured_stderr_to_emit is pure, so the tests assert the exact bytes emitted for both the
suppressed and the relayed case.
P13b was a false green: its inheritDefaultEnv=false case supplied a verbatim env block missing
SYSTEMROOT/LOCALAPPDATA, which was always refused — the -1 exit hid it. The fixture now uses Get-MinimalEnv, so the pair varies only by the field under test.
Exit-code expectations move from -1 to 1 where a refusal is now typed: wslc_destroy_on_exit_false_rejected, two IsolationSession one-shot cases, and the Node Seatbelt blockedHosts test (backend_error -> policy_validation).
Concistency Fixes
The mapping above only bound the surfaces the bug touched. A sweep of every backend found six
paths that still disagreed about what a failure means:
Executor entry. A failure raised before a ScriptResponse exists printed a bare stderr line
and exited 1 — indistinguishable from a policy refusal. wxc, lxc and mxc_darwin now emit a
typed envelope through emit_mxc_error_exit.
IsolationSession timeout. One-shot never applied with_service_timeout_grace, so the service
timer won the race and a timeout arrived as an ordinary exit 1, colliding with the rejection code. create_process now returns ExecOutcome with a split deadline, matching piped_exec_handle.
nanvix / hyperlight. A single Preflight variant collapsed policy refusals, missing shipped
artifacts and environment faults into one code; each splits into Policy / Unavailable / Setup.
Spent deadlines. wslc, LXC and Windows Sandbox reported a watchdog kill as the workload's own
exit. LXC gains a typed AttachError and Windows Sandbox a Termination tag on ExecResult, so
the kill survives the call instead of being flattened into a sentinel.
Availability. Explicitly selecting LXC skipped the probe discovery runs; macOS claimed
Seatbelt without checking sandbox-exec. Both now fail as backend_unavailable up front.
Feature gates. A disabled build feature is backend_unavailable; a missing experimental
opt-in is malformed_request. These were previously swapped in places.
cargo fmt --all -- --check and cargo clippy --workspace --all-targets -- -D warnings are clean. cargo test on the touched crates passes against a stashed baseline: wxc_common 905/49 and windows_sandbox_lifecycle 184/0 each gain exactly the one test added here, and process_container_common matches baseline at 315/1. isolation_session (195), lxc_common (355), mxc_engine (153), nanvix_runner (44) and wslc_common (263) are fully green. The remaining
local failures are pre-existing and environment-dependent (telemetry/registry, PSEC RPC,
reparse-point tests) in modules this PR does not touch. Seatbelt and hyperlight do not build on
this host, so their changes are covered by CI. Host suites run separately (P8a/P8c/P8d/P8e, P13b).
Audit every ScriptResponse failure-construction site across the backends and
tag each with the phase that matches its cause, so the wire error code derived
from FailurePhase is truthful:
- rejected: the request itself is refusable; only a changed request succeeds.
- unavailable: the host cannot run this backend, so a caller may fall back.
- error: a runtime failure (spawn, syscall, IO, crash, invariant).
Covers bubblewrap, lxc, seatbelt, process_container, isolation_session, and
windows_sandbox. Genuinely mixed cases stay as error, since mislabelling
infrastructure as policy_validation is the failure this contract exists to
prevent.
Also addresses PR review:
- FailurePhase::Rejected is now policy-only. Windows Sandbox mapped
SandboxUnavailable onto it, which reported a disabled optional feature as
policy_validation; it now uses BackendUnavailable.
- emit_captured_stderr no longer suppresses a workload's own stderr. Windows
Sandbox mirrors a failing guest's stderr into error_message, so the
de-duplication swallowed real output; it now applies only to phases where no
workload ran.
- validate_legacy_host_lists refusals are tagged as rejections, matching the
rest of the shared validator.
- captured_stderr_to_emit is pure, so the tests assert the exact bytes emitted
for both the suppressed and relayed cases.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
The reason will be displayed to describe this comment to others. Learn more.
Copilot review overview
🟡 Changes recommended
Seatbelt currently classifies the same invalid working directory differently by launch method, and stderr relay introduces an avoidable unbounded copy.
An embedded NUL can only come from the caller-supplied/generated profile (not from a runtime failure), so this still reports a policy refusal as backend_error. A seatbelt.profileOverride containing \u0000, for example, can only succeed after the request changes and should produce policy_validation under the new contract.
Emit policy validation envelope before rejected dry-run exit
src/core/wxc_common/src/script_runner.rs:69
A rejected dry run now exits with 1, but this function terminates before every executor reaches emit_backend_error_envelope. Because validation itself does not write the rejection reason to the logger, dry-run callers receive neither the promised policy_validation envelope nor the reason. Emit the response stderr/envelope here before exiting (the helpers are no-ops for a successful dry run).
Avoid appending newline when relaying workload stderr
src/core/wxc_common/src/script_runner.rs:115
This appends a newline to captured workload stderr even when no diagnostic follows. For example, a successful LXC workload that writes progress without a newline has error_message empty, so the old eprint! preserved progress but this helper emits progress\n. Only request an added newline when envelope_applies(response); otherwise relay the workload bytes unchanged, and update the helper test accordingly.
- A rejected dry run exited 1 with neither the reason nor its envelope: every
executor calls handle_dry_run_exit before reaching the relay helpers, and it
diverges. It now emits both before exiting; each is a no-op for a dry run
that passed.
- Captured workload stderr is relayed exactly as written. The terminating
newline is added only when the error envelope follows and would otherwise
start mid-line, so unterminated progress output no longer gains a newline it
never wrote.
- An embedded NUL in a Seatbelt profile is a rejection: the profile is built
from the request, verbatim for seatbelt.profileOverride, so only a changed
request can succeed.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
This rejection is not reached for BaseContainerRunner::new(): on an API-capable host, can_backend_service_request returns PolicyIncompatible for least_privilege_mode, so the preceding branch reports backend_unavailable; dry-run also returns before this check. Move the caller-fixable policy check ahead of both branches so direct and dry-run callers receive policy_validation.
Preserve workload stderr after post-launch relay failures
src/core/wxc_common/src/script_runner.rs:100
PostLaunchFailed explicitly includes failures “while running user code” (models.rs:1314-1317), so this predicate can say the workload did not run after it already emitted stderr. If a relay failure preserves and mirrors that partial stderr into error_message, captured_stderr_to_emit drops the workload bytes. The response needs an explicit stderr-origin/workload-started signal (and diagnostic constructors should use it) rather than inferring this from an ambiguous phase.
Classify empty commands as malformed, not policy validation
src/core/wxc_common/src/validator.rs:215
Rejected is documented as a policy refusal, but an empty command is structural malformed input. The state-aware equivalent already returns malformed_request in validator.rs:246-250; this path will instead emit policy_validation. Keep this unclassified until ScriptResponse can represent malformed input, rather than mislabeling it as policy.
BaseContainerRunner::validate ran the leastPrivilege refusal after both the
dry-run return and the serviceability probe, so a caller-fixable policy was
reported as backend_unavailable on an API-capable host and not reported at
all on a dry run. BaseContainerRequestDecision already separates
PolicyIncompatible from the host-capability verdicts; honour that split
instead of collapsing every non-serviceable decision into unavailable.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Both backends funnelled every error into a ScriptResponse carrying only
exit -1, so a watchdog kill was indistinguishable from a spawn failure --
the inference FailurePhase exists to prevent. Derive the phase from each
error variant at the single to_response funnel.
Hyperlight folded RunError::TimedOut into a generic Runtime error, losing
the timeout outright; give it a Timeout variant that renders identical text.
Preflight stays unclassified in both: it mixes missing host artifacts with
refused policy, so it cannot honestly claim either phase.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
A caller-supplied Seatbelt filesystem path containing .. is rejected only inside build_profile_with_proxy (profile_builder.rs:729-755), but this maps every builder failure to an unclassified ScriptResponse::error. That request therefore still exits -1 with backend_error, contrary to the new rejection contract. Because this builder also reports host/environment failures such as a missing HOME, preserve those as backend errors while giving policy-build failures a typed result that maps to ScriptResponse::rejected.
reject_unhonorable_environment raises IsolationSessionError::Policy, which
now exits 1 like the ui and lifecycle refusals beside it. The expectation
was left at -1 when those two were updated, so the case would have failed
deterministically on a host run.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
With --audit, audit::finalize can fail after the runner returned Rejected; the error branch at lines 1686-1694 changes exit_code and error_message but leaves failure_phase as Rejected. This mapping then exits 1 (and the preceding envelope emits policy_validation) even though the final failure includes audit infrastructure. Reset the phase to an infrastructure phase when audit finalization fails so this mixed case remains backend_error/-1.
lxc_runner.rs conflicted wholesale because main normalized it from CRLF to
LF (#1334) while this branch changed its content. Main made no content
change to the file, so the resolution keeps this branch's rejection
classifications and adopts main's LF line endings.
A rejection reached the OS exit code only through process_exit_code, so
display_script_results and completion telemetry read -1 from the response
while the process exited 1.
FailurePhase::mxc_exit_code is now the single source for the code MXC
reports for its own failures, and every producer that builds a Rejected
response uses it.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
The phase-to-code mapping landed earlier only bound the surfaces the bug
touched. Six paths still disagreed about what a given failure means.
- Executors that fail before a ScriptResponse exists now emit a typed
envelope via emit_mxc_error_exit instead of a bare stderr line + exit 1.
- IsolationSession one-shot never applied the service-timeout grace, so
the service timer won the race and a timeout surfaced as exit 1 --
colliding with the rejection code. create_process now returns
ExecOutcome with a split deadline, matching piped_exec_handle.
- nanvix and hyperlight collapsed policy refusals, missing artifacts and
environment faults into one Preflight variant; each now splits into
Policy / Unavailable / Setup.
- wslc, lxc and windows_sandbox reported a spent deadline as an ordinary
workload exit. LXC gains a typed AttachError and windows_sandbox a
Termination tag on ExecResult so the kill survives the call.
- Explicit LXC selection now runs the same probe discovery uses, and
macOS validates sandbox-exec before claiming the backend.
- A disabled feature is backend_unavailable; a missing experimental
opt-in is malformed_request.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
The in-process sandbox handle merged from main carried two failure sites
that predate the classification contract.
- Refusing StdioMode::Inherit is a refusal, not a fault: only a changed
call can succeed, and it does so on every host. IsolationSession
already returns `rejected` for the identical condition.
- A non-Linux host cannot launch an LXC sandbox at all, which is a host
prerequisite rather than a bad request, so it reports
backend_unavailable and a caller may fall back a tier.
The two junction deny tests asserted the pre-contract exit code. They
were passing only because their `mklink /J` guard was skipping them on
an unprivileged host; both now run and assert the phase.
Document the constructor-to-code-to-exit mapping in the Copilot
instructions so a new failure site is classified rather than defaulted.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Elliot (theelliotm)
changed the title
Exit 1 with a typed code when a backend rejects a policy
Exit codes are consistent cross-backend w/ structured ScriptResponse JSON
Sep 30, 2026
- A hyperlight mount path whose parent is missing is host-filesystem
state, not a refusal: the message itself says `mkdir -p` fixes it, so
the same request succeeds elsewhere. It moves from Policy to Setup.
- The IsolationSession and LXC timeout branches built their diagnostic
with ScriptResponse::error, which copies the message into standard_err.
Because a timed-out workload did run, the executor relays standard_err
bare and the envelope then printed the same text again. Both now keep
the diagnostic in error_message only, matching the timeout responses in
mxc_engine and sandbox_process.
- emit_mxc_error_exit wrote the buffered log to stderr, ahead of the
envelope, leaving stderr unparseable whenever anything had been logged.
The buffer goes to stdout, as handle_dry_run_exit already does.
- An empty command reported policy_validation from the one-shot path and
malformed_request from the state-aware one. FailurePhase cannot gain a
variant -- it crosses the WSLc daemon IPC boundary, where a cached image
may predate the host -- so ScriptResponse carries an optional code that
outranks the phase, and the two surfaces now agree.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…he wire code
Every branch of the seatbelt working-directory check reports host state -- a
path that is a file, a missing search bit, or an absent directory -- so the
same request can succeed once the host is repaired. That is a mixed cause,
which the contract keeps as a backend error rather than a policy refusal.
Completion telemetry now classifies from a response's explicit wire code when
it carries one, so ScriptResponse::malformed records ConfigError to match the
state-aware path instead of the phase's generic PolicyError.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Aligning the state-aware gate with run-to-completion changed a missing
experimental opt-in from backend_unavailable to malformed_request, but
only the engine's own tests were updated. mxc_ffi and mxc-sdk assert the
same gate through the C ABI and the Rust SDK, so every platform job
failed on six stale expectations.
The SDK rustdoc documented the old code as well: the exec_sandbox
opt-in refusal, and a TimedOut note that collapsed the opt-in and the
feature-off refusals into one code. They are now distinct.
Drops two rationales that the alignment made obsolete -- a wslc id no
longer shares a code with the gate, and the FFI test can no longer
separate "passed the gate" from "parsed the request".
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
ExecOutcome::Exited also carries nonzero workload exits, but this response defaults every such result to FailurePhase::None. Set ProcessExited when exit_code != 0; otherwise the serialized response says there was no failure even though the workload failed.
Report trusted directory API failures as backend errors
find_bfscfg_exe() returns Err only when the trusted Windows-directory API itself fails; a missing binary is Ok(None) and is handled separately. Reclassifying this runtime API failure as backend_unavailable incorrectly invites tier fallback instead of reporting the infrastructure fault as backend_error.
A normal exit is not always successful: collect_output already has a test passing exit code 42, but this new match labels it FailurePhase::None. That breaks the stated phase contract and prevents consumers from distinguishing a workload failure from success; use ProcessExited for nonzero exited outcomes.
Distinguish disabled x86_64 feature from unsupported architecture
src/core/mxc_engine/src/run.rs:557
This branch combines two different cases and gives an incorrect message on x86_64 builds where only the feature is disabled: the host already is x86_64, yet the error says x86_64 is required. Split feature-disabled x86_64 from unsupported architectures so the former reports how to enable the build feature and the latter remains unsupported_containment as documented above.
The reason will be displayed to describe this comment to others. Learn more.
This branch has not been deployed
No deployments
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
📖 Description
A backend that refused a policy exited -1 with a generic
backend_error, so a caller could nottell a refused policy from a crash, a launch failure or a timeout.
FailurePhasenow derives the wire error code, and the executor envelope and the engine sharethat one mapping:
Rejectedpolicy_validationBackendUnavailablebackend_unavailablebackend_errorEvery
ScriptResponsefailure site across Bubblewrap, LXC, Seatbelt, process container,IsolationSession, Nanvix, Hyperlight, and Windows Sandbox is tagged with the phase that matches its cause:
Genuinely mixed cases stay
error. Mislabelling infrastructure aspolicy_validationis thefailure this contract exists to prevent.
From review:
Rejectedis policy-only. Windows Sandbox mappedSandboxUnavailableonto it, reporting adisabled optional feature as
policy_validation; it now usesBackendUnavailable.emit_captured_stderrno longer suppresses a workload's own stderr. Windows Sandbox mirrors afailing guest's stderr into
error_message, so the de-duplication swallowed real output; it nowapplies only to phases where no workload ran.
validate_legacy_host_listsrefusals are tagged, matching the rest of the shared validator.captured_stderr_to_emitis pure, so the tests assert the exact bytes emitted for both thesuppressed and the relayed case.
P13b was a false green: its
inheritDefaultEnv=falsecase supplied a verbatim env block missingSYSTEMROOT/LOCALAPPDATA, which was always refused — the -1 exit hid it. The fixture now uses
Get-MinimalEnv, so the pair varies only by the field under test.Exit-code expectations move from -1 to 1 where a refusal is now typed:
wslc_destroy_on_exit_false_rejected, two IsolationSession one-shot cases, and the Node SeatbeltblockedHoststest (backend_error->policy_validation).Concistency Fixes
The mapping above only bound the surfaces the bug touched. A sweep of every backend found six
paths that still disagreed about what a failure means:
ScriptResponseexists printed a bare stderr lineand exited 1 — indistinguishable from a policy refusal.
wxc,lxcandmxc_darwinnow emit atyped envelope through
emit_mxc_error_exit.with_service_timeout_grace, so the servicetimer won the race and a timeout arrived as an ordinary exit 1, colliding with the rejection code.
create_processnow returnsExecOutcomewith a split deadline, matchingpiped_exec_handle.Preflightvariant collapsed policy refusals, missing shippedartifacts and environment faults into one code; each splits into
Policy/Unavailable/Setup.exit. LXC gains a typed
AttachErrorand Windows Sandbox aTerminationtag onExecResult, sothe kill survives the call instead of being flattened into a sentinel.
Seatbelt without checking
sandbox-exec. Both now fail asbackend_unavailableup front.backend_unavailable; a missing experimentalopt-in is
malformed_request. These were previously swapped in places.🔗 References
Contributes to #612.
🔍 Validation
cargo fmt --all -- --checkandcargo clippy --workspace --all-targets -- -D warningsare clean.cargo teston the touched crates passes against a stashed baseline:wxc_common905/49 andwindows_sandbox_lifecycle184/0 each gain exactly the one test added here, andprocess_container_commonmatches baseline at 315/1.isolation_session(195),lxc_common(355),mxc_engine(153),nanvix_runner(44) andwslc_common(263) are fully green. The remaininglocal failures are pre-existing and environment-dependent (telemetry/registry, PSEC RPC,
reparse-point tests) in modules this PR does not touch. Seatbelt and hyperlight do not build on
this host, so their changes are covered by CI. Host suites run separately (P8a/P8c/P8d/P8e, P13b).
✅ Checklist
Cargo.lock, thedependency-feed-checkcheck passes (see docs/pull-requests.md)📋 Issue Type
Microsoft Reviewers: Open in CodeFlow