Where does OpenShell draw the boundary between policy correctness and operational correctness? #2936
Replies: 1 comment
|
Hi. I'm the solo developer of Agentic Security Harness. The question in this discussion is close to what I'm working on: how to check not just that a policy is correct, but that the system actually follows it while an agent is running. I've run multi-step experiments with a local Prometheus model: it chooses an action, receives a result or a denial, and continues the task. The test rig records the model's proposal, the control decision, and actual execution separately. In one series, the research wrapper around Harness handled 43 proposals: 39 pure operations executed and four premature actions were blocked. Two of eight tasks completed; repeated permitted steps without completing the task were a noticeable problem in the others. That's why I want to check both sides: the forbidden action did not happen, and a permitted path remained usable. This series is still kept locally; I haven't tested OpenShell with my rig yet. What integration point would you recommend for an external verification rig? Are OCSF events and a snapshot of the effective policy enough to link a particular request, decision, and result, or is another interface intended for that? I'd be interested in starting with a small reproducible example: a permitted operation, a denial, and continuation after that denial, with no external actions. If such a test set already exists, I'd appreciate a pointer so I can build on it. Thanks! |
Uh oh!
There was an error while loading. Please reload this page.
I have been looking through OpenShell’s current architecture, particularly the relationship between the policy engine, the policy advisor/prover, the sandbox supervisor, and the actual runtime enforcement path.
One distinction in the documentation caught my attention.
The policy prover evaluates the policy as written and can identify changes in reachability or capability, while the runtime proxy and sandbox are responsible for enforcing that policy during operation. As I understand it, the prover is intentionally a change-review mechanism rather than a claim that the as-operated runtime has independently demonstrated compliance with the policy model.
That seems like an important and correct separation.
It also raises a question I would be interested in hearing the maintainers’ view on:
Where does OpenShell consider the responsibility for verifying that the operating system actually behaves according to the declared policy to reside?
For example, there appear to be at least two different engineering claims:
The first seems increasingly well addressed by policy validation and the prover.
The second seems to require evidence from the operating system itself: deliberate denial tests, restart/recovery cases, policy-generation changes, credential-path tests, bypass attempts, and comparison of the effective runtime state against the declared state.
I am not suggesting that OpenShell itself must also be the independent evaluator of that second claim. In fact, having the implementation certify its own correctness would seem to weaken the value of the evidence.
So the architectural question is really about ownership:
I am asking because OpenShell already appears to make a useful distinction between what the policy permits and what the runtime enforces.
Understanding where the verification of that relationship belongs would help clarify the boundary between policy assurance, runtime assurance, and the governance system around both.
I would welcome corrections if I have misunderstood the intended responsibility split.I
All reactions