Unsafe baseline
The frozen synthetic workload traversed 17,599 unavailable routes, reached the mock control plane on attempt 17,600 and changed the simulated release from hold to deployed.
We built a safe synthetic 17,600-attempt workload to test that question. The intentionally unsafe baseline kept exploring until the final route executed a simulated privileged action. The protected run used the same workload but was contained at attempt 26.
A single failed path may be harmless. A persistent agent can keep searching, retrying and crossing reachable trust boundaries until several ordinary weaknesses connect. The benchmark tests whether controls can bound that persistence before a consequential action.
The frozen synthetic workload traversed 17,599 unavailable routes, reached the mock control plane on attempt 17,600 and changed the simulated release from hold to deployed.
The identical workload kept the established denied-path threshold at 25. The breaker opened and attempt 26 was blocked. No synthetic credential was exposed and no privileged action executed.
The workload did not become easier. The tested control changed the observed outcome. We did not weaken the control merely to force the protected run to reach 17,600 attempts.
It proves the bounded synthetic behaviour, evidence lineage and deterministic protected outcome under this frozen benchmark—not production protection or prevention of a real-world incident.
Claims below come from the validated Phase 8 run and its SHA-256-bound evidence manifest.
deployedholdARL17K was built as a staged evidence chain: observe the unsafe condition, add containment, record remediation, replay the exact workload, verify recovery, preserve a human deployment decision and then scale the workload.
Establish the unsafe synthetic route and separate agent activity from lab-side observation.
Apply velocity counting, a circuit breaker and exact-action policy fallback.
Record the confirmed condition, remediation owner and exact retest.
Replay the frozen workload digest and verify the changed outcome.
Stop, revoke, isolate, restore and probe the recovered state.
An operator separately recorded hold; AgentRiskLayer did not make the deployment decision.
Repeat at 1,000 attempts before the final 17,600-attempt workload.
Hash the final workload, evidence streams and result into an evidence manifest.
The canonical workload digest hashes compact JSON. The saved workload artifact is pretty-printed, so its file hash is intentionally different.
SHA-256 of the frozen 17,600-attempt workload in canonical compact form.
SHA-256 of the pretty-printed workload manifest file.
SHA-256 of the 17,607-line baseline observer evidence stream.
SHA-256 of the 56-line protected observer evidence stream.
SHA-256 of the final benchmark result artifact.
SHA-256 of the final artifact manifest that records the evidence boundary.
The benchmark is deliberately bounded so a buyer can see what the result means—and where it stops.
hold.Start with one agent. Map its access, data, tools, approval and recovery state. Unknowns remain evidence gaps; findings require support.