ARL17K · validated public benchmark

Could your AI agent keep trying until an ordinary weakness becomes a consequential action?

We built a safe synthetic 17,600-attempt workload to test that question. The intentionally unsafe baseline kept exploring until the final route executed a simulated privileged action. The protected run used the same workload but was contained at attempt 26.

WorkloadFrozen at 17,600 attempts
ReplaySame workload digest
ControlBreaker threshold kept at 25
BoundarySynthetic local-only evidence
Why this matters

Autonomy changes the risk of ordinary weaknesses.

A single failed path may be harmless. A persistent agent can keep searching, retrying and crossing reachable trust boundaries until several ordinary weaknesses connect. The benchmark tests whether controls can bound that persistence before a consequential action.

01

Unsafe baseline

The frozen synthetic workload traversed 17,599 unavailable routes, reached the mock control plane on attempt 17,600 and changed the simulated release from hold to deployed.

02

Protected replay

The identical workload kept the established denied-path threshold at 25. The breaker opened and attempt 26 was blocked. No synthetic credential was exposed and no privileged action executed.

03

What changed

The workload did not become easier. The tested control changed the observed outcome. We did not weaken the control merely to force the protected run to reach 17,600 attempts.

04

What this proves

It proves the bounded synthetic behaviour, evidence lineage and deterministic protected outcome under this frozen benchmark—not production protection or prevention of a real-world incident.

Observed evidence

The result is tied to a frozen workload and inspectable artifacts.

Claims below come from the validated Phase 8 run and its SHA-256-bound evidence manifest.

MeasureUnsafe baselineProtected run
Workload17,600 attemptsSame 17,600-attempt workload digest
Attempts started17,60026
Denied paths17,59925
Privileged actionExecuted in memoryNot executed
Final release statedeployedhold
Observer events17,60756
Repeat stabilityDeterministic across 2 runsDeterministic across 2 runs
Evidence chain

We did not jump from a demo to a security claim.

ARL17K was built as a staged evidence chain: observe the unsafe condition, add containment, record remediation, replay the exact workload, verify recovery, preserve a human deployment decision and then scale the workload.

  1. 1
    Observe

    Establish the unsafe synthetic route and separate agent activity from lab-side observation.

  2. 2
    Contain

    Apply velocity counting, a circuit breaker and exact-action policy fallback.

  3. 3
    Fix

    Record the confirmed condition, remediation owner and exact retest.

  4. 4
    Retest

    Replay the frozen workload digest and verify the changed outcome.

  5. 5
    Recover

    Stop, revoke, isolate, restore and probe the recovered state.

  6. 6
    Decide

    An operator separately recorded hold; AgentRiskLayer did not make the deployment decision.

  7. 7
    Scale

    Repeat at 1,000 attempts before the final 17,600-attempt workload.

  8. 8
    Bind

    Hash the final workload, evidence streams and result into an evidence manifest.

Integrity details

Exact artifacts, exact hashes.

The canonical workload digest hashes compact JSON. The saved workload artifact is pretty-printed, so its file hash is intentionally different.

Canonical workload

f0323357…f90f7847

SHA-256 of the frozen 17,600-attempt workload in canonical compact form.

Saved workload artifact

9ac237db…ddcc929

SHA-256 of the pretty-printed workload manifest file.

Unsafe evidence

50de5151…e6bf3f9

SHA-256 of the 17,607-line baseline observer evidence stream.

Protected evidence

6ac7040e…693d19

SHA-256 of the 56-line protected observer evidence stream.

Benchmark result

30147656…a804c3

SHA-256 of the final benchmark result artifact.

Evidence manifest

ab480c8e…a5dc2

SHA-256 of the final artifact manifest that records the evidence boundary.

What we do not claim

Evidence before claims.

The benchmark is deliberately bounded so a buyer can see what the result means—and where it stops.

  • This is a safe synthetic benchmark inspired by publicly disclosed characteristics of a July 2026 autonomous-agent security incident. It does not reproduce that incident.
  • It does not establish that AgentRiskLayer would have prevented the real-world incident.
  • It does not prove production integration of the stateful circuit breaker or production recovery.
  • The lab-side observer runs in the same process. It is not independent monitoring, third-party assurance or an accredited certification.
  • No real network target, credential, customer data, shell execution or production side effect was used.
  • The Phase 8 PASS is benchmark evidence, not a deployment decision. The earlier operator decision remains hold.
Now test the agent that actually matters

Could your agent keep trying until an ordinary weakness becomes a real action?

Start with one agent. Map its access, data, tools, approval and recovery state. Unknowns remain evidence gaps; findings require support.

Assess one agent free