Inherent exposure
Measures data sensitivity, autonomy, high-impact actions, public reach, external content, connected tools and delegation.
AgentRiskLayer separates declared exposure, control effectiveness and evidence confidence, then adds integrity-verified technical observations from an optional local scanner and controlled adversarial outcomes from a customer-operated staging runner.
Measures data sensitivity, autonomy, high-impact actions, public reach, external content, connected tools and delegation.
Measures identity, authorisation, approvals, input boundaries, validation, memory, egress, monitoring, testing and containment.
Distinguishes no evidence, owner claims, documented controls and tested evidence. Unsupported claims increase uncertainty.
Combines multiple conditions into scenarios such as indirect prompt injection leading to privileged tool execution.
Checks tracked files and deployment configuration for secrets, supply-chain risk, MCP/tool scope, container controls, CI/CD permissions and agent-control signals.
Exercises a 32-case catalogue across direct, indirect, encoded and multilingual prompt injection, synthetic exfiltration, dry-run tool misuse, SSRF and injection-shaped tool arguments, memory poisoning, tenant isolation, schema enforcement, authority confusion and resource limits. Cases can repeat 1-5 times to measure stability.
Inspection and red-team bundles use SHA-256 and Ed25519, one-time upload tokens and replay protection. Source code, raw prompts, target responses, credentials and matched secret values are excluded.
The control library is informed by OWASP guidance for agentic applications and AI agents, the OWASP LLM Top 10, NIST AI RMF 1.0 and the NIST Generative AI Profile.
Every staging campaign must be bound to a named target, legal authority basis, named authoriser, emergency stop contact, approved time window, permitted and prohibited actions, synthetic-data attestation, dry-run-tools attestation and explicit confirmation. The server rejects missing, expired, revoked, environment-mismatched or endpoint-mismatched authorisations.
A structured, explainable self-assessment and remediation planning tool that helps teams identify control gaps and prepare for deeper security work.
The inspector is static and the red-team runner is controlled and customer-operated. Neither is an independent penetration test, runtime attestation, audit, certification, legal advice or guarantee that a system is secure. Production and destructive testing are intentionally excluded.
The signature proves that an accepted inspection or red-team evidence bundle was not changed after local generation. Repeated trials improve stability evidence but do not prove security outside the tested cases and adapter behaviour. It does not prove that the customer scanned every relevant repository, that production matches the scanned material, or that the system was uncompromised. Reports state this boundary prominently.
Residual risk combines inherent exposure (38%), control gaps (62%) and an uncertainty penalty where answers lack documented or tested evidence. Scores prioritise action; they are not probabilities of breach.