Agent outcome benchmark · v0.5
Did the agent succeed, or only say it did?
Browser, coding, data, and messaging agents run bounded tasks. A held-out evaluator checks persisted outcomes, policy compliance, and evidence after execution.
The archive includes retained results, public keys, and an offline verifier. It does not run a fresh study. Source and SHA-256 manifest
Historical regression
Every release keeps the prior result visible.
Results are append-only research records. New runs do not overwrite a previous benchmark version.
No blocking regression. Absolute Assurance Gap increased by 3.3%, crossing the review threshold.
Real harnesses
Named, pinned open-source agent frameworks execute deterministic synthetic systems.
Held-out outcomes
The agent cannot read the private predicate or evaluator signing key.
Signed artifacts
Task results, provenance, costs, manifests, and evaluator receipts remain independently verifiable.
Explicit limits
This measures declared fixtures. It is not a general safety certificate or production guarantee.