Skip to contentWitnora
Menu

Agent outcome benchmark · v0.5

Did the agent succeed, or only say it did?

Browser, coding, data, and messaging agents run bounded tasks. A held-out evaluator checks persisted outcomes, policy compliance, and evidence after execution.

The archive includes retained results, public keys, and an offline verifier. It does not run a fresh study. Source and SHA-256 manifest

Apparent success91.9%480 total task runs
Verified success87.5%Independent observed state
Assurance Gap+4.4%15.6% directional disagreement
Run variance4.7% SD5 repetitions per cell

Historical regression

Every release keeps the prior result visible.

Results are append-only research records. New runs do not overwrite a previous benchmark version.

Automated regression gatewarning

No blocking regression. Absolute Assurance Gap increased by 3.3%, crossing the review threshold.

Apparent successVerified success
v0.4
87.5%88.5%
v0.5
91.9%87.5%
VersionRunsRepsApparentVerifiedGapControl findings
v0.52026-08-03480591.9%87.5%+4.4%0/8repetition sign-flip, Holm-adjusted
v0.42026-08-0296187.5%88.5%-1.0%0/8paired bootstrap, unadjusted
01

Real harnesses

Named, pinned open-source agent frameworks execute deterministic synthetic systems.

02

Held-out outcomes

The agent cannot read the private predicate or evaluator signing key.

03

Signed artifacts

Task results, provenance, costs, manifests, and evaluator receipts remain independently verifiable.

04

Explicit limits

This measures declared fixtures. It is not a general safety certificate or production guarantee.