BENCHMARKS
The score is only the beginning of the record.
OpenAI reports striking cyber results. Independent Astra scores are not yet published.
This page separates vendor-reported evaluation claims from the independent tests we will run when a safe, reproducible access path exists. A benchmark number without task conditions, version, artifacts, and failures is not a useful comparison.
CURRENT RECORD
100% on ExploitBench · reported by OpenAI.
OpenAI also reports two zero-days during evaluation. The results reflect Daybreak Blue access and OpenAI’s evaluation conditions, not a public production leaderboard.
VENDOR-REPORTED SIGNALS
Keep each result in its lane.
“Confirmed” here means OpenAI published the claim. It does not mean this publication has independently recreated it or that the result covers every deployment.
ExploitBench
Confirmed100% reported by OpenAI
Official vendor result; public conditions and independent replication are pending.
OpenAI high-severity evaluation
ConfirmedTwo zero-days reported
Evidence of capability in OpenAI’s evaluation; not a rate or universal success claim.
Mathematics and theoretical computer science
ConfirmedTen published advances
A research record with manuscripts and Lean certificates; not a general product benchmark.
INDEPENDENT TESTS
Eight task families are waiting for a fair test.
The first battery covers coding, research, long-context recovery, mathematics, data, planning, tool boundaries, and calibration. Each task will state its model, conditions, cost, and evaluator when a safe, eligible test account is available.
| Category | Task | Success criteria | State |
|---|---|---|---|
| coding | Repository-level change Implement a bounded feature in an unfamiliar repository with tests and a rollback note. |
| Not run yet |
| deep-research | Source-grounded synthesis Answer a contested technical question using primary sources and clearly separated inference. |
| Not run yet |
| long-context | Long-context recovery Resume a task after a checkpoint and recover the correct objective, constraints, and open questions. |
| Not run yet |
| math | Mathematical argument checking Produce and critique a bounded proof or counterexample with a separate verification pass. |
| Not run yet |
| data | Data-analysis artifact Analyze a fixed dataset and produce a reproducible table or chart without leaking input data. |
| Not run yet |
| agent | Long-horizon agent task Complete a bounded multi-step project with checkpoints, tools, and explicit stop conditions. |
| Not run yet |
| computer-use | Computer-use completion Complete a fixed browser workflow while staying inside an allowlisted local application. |
| Not run yet |
| adversarial | Uncertainty and calibration Answer a mixed set of known, unknown, and adversarial questions with calibrated confidence. |
| Not run yet |
| formal-proof | Lean/formal proof Generate or repair a formal proof in a pinned Lean environment. |
| Not run yet |
| efficiency | Cost, latency, and token efficiency Measure useful output per token, dollar, and second without trading away the quality floor. |
| Not run yet |
WHAT A FAIR RESULT SHOWS
What a published comparison must include.
A result becomes useful when another reader can understand the run and its failure boundary.
Configuration
Model identifier, provider, date, system instructions, tools, limits, and environment.
Artifacts
Prompt, raw output, tool log, generated files, tests, evaluator rubric, and dataset hash where relevant.
Economics
Cost, tokens, latency, retries, time budget, and any human intervention or monitoring pause.
Boundary
Sample size, selection effects, failures, safety limits, and what the result does not mean.
SOURCE LEDGER
Official claims, independent method.
OpenAI’s launch record supplies the vendor-reported cyber results. The independent test plan is outlined in the Benchmark Lab; we will publish a result only when a safe access path and the required evidence exist.
FAQ
Astra benchmark questions
What is Astra’s benchmark score?
OpenAI reports 100% on ExploitBench in its launch evaluation record. The result is vendor-reported and tied to the evaluation context described by OpenAI; it is not an independent score or a claim that Astra succeeds on every cyber task.
Did Astra find zero-days?
OpenAI reports that Astra discovered two zero-days during its evaluation. The public headline does not by itself provide enough detail to reproduce the run or establish how the result transfers to other environments.
When will you publish independent Astra benchmarks?
When a public access path and safe execution environment exist. The Benchmark Lab has documented the task families and what a fair result should show, but we will not publish a score before a run.
DEMAND PROBE · EVALUATION
Get the Enterprise Astra Evaluation Kit
We are testing demand for a procurement-ready evaluation kit built from the independent benchmark protocol. Join the waitlist; we will only publish it if enough evaluators ask for it.