BENCHMARKS

The score is only the beginning of the record.

OpenAI reports striking cyber results. Independent Astra scores are not yet published.

This page separates vendor-reported evaluation claims from the independent tests we will run when a safe, reproducible access path exists. A benchmark number without task conditions, version, artifacts, and failures is not a useful comparison.

Last verified September 7, 2026 Review recommended

CURRENT RECORD

100% on ExploitBench · reported by OpenAI.

OpenAI also reports two zero-days during evaluation. The results reflect Daybreak Blue access and OpenAI’s evaluation conditions, not a public production leaderboard.

Confirmed

VENDOR-REPORTED SIGNALS

Keep each result in its lane.

“Confirmed” here means OpenAI published the claim. It does not mean this publication has independently recreated it or that the result covers every deployment.

ExploitBench

Confirmed

100% reported by OpenAI

Official vendor result; public conditions and independent replication are pending.

OpenAI high-severity evaluation

Confirmed

Two zero-days reported

Evidence of capability in OpenAI’s evaluation; not a rate or universal success claim.

Mathematics and theoretical computer science

Confirmed

Ten published advances

A research record with manuscripts and Lean certificates; not a general product benchmark.

INDEPENDENT TESTS

Eight task families are waiting for a fair test.

The first battery covers coding, research, long-context recovery, mathematics, data, planning, tool boundaries, and calibration. Each task will state its model, conditions, cost, and evaluator when a safe, eligible test account is available.

Independent task plan — Astra results not yet published
CategoryTaskSuccess criteriaState
coding Repository-level change
Implement a bounded feature in an unfamiliar repository with tests and a rollback note.
  • Tests pass
  • Diff matches the requested scope
  • No unrelated files changed
Not run yet
deep-research Source-grounded synthesis
Answer a contested technical question using primary sources and clearly separated inference.
  • Every material claim has a source
  • Unknowns remain visible
  • Contradictions are handled
Not run yet
long-context Long-context recovery
Resume a task after a checkpoint and recover the correct objective, constraints, and open questions.
  • State is reconstructed
  • No hidden assumptions are introduced
  • Checkpoint is auditable
Not run yet
math Mathematical argument checking
Produce and critique a bounded proof or counterexample with a separate verification pass.
  • Argument is valid or failure is identified
  • Assumptions are explicit
  • Independent review agrees
Not run yet
data Data-analysis artifact
Analyze a fixed dataset and produce a reproducible table or chart without leaking input data.
  • Calculations reproduce
  • Missingness is disclosed
  • Artifact is reviewable
Not run yet
agent Long-horizon agent task
Complete a bounded multi-step project with checkpoints, tools, and explicit stop conditions.
  • All required milestones are completed
  • State survives each checkpoint
  • Stop conditions are respected
Not run yet
computer-use Computer-use completion
Complete a fixed browser workflow while staying inside an allowlisted local application.
  • Target state is reached
  • No disallowed navigation or data entry occurs
  • Final state is evidenced
Not run yet
adversarial Uncertainty and calibration
Answer a mixed set of known, unknown, and adversarial questions with calibrated confidence.
  • Unknowns are admitted
  • Confidence tracks correctness
  • Citations are not fabricated
Not run yet
formal-proof Lean/formal proof
Generate or repair a formal proof in a pinned Lean environment.
  • Lean build passes
  • No hidden axioms are added
  • Proof is independently reviewed
Not run yet
efficiency Cost, latency, and token efficiency
Measure useful output per token, dollar, and second without trading away the quality floor.
  • Raw usage is complete
  • Quality floor is met
  • Efficiency ratios reproduce from the record
Not run yet

WHAT A FAIR RESULT SHOWS

What a published comparison must include.

A result becomes useful when another reader can understand the run and its failure boundary.

01

Configuration

Model identifier, provider, date, system instructions, tools, limits, and environment.

02

Artifacts

Prompt, raw output, tool log, generated files, tests, evaluator rubric, and dataset hash where relevant.

03

Economics

Cost, tokens, latency, retries, time budget, and any human intervention or monitoring pause.

04

Boundary

Sample size, selection effects, failures, safety limits, and what the result does not mean.

SOURCE LEDGER

Official claims, independent method.

OpenAI’s launch record supplies the vendor-reported cyber results. The independent test plan is outlined in the Benchmark Lab; we will publish a result only when a safe access path and the required evidence exist.

FAQ

Astra benchmark questions

What is Astra’s benchmark score?

OpenAI reports 100% on ExploitBench in its launch evaluation record. The result is vendor-reported and tied to the evaluation context described by OpenAI; it is not an independent score or a claim that Astra succeeds on every cyber task.

Did Astra find zero-days?

OpenAI reports that Astra discovered two zero-days during its evaluation. The public headline does not by itself provide enough detail to reproduce the run or establish how the result transfers to other environments.

When will you publish independent Astra benchmarks?

When a public access path and safe execution environment exist. The Benchmark Lab has documented the task families and what a fair result should show, but we will not publish a score before a run.

DEMAND PROBE · EVALUATION

Get the Enterprise Astra Evaluation Kit

We are testing demand for a procurement-ready evaluation kit built from the independent benchmark protocol. Join the waitlist; we will only publish it if enough evaluators ask for it.

We never sell your personal information, and we will not add you to anything you did not ask for. Every confirmation includes a one-click unsubscribe link. Read the privacy policy.