CAN ASTRA DO X?

The questions Astra still needs to answer.

A practical test library, with success criteria in plain sight.

Browse the questions we plan to test across coding, research, analysis, and long-running work. Results will appear only after a documented run and review.

Last verified September 20, 2026

TEST STATUS

Eight test questions are defined. Results are not available yet.

These outlines show what a useful answer would need to demonstrate. They are not evidence of Astra’s current performance.

See the benchmark plan

TEST QUESTIONS

The first questions worth answering.

Open a question to see the task, what counts as success, and what would make the result inconclusive.

Not run yet Can Astra build a small app from a written brief? Planned test · no result yet
What we will test
A fixed task brief will be published with the first test.
Success criteria
  • Build runs
  • Core user journey works
  • Tests and limitations are documented
Stop conditions
  • No progress after the time budget
  • Unsafe external action requested
  • Required access unavailable
Last tested
No result yet. The first test needs an eligible account and a controlled environment.
Not run yet Can Astra analyze a 10-K with traceable evidence? Planned test · no result yet
What we will test
A fixed document set, extraction rule, and citation format will be published with the first test.
Success criteria
  • Material claims cite page evidence
  • Arithmetic reproduces
  • Unknowns are disclosed
Stop conditions
  • Document cannot be accessed
  • Evidence cannot be traced
  • Output exceeds review budget
Last tested
No result yet. The first test needs an eligible account and a controlled environment.
Not run yet Can Astra debug a repository without scope drift? Planned test · no result yet
What we will test
A fixed repository, issue, and tool boundary will be published with the first test.
Success criteria
  • Root cause is demonstrated
  • Patch is minimal
  • Regression test passes
Stop conditions
  • Patch would alter unrelated behavior
  • Credentials or network access are required
  • Evidence is inconclusive
Last tested
No result yet. The first test needs an eligible account and a controlled environment.
Not run yet Can Astra solve a graduate-level mathematics task and verify it? Planned test · no result yet
What we will test
A fixed problem, allowed references, and independent review method will be published with the first test.
Success criteria
  • Solution is correct
  • Assumptions are explicit
  • Independent checker agrees
Stop conditions
  • No proof or counterexample is established
  • Verification is unavailable
  • Time budget is exhausted
Last tested
No result yet. The first test needs an eligible account and a controlled environment.
Not run yet Can Astra review scientific literature without inventing evidence? Planned test · no result yet
What we will test
A fixed paper set and citation review method will be published with the first test.
Success criteria
  • Claims map to papers
  • Contradictions are visible
  • Missing evidence is labeled
Stop conditions
  • Source access fails
  • Citation identity cannot be verified
  • Scope expands beyond the paper set
Last tested
No result yet. The first test needs an eligible account and a controlled environment.
Not run yet Can Astra analyze a large dataset with reproducible outputs? Planned test · no result yet
What we will test
A fixed dataset, task, and output checks will be published with the first test.
Success criteria
  • Output reproduces
  • Memory/time behavior is recorded
  • Data handling is safe
Stop conditions
  • Data cannot be processed safely
  • Artifact cannot be reproduced
  • Tool budget is exhausted
Last tested
No result yet. The first test needs an eligible account and a controlled environment.
Not run yet Can Astra use a browser reliably while respecting boundaries? Planned test · no result yet
What we will test
A sandboxed site, allowed actions, and test data will be published with the first test.
Success criteria
  • Actions stay within the allowlist
  • State is verified after each action
  • No sensitive data is requested
Stop conditions
  • Site state is ambiguous
  • Action would be irreversible
  • Browser access is not officially supported
Last tested
No result yet. The first test needs an eligible account and a controlled environment.
Not run yet Can Astra generate or repair a Lean proof with a clean build? Planned test · no result yet
What we will test
A fixed Lean environment and proof task will be published with the first test.
Success criteria
  • Lean build passes
  • No hidden axioms are added
  • Proof is reviewed
Stop conditions
  • Environment cannot be reproduced
  • Proof relies on unapproved axioms
  • Time budget is exhausted
Last tested
No result yet. The first test needs an eligible account and a controlled environment.

WHEN RESULTS ARRIVE

A result should include the task, evidence, cost, comparison, and review status so readers can judge it for themselves.

Task

The question, inputs, tools, and success criteria are stated clearly.

Evidence

The output, checks, failures, and any relevant comparison are available to inspect.

Context

The date, limitations, and review status make the result easier to interpret.