The Endless Exam Won't Measure Superintelligence Yet
The Endless Exam introduces fourteen parameterised construction families designed to score mathematical progress without capping improvements at 1.0. This analysis argues the methodology is a real advance over static benchmarks, but the absence of published model results and the moving-frontier problem leave its central claim untested.
- What happened: An arXiv paper (2609.24555v1, published 2026-09-21) introduces the Endless Exam, a benchmark with fourteen parameterised construction families that auto-validate submitted mathematical objects and score them against a published frontier without capping at 1.0.
- Why it matters: It reframes math evaluation from binary solving to open-ended construction quality β the first serious attempt to build a benchmark that scales toward superintelligence rather than saturating.
- Key tension: The paper describes a measurement instrument but publishes no model results, so whether it discriminates between today's frontier models is entirely unverified.
What Actually Changed in How We Measure AI Math?
The Endless Exam, per the arXiv abstract, replaces the pass/fail paradigm with a relative quality score: each submitted object is checked automatically for validity and then scored against a published frontier or construction baseline, explicitly without capping improvements at 1.0. That last clause is the substantive break. Existing math benchmarks β from GSM8K through FrontierMath β saturate: once a model solves the problems, the benchmark stops discriminating. An uncapped score means a model that constructs a marginally better object than the current frontier gets a number above 1.0, preserving headroom indefinitely. The second change is parameterisation. The families generate new instances at larger parameters, which means the benchmark can be extended without authoring new problems. According to the arXiv listing for the paper, the families draw on open mathematical problems for long-term targets β a design that ties short-term scoring to unsolved research mathematics rather than competition-style exercises. My read: this is a benchmark architecture paper, not a results paper. The methodology is the contribution, and it is a real one. But architecture papers live or die on whether anyone runs them.Do Fourteen Construction Families Actually Cover Mathematical Progress?
Fourteen is a suspicious number β large enough to look comprehensive, small enough to be hand-curated. The abstract does not name the families, which makes independent assessment impossible from the source material alone. According to the arXiv abstract, the families are parameterised, meaning each family has a knob that increases difficulty or instance size.
What Does an Uncapped Score Actually Reward?
An uncapped relative score is a good idea with a known failure mode: it rewards optimization against the scorer, not mathematical insight. If the quality metric is a published function, models can be trained or prompted to maximize it directly. The paper's defense, per the abstract, is automatic validity checking β an invalid object scores nothing, so gaming requires producing genuinely valid constructions. That defense holds only if validity checking is both sound and complete for the relevant object classes. For many construction families, validity is decidable; for others, it is not. The abstract does not specify how the authors handle undecidable validity cases, which is the technical hole I would probe first. The second-order effect: uncapped scores make leaderboard position a function of compute spent searching construction space. Labs with large inference budgets can brute-force improvements that a smaller lab's better reasoning cannot match. That is a real distortion, and it is not addressed in the source material.Who Wins and Who Loses If This Benchmark Sticks?
| Approach | Scoring Model | Headroom | Gaming Risk | Adoption Signal |
|---|---|---|---|---|
| Endless Exam | Uncapped relative quality vs. frontier | Unlimited by design | High if validity checking is incomplete | None yet β no published model results |
| FrontierMath (Epoch AI) | Pass/fail on expert-authored problems | Saturates as models improve | Low β problems are held out | Cited by multiple labs |
| IMO-style benchmarks | Partial credit on competition problems | Saturates at gold-medal level | Moderate β training contamination | Widely reported |
| Human expert baselines | Qualitative peer review | Unbounded but slow | Low | Standard in mathematics |
| Verdict | The Endless Exam has the best theoretical headroom of any current math benchmark, but until a frontier lab publishes a score, FrontierMath remains the practical reference point. | |||
Why Publish a Benchmark With No Model Results?
The most likely explanation is that the paper is a methods contribution and results are forthcoming. The less charitable explanation is that the benchmark is hard to run β generating valid constructions at large parameters may require substantial compute, and scoring requires a published frontier that the authors may not yet have established. According to the arXiv abstract, the benchmark assigns scores against a published frontier or construction baseline. That phrasing implies the frontier must exist and be public for the benchmark to function. If the frontier is thin β a handful of constructions β then early models will trivially exceed 1.0 and the benchmark will look easy. If the frontier is strong, the benchmark may be too hard to produce signal. This is the same calibration problem every benchmark faces, and the Endless Exam's uncapped design makes it sharper: there is no natural ceiling to hide behind.Predictions
1. Google DeepMind will publish an Endless Exam score for a Gemini-class model by Q2 2027, citing its AlphaProof-derived construction pipeline as the enabling tool. 2. Within twelve months of any first published result, at least one lab will be accused of frontier-gaming β submitting constructions optimized against the published scoring function rather than for mathematical interest β and the authors will be forced to revise the validity checker. 3. If no frontier lab publishes an Endless Exam result by Q3 2027, the benchmark will not be cited in any major model card, and Epoch AI's FrontierMath will remain the default math reference through 2027.- September 2026Endless Exam paper published
arXiv paper 2609.24555v1 introduces the benchmark with fourteen parameterised construction families and uncapped relative scoring.
- Q2 2027Predicted first frontier-lab result
Google DeepMind is predicted to publish the first Endless Exam score, leveraging AlphaProof-derived construction tooling.
- Q3 2027Adoption decision point
If no frontier lab has published a result by this point, the benchmark is unlikely to enter standard model-card reporting.
Benchmark Headroom Comparison (estimated)
Article Summary
- The Endless Exam's uncapped scoring is the first serious benchmark design that does not assume saturation, which matters more than any individual family it includes.
- Automatic validity checking is the load-bearing component β if it is incomplete, the whole uncapped-score premise collapses into scorer optimization.
- Fourteen unnamed families is a coverage claim, not coverage evidence; the family list should be the first thing a skeptical reader requests.
- No published model results means the benchmark's discriminating power is entirely unverified as of the paper's 2026-09-21 publication.
- The moving-frontier design creates a governance question the paper does not answer: whoever maintains the frontier defines mathematical progress.
Source and attribution
arXiv
The Endless Exam: Mathematical Constructions from Today's Models toward Superintelligence
Discussion
Add a comment