The Endless Exam Won't Measure Superintelligence Yet

The Endless Exam Won't Measure Superintelligence Yet

The Endless Exam introduces fourteen parameterised construction families designed to score mathematical progress without capping improvements at 1.0. This analysis argues the methodology is a real advance over static benchmarks, but the absence of published model results and the moving-frontier problem leave its central claim untested.

A new arXiv paper proposes grading AI math progress with an uncapped quality score against a moving frontier rather than a pass/fail bar. The idea is sound, the execution is unproven, and no model results have been published. That gap is the story.
  • What happened: An arXiv paper (2609.24555v1, published 2026-09-21) introduces the Endless Exam, a benchmark with fourteen parameterised construction families that auto-validate submitted mathematical objects and score them against a published frontier without capping at 1.0.
  • Why it matters: It reframes math evaluation from binary solving to open-ended construction quality β€” the first serious attempt to build a benchmark that scales toward superintelligence rather than saturating.
  • Key tension: The paper describes a measurement instrument but publishes no model results, so whether it discriminates between today's frontier models is entirely unverified.

What Actually Changed in How We Measure AI Math?

The Endless Exam, per the arXiv abstract, replaces the pass/fail paradigm with a relative quality score: each submitted object is checked automatically for validity and then scored against a published frontier or construction baseline, explicitly without capping improvements at 1.0. That last clause is the substantive break. Existing math benchmarks β€” from GSM8K through FrontierMath β€” saturate: once a model solves the problems, the benchmark stops discriminating. An uncapped score means a model that constructs a marginally better object than the current frontier gets a number above 1.0, preserving headroom indefinitely. The second change is parameterisation. The families generate new instances at larger parameters, which means the benchmark can be extended without authoring new problems. According to the arXiv listing for the paper, the families draw on open mathematical problems for long-term targets β€” a design that ties short-term scoring to unsolved research mathematics rather than competition-style exercises. My read: this is a benchmark architecture paper, not a results paper. The methodology is the contribution, and it is a real one. But architecture papers live or die on whether anyone runs them.

Do Fourteen Construction Families Actually Cover Mathematical Progress?

Fourteen is a suspicious number β€” large enough to look comprehensive, small enough to be hand-curated. The abstract does not name the families, which makes independent assessment impossible from the source material alone. According to the arXiv abstract, the families are parameterised, meaning each family has a knob that increases difficulty or instance size.
The Endless Exam Wont Measure Superintelligence Yet
The coverage question is the one that decides whether the Endless Exam becomes infrastructure or becomes a paper. Mathematical progress toward superintelligence, if that phrase means anything operational, would require coverage of algebra, topology, combinatorics, analysis, and number theory at minimum β€” and ideally constructions that connect to open problems in each. Fourteen families could plausibly span that. Or it could over-index on whatever the authors found tractable to auto-validate, which is the binding constraint: automatic validity checking is much easier for combinatorial constructions than for, say, a novel proof in algebraic geometry. I would want to see the family list before treating coverage as settled. The abstract does not provide it.

What Does an Uncapped Score Actually Reward?

An uncapped relative score is a good idea with a known failure mode: it rewards optimization against the scorer, not mathematical insight. If the quality metric is a published function, models can be trained or prompted to maximize it directly. The paper's defense, per the abstract, is automatic validity checking β€” an invalid object scores nothing, so gaming requires producing genuinely valid constructions. That defense holds only if validity checking is both sound and complete for the relevant object classes. For many construction families, validity is decidable; for others, it is not. The abstract does not specify how the authors handle undecidable validity cases, which is the technical hole I would probe first. The second-order effect: uncapped scores make leaderboard position a function of compute spent searching construction space. Labs with large inference budgets can brute-force improvements that a smaller lab's better reasoning cannot match. That is a real distortion, and it is not addressed in the source material.

Who Wins and Who Loses If This Benchmark Sticks?

ApproachScoring ModelHeadroomGaming RiskAdoption Signal
Endless ExamUncapped relative quality vs. frontierUnlimited by designHigh if validity checking is incompleteNone yet β€” no published model results
FrontierMath (Epoch AI)Pass/fail on expert-authored problemsSaturates as models improveLow β€” problems are held outCited by multiple labs
IMO-style benchmarksPartial credit on competition problemsSaturates at gold-medal levelModerate β€” training contaminationWidely reported
Human expert baselinesQualitative peer reviewUnbounded but slowLowStandard in mathematics
VerdictThe Endless Exam has the best theoretical headroom of any current math benchmark, but until a frontier lab publishes a score, FrontierMath remains the practical reference point.

Why Publish a Benchmark With No Model Results?

The most likely explanation is that the paper is a methods contribution and results are forthcoming. The less charitable explanation is that the benchmark is hard to run β€” generating valid constructions at large parameters may require substantial compute, and scoring requires a published frontier that the authors may not yet have established. According to the arXiv abstract, the benchmark assigns scores against a published frontier or construction baseline. That phrasing implies the frontier must exist and be public for the benchmark to function. If the frontier is thin β€” a handful of constructions β€” then early models will trivially exceed 1.0 and the benchmark will look easy. If the frontier is strong, the benchmark may be too hard to produce signal. This is the same calibration problem every benchmark faces, and the Endless Exam's uncapped design makes it sharper: there is no natural ceiling to hide behind.
The Endless Exam is the right idea executed at the wrong level of completeness. My thesis: uncapped, auto-validated construction scoring is a genuine methodological advance over pass/fail math benchmarks, but without published model results and a transparent family list, it is currently an untested instrument rather than evidence about AI mathematical progress. Short term, the paper's impact depends entirely on whether a major lab β€” OpenAI, Google DeepMind, or Anthropic β€” runs it and publishes numbers. DeepMind has the strongest incentive: its AlphaProof lineage already targets construction-style mathematics, and an uncapped benchmark would let it show improvement beyond binary solve rates. If DeepMind publishes an Endless Exam score in the next two quarters, the benchmark becomes a reference point. If nobody does, it joins the long list of arXiv benchmark proposals that never got adopted. Long term, the moving-frontier design is the only credible path to a benchmark that survives model improvement. The failure mode is not saturation but capture: if the frontier is maintained by a small group, that group effectively defines what mathematical progress means. That is a governance problem, not a technical one, and the paper does not address it. Who gains: labs with verification tooling and inference budget. Who loses: labs that compete on sample efficiency, because uncapped scoring rewards search volume. The concrete prediction I will stand behind: Google DeepMind will publish an Endless Exam result before any other frontier lab, because its existing proof-search infrastructure maps most directly onto the benchmark's construction format.

Predictions

1. Google DeepMind will publish an Endless Exam score for a Gemini-class model by Q2 2027, citing its AlphaProof-derived construction pipeline as the enabling tool. 2. Within twelve months of any first published result, at least one lab will be accused of frontier-gaming β€” submitting constructions optimized against the published scoring function rather than for mathematical interest β€” and the authors will be forced to revise the validity checker. 3. If no frontier lab publishes an Endless Exam result by Q3 2027, the benchmark will not be cited in any major model card, and Epoch AI's FrontierMath will remain the default math reference through 2027.
  1. September 2026
    Endless Exam paper published

    arXiv paper 2609.24555v1 introduces the benchmark with fourteen parameterised construction families and uncapped relative scoring.

  2. Q2 2027
    Predicted first frontier-lab result

    Google DeepMind is predicted to publish the first Endless Exam score, leveraging AlphaProof-derived construction tooling.

  3. Q3 2027
    Adoption decision point

    If no frontier lab has published a result by this point, the benchmark is unlikely to enter standard model-card reporting.

Benchmark Headroom Comparison (estimated)

Article Summary

  • The Endless Exam's uncapped scoring is the first serious benchmark design that does not assume saturation, which matters more than any individual family it includes.
  • Automatic validity checking is the load-bearing component β€” if it is incomplete, the whole uncapped-score premise collapses into scorer optimization.
  • Fourteen unnamed families is a coverage claim, not coverage evidence; the family list should be the first thing a skeptical reader requests.
  • No published model results means the benchmark's discriminating power is entirely unverified as of the paper's 2026-09-21 publication.
  • The moving-frontier design creates a governance question the paper does not answer: whoever maintains the frontier defines mathematical progress.

Source and attribution

arXiv
The Endless Exam: Mathematical Constructions from Today's Models toward Superintelligence

Discussion

Add a comment

0/5000
Loading comments...