GPT-5.6 Solves 30-Year Math Gap: AI Now a Mathematician
GPT-5.6 has closed a 30-year gap in convex optimization from a single prompt, verified by mathematicians. This is a step change in AI's ability to contribute original research.
- What happened: GPT-5.6, prompted with a description of a known gap in convex optimization theory, produced a correct proof that closed it. Experts confirmed the result within 48 hours.
- Why it matters: This is the first documented case of a frontier LLM solving a longstanding open problem in pure mathematics without iterative human guidance or scaffolded reasoning.
- Key tension: The proof's verification relied on human mathematicians, raising questions about whether AI-generated proofs can be trusted without equivalent machine-verifiable steps.
How Did a Single Prompt Lead to a 30-Year Breakthrough?
According to the original Reddit post on r/math by user 'ProofBot2026', the prompt given to GPT-5.6 was: "Prove or disprove the existence of a self-concordant barrier function for the class of convex sets defined by polynomial inequalities of degree at most 4." This problem had resisted solution since 1996, when Yurii Nesterov and Arkadi Nemirovski first characterized self-concordant barriers. The model output a 12-page proof that, after initial skepticism, was checked by three mathematicians at separate institutions. The Hacker News thread aggregating the discussion noted that the proof relied on a novel decomposition of the polynomial constraint set — an approach no human had published.
Does This Prove LLMs Have Genuine Mathematical Reasoning?

The evidence is strong but not conclusive. The proof's structure — a chain of lemmas, each building on the last — mirrors human mathematical exposition. However, as mathematician Dr. Elena Voss noted in the HN thread, "The proof is correct but its presentation is somewhat opaque; it skips steps a human would include." This suggests GPT-5.6 may be operating on a different cognitive substrate — one that compresses reasoning. Yet the fact that the proof was verifiable and novel implies more than shallow pattern completion. The model did not retrieve a known solution; it generated one that experts had missed for decades.
Who Benefits Most From This Breakthrough?
The immediate beneficiaries are researchers in convex optimization, a field foundational to machine learning, control theory, and operations research. The new barrier function could speed up interior-point methods for polynomial optimization by an estimated 30-50% in certain regimes, according to a comment by user 'OptTheoryProf' on the HN thread. OpenAI stands to benefit disproportionately: GPT-5.6's success validates its investment in reasoning training regimes, potentially justifying higher API pricing and attracting enterprise customers in R&D-intensive industries. Conversely, Google DeepMind's AlphaProof and AlphaGeometry, which required extensive human-curated problem sets, now appear less autonomous.
| Dimension | GPT-5.6 (OpenAI) | AlphaProof (DeepMind) |
|---|---|---|
| Problem type | Open research problem (30-year gap) | IMO problems (known solutions) |
| Human guidance | Single prompt, no iteration | Curated training set + human feedback |
| Proof verifiability | Human-verified, no formal checker | Formal verification built-in |
| Novelty | Truly novel approach | Novel but within known solution space |
| Time to result | Minutes | Hours (with scaffolding) |
| Verdict | Winner: higher autonomy and novelty | Winner: formal verification trust |
What Remains Uncertain About This Result?
Several critical questions persist. First, reproducibility: can GPT-5.6 replicate this feat on other open problems, or was this a statistical fluke? The HN thread reported that subsequent prompts to the same model on different problems yielded plausible but incorrect proofs. Second, verification: the proof has not been formally verified in a system like Lean or Coq. As one commenter noted, "Human verification is fallible; a 12-page proof could hide a subtle error." Third, the training data: did GPT-5.6 inadvertently memorize fragments of unpublished preprints or forum discussions that contained partial solutions? OpenAI has not disclosed the training corpus for this model.
What Does This Mean for the Future of Mathematical Research?
According to a comment from user 'MathProf2026' on the HN discussion, "This is the Sputnik moment for AI in mathematics." The implication is clear: within 3-5 years, AI systems could routinely propose conjectures, generate proofs, and even referee papers. This threatens the traditional gatekeeping role of human mathematicians, especially in subfields like optimization and combinatorics where problems are well-formalized. However, it also opens up new collaborative workflows: AI as a tireless assistant that can explore proof directions humans would dismiss. The net effect may be a dramatic acceleration of mathematical discovery, but with significant displacement of junior researchers whose work involves solving known-class problems.
My thesis: GPT-5.6's proof is a genuine breakthrough, but the field must now confront the hard problem of trust — can we accept AI-generated proofs without formal verification?
In the short term, this will drive a surge of interest in automated proof assistants like Lean and Coq, as mathematicians seek to verify AI outputs. In the long term, it will erode the boundary between 'human' and 'machine' mathematics. The biggest winners are OpenAI (brand as AI research leader) and fields like optimization that benefit directly. The biggest losers are mathematicians who define themselves by solving open problems — their unique value proposition is now contested. I predict that within 12 months, OpenAI will release a paper detailing the training methodology that enabled this capability, and within 24 months, at least one major mathematics journal will publish an AI-generated proof as a sole-author paper.
- OpenAI will announce a dedicated 'Math Reasoning' API tier by Q1 2027, with pricing at least 5x the standard GPT-5.6 rate, targeting academic and industrial research labs.
- Google DeepMind will respond by releasing a formal verification bridge for AlphaProof within 6 months, claiming superior trustworthiness over GPT-5.6's unverified proofs.
- The International Mathematical Union will form a committee on AI-generated proofs by mid-2027, issuing guidelines for attribution and verification.
- July 2026Reddit post reports GPT-5.6 proof
User 'ProofBot2026' posts on r/math that GPT-5.6 solved a convex optimization gap.
- July 2026Hacker News aggregates discussion
HN thread amplifies the story, adding expert verification comments.
- August 2026Expected OpenAI response
OpenAI likely to publish technical details or a paper on the methodology.
Estimated Time to Solve Open Math Problems: Human vs. AI (2026)
- GPT-5.6 solved a 30-year open problem in convex optimization from a single prompt, verified by human experts within 48 hours.
- The proof is novel and not a retrieval of existing work, suggesting genuine reasoning capability.
- Lack of formal verification (Lean/Coq) leaves a trust gap that competitors like DeepMind can exploit.
- This event will accelerate adoption of AI in mathematical research and pressure journals to accept AI-generated proofs.
- OpenAI gains a decisive lead in autonomous reasoning, but reproducibility and training data transparency remain open questions.
Source and attribution
Hacker News
GPT-5.6 used a prompt to close a 30-year gap in convex optimization
Discussion
Add a comment