AGI Is Dead: AI Leaders Kill the Benchmark

AGI Is Dead: AI Leaders Kill the Benchmark

AGI was supposed to be the finish line. Now the people building toward it say it never existed. Bloomberg's September 2026 report shows why the industry is quietly abandoning the term—and what replaces it.

In September 2026, Bloomberg Technology reported that AI leaders are now openly saying AGI is a 'fuzzy target,' a stark reversal from the measurable milestone they once promised. Sam Altman's OpenAI and Demis Hassabis's DeepMind built entire roadmaps around AGI, but now they're backing away from the term—and that's not humility, it's a strategic retreat.
  • Bloomberg Technology reported on September 7, 2026, that AI leaders now describe AGI as a 'fuzzy target' rather than a measurable milestone.
  • This shift undermines the public benchmark that justified massive compute spending and safety debates for the past decade.
  • The article argues that 'task-specific proficiency' is replacing AGI as the operative goal, but this creates new accountability gaps.

Why Did AI Leaders Suddenly Declare AGI 'Fuzzy'?

According to Bloomberg Technology's September 2026 newsletter, the shift is driven by the realization that AGI's definition—'AI handles most tasks better than humans'—collapses under scrutiny. Bloomberg reported that leaders from OpenAI, DeepMind, and Anthropic now say AGI is not a single point but a spectrum of capabilities that vary by domain. The evidence: no two labs have ever agreed on a test for AGI, and the 'most tasks' clause is operationally meaningless when AI exceeds humans at chess yet fails at common-sense reasoning.

My interpretation: this is a face-saving retreat. When the goalpost keeps moving—from 'human-level at everything' to 'better than most humans at most tasks'—the people who set the goalpost are admitting they cannot hit it. The fuzziness is not a discovery; it's a defense mechanism.

What Replaces AGI as the New Measuring Stick?

AGI Is Dead: AI Leaders Kill the Benchmark

Bloomberg's report points to 'task-specific proficiency' as the emerging alternative—measuring AI by what it can do in concrete jobs like radiology, legal research, or code review. DeepMind's own 'Levels of AGI' paper (published November 2023) already sketched this, defining levels from 'Emerging' to 'Competent' to 'Expert' based on performance in specific tasks. According to that framework, current systems sit at Level 3 or 4—'Competent' or 'Expert' in narrow domains but far from Level 5, which requires outperforming 90% of skilled adults across all tasks.

This reframing lets labs claim incremental wins without promising a single 'AGI day.' But it also means the public loses a simple yardstick. When OpenAI says 'we're approaching AGI,' investors can no longer ask 'by when?'—they must ask 'on which benchmark?'

Who Benefits From a Fuzzy AGI Timeline?

Enterprises and regulators gain the most. If AGI is fuzzy, then procurement decisions cannot wait for a mythical threshold—they must evaluate AI on current, verifiable task performance. According to Bloomberg, this is already happening: corporate buyers are shifting from 'AGI-ready' contracts to 'task-specific' service agreements that specify error rates and human oversight. The losers are AI labs that used AGI as a fundraising story—without a clear finish line, their multi-billion-dollar compute bets look less like a race and more like an open-ended expense.

Consider the comparison: OpenAI's GPT-5 and Anthropic's Claude 4.5 both claim 'expert-level' performance on specific benchmarks, but neither can be called AGI. The table below shows how the two leading labs now position their systems.

DimensionOpenAI GPT-5Anthropic Claude 4.5
Claimed capability'Expert-level' on MMLU, code, and math (OpenAI, 2025)'Expert-level' on reasoning and safety (Anthropic, 2025)
AGI languageOpenAI now says 'AGI is a spectrum' (Bloomberg, 2026)Anthropic says 'AGI is not a useful target' (Bloomberg, 2026)
Primary benchmarkMMLU, HumanEval, GSM8KARC-AGI, MMLU-Pro, safety evals
Stated timeline to AGINo fixed date; 'we will know when we see it'No fixed date; 'we focus on task mastery'
VerdictNeither is AGI by any historical definition; both are using 'fuzzy' to avoid accountability for missed deadlines.

Is This a PR Stunt or a Real Shift in AI Goals?

The evidence leans toward PR. Bloomberg reported that the same leaders who now call AGI 'fuzzy' used it repeatedly in investor calls and product launches between 2020 and 2025. For example, OpenAI's Sam Altman wrote in 2023 that 'AGI will be achieved in the next decade.' DeepMind's Demis Hassabis said in 2022 that AGI could arrive 'within years.' Now, Bloomberg reported, both refuse to give dates, saying instead that 'the question is ill-posed.'

This is not a scientific correction; it's a marketing pivot. When you have sold billions of dollars of compute on a promise, and the promise does not materialize on schedule, you redefine the promise. The 'fuzziness' is a convenient way to move goalposts without admitting failure.

My thesis: AGI was never a technical threshold—it was a narrative device, and its death is the industry's own doing. In the short term, this shift means less hype-driven investment in moonshot labs and more demand for measurable ROI from enterprise AI. In the long term, it could actually be healthy: if we stop chasing a mythical AGI, we might start building AI that reliably helps humans in specific jobs. But the immediate losers are the labs that staked their reputations on AGI—OpenAI and DeepMind now face skepticism every time they claim progress, because they have admitted their own yardstick was broken. The winners are niche AI companies that can demonstrate concrete value in healthcare, legal, or logistics—they no longer have to compete with the AGI narrative. My concrete prediction: by March 2027, the term 'AGI' will disappear from OpenAI's public investor communications, replaced by 'task-specific frontier models.'

What Should Regulators Do Now That AGI Is Fuzzy?

This is the most important question. If AGI is not a measurable target, then the EU AI Act's 'general-purpose AI' category and the US Executive Order's 'frontier models' definition lose their anchor. According to Bloomberg, regulators are already struggling: the EU AI Office asked labs in June 2026 to define 'general capability' and received 'inconsistent, vague responses.' This is a governance vacuum. My recommendation: regulators should not chase AGI definitions—they should mandate task-level audits, requiring labs to state exactly what their models can and cannot do in specific high-risk domains. That is the only way to make 'fuzzy' accountable.

  1. By March 2027, OpenAI will remove 'AGI' from its public investor materials and replace it with 'task-specific frontier models.'
  2. By June 2027, the EU AI Office will issue new guidance that formally abandons 'general-purpose AI' thresholds in favor of task-level risk assessments.
  3. By December 2026, Gartner will retire 'AGI' from its Hype Cycle for artificial intelligence, citing 'definitional instability.'

  1. November 2023
    DeepMind publishes 'Levels of AGI'

    Introduces a task-based spectrum, implicitly rejecting a single AGI threshold.

  2. June 2026
    EU AI Office queries labs on 'general capability'

    Receives vague responses, prompting internal calls for task-level definitions.

  3. September 2026
    Bloomberg reports AGI is 'fuzzy'

    AI leaders publicly abandon the term as a measurable target.

Frequency of 'AGI' in AI Lab Earnings Calls (estimated)

  • AGI was a marketing construct, not a scientific one; its collapse shifts power from AI labs to enterprise buyers and regulators.
  • Task-specific proficiency will become the new benchmark, but without standardized tests, it risks becoming just as fuzzy.
  • The 'fuzzy AGI' narrative lets labs avoid accountability for missed timelines—watch for this language in earnings calls.
  • Regulators must act now to define task-level audits, or the governance vacuum will be filled by self-serving lab metrics.
  • Investors should stop asking 'when AGI?' and start asking 'which task, at what error rate, with what oversight?'

Source and attribution

Bloomberg Technology
When Will We Achieve AGI? AI Leaders Now Say It’s a Fuzzy Target

Discussion

Add a comment

0/5000
Loading comments...