,

AGI Is General, Until the Real World Opens the Wrong File

Frontier AI keeps smashing benchmarks while still stumbling over ordinary real-world work. Maybe AGI needs reality, not another scorecard, before the parade.

By

Published


An AI terminal stands beside a high benchmark score while a worker faces several incompatible files and applications.

Lede

AGI keeps arriving in benchmarks while ordinary reality keeps forgetting to read the press release.

Words used

  • AGI: Artificial general intelligence. There is no single universally accepted test or definition of when it has been reached.
  • ASI: Artificial superintelligence. A hypothetical intelligence substantially beyond human ability across many fields. It is a stronger claim than AGI.
  • Benchmark: A defined test or task suite used to compare model performance under measurable conditions.

Hermit Off Script

AGI has become a private dictionary. Every camp seems tempted to define it in the way that best fits what it wants to build, sell, regulate or fear. I don’t deny the progress for a second. AI in 2026 is clearly better than it was a year ago, and dramatically better than two years ago. But doing everything better than humans? Not a chance, at least not in the world I use. It speaks like us, sometimes quacks convincingly like us, codes, researches, browses and can occasionally perform work that would take a person hours. Then you hand it the wrong file, the wrong tool, a messy website, a broken workflow or something outside the neat environment prepared for it, and suddenly the general intelligence needs a chaperone. Benchmarks matter, but benchmark victory is not the same thing as surviving reality. My own line for AGI is deliberately brutal and practical. If you want to call a model general, it shouldn’t cough at the ordinary digital world. Text, audio, video, spreadsheets, PDFs, old formats, strange extensions, version mismatches, websites, operating systems – basically the things humans already touch every day. It should know how to approach them, find the right tool, convert what needs converting and carry on. I am not saying every file extension is an intelligence test. I am saying “general” starts looking suspiciously narrow when reality has to be cleaned and prepared before the genius enters the room. Then comes the control theatre. We hear that models may become uncontrollable, while the same companies decide who gets which model, which tools, which safety restrictions, which usage allowance and which price. For a hosted model, a provider can stop access. That isn’t a magical red button for every AI system on Earth, especially once model weights or copies exist elsewhere, but it does suggest that the immediate control problem is still heavily about deployment, permissions and humans. Regulation can make sense as well. Verified identity could make some online enforcement easier, but banking already proves that knowing someone’s identity doesn’t abolish fraud or unlawful behaviour. And I don’t believe the race simply stops. Some engineers will watch models absorb more coding work. Lab leaders will worry about misuse, regulators, competitors, reputation and public reaction. Workers will worry about becoming somebody else’s efficiency saving. My worry is simpler: frontier intelligence stays concentrated in a few hands, everyone else pays for smaller tiers, and we call the resulting dependence abundance. Maybe that future never arrives. But if it does, the joke won’t be that AGI took everyone’s jobs. It will be that we pay a monthly subscription for permission to use the machine that took them.

What does not make sense

  • Calling intelligence “general” while every organisation is still free to choose the definition, benchmark and finishing line.
  • Treating a near-perfect benchmark score as proof of universal ability when performance can fall sharply on another test involving ordinary computer work.
  • Advertising multimodality as if accepting text, images and audio automatically means dealing reliably with the untidy formats, applications and versions people actually use.
  • Talking about one mythical AGI off-switch when hosted services, open models, local copies and tool-connected agents create very different control problems.
  • Predicting mass unemployment with certainty when current labour evidence still points more strongly towards uneven task transformation than instant replacement.
  • Demanding either zero regulation or maximum regulation before agreeing on what capability is actually being regulated.

Sense check / The numbers

  1. OpenAI’s GPT-6 Astra scores 99.9 per cent on ARC-AGI-3, but 72.6 per cent on OSWorld 2.0 and 41.4 per cent on AutomationBench. Extraordinary capability, three very different measures of what “intelligent” means. [OpenAI]
  2. Stanford’s 2026 AI Index says frontier performance on Humanity’s Last Exam improved by about 30 points in a year. Yet its review also reports OSWorld computer-use performance at 66.3 per cent and real household robot success at only 12 per cent. The capability curve is rising quickly, but it is still jagged. [Stanford HAI]
  3. METR’s current time-horizon work uses more than 100 mainly software engineering, machine-learning and cybersecurity tasks. METR explicitly warns that estimates above 16 hours are unreliable with its current suite and says agents perform worse on messier tasks scored more holistically. [METR]
  4. OpenAI’s Charter defines AGI around outperforming humans at “most economically valuable work”. Google DeepMind instead built an AGI framework from 6 principles and separates performance, generality and autonomy. Even the people building towards AGI do not begin with one identical ruler. [OpenAI] [Google DeepMind]
  5. The ILO’s 2025 global study found 25 per cent of employment had some exposure to generative AI, while only 3.3 per cent fell into its highest exposure category. Its conclusion was that job transformation was more likely than wholesale replacement because many occupations still require human input. [ILO]

The sketch

Scene 1: The AGI certificate
An AI silhouette stands on a spotless laboratory podium beside an enormous benchmark score while executives applaud.
Dialogue:
Executive: “AGI. Look at the score.”
Model: “Which test?”

Scene 2: The real world
The same AI sits at a cluttered desk facing an old spreadsheet, an unfamiliar file and a half-working browser.
Dialogue:
User: “Open this old file.”
Model: “Unsupported format.”
User: “But you’re general.”

Scene 3: Abundance, monthly
The AI now sits behind a locked glass door marked with several access tiers while a worker holds an empty office box.
Dialogue:
Executive: “Frontier model. Premium access.”
Worker: “And my job?”
Executive: “Basic tier available.”



What to watch, not the show

  • Benchmark incentives. Once a score becomes prestigious, companies have every reason to optimise models, harnesses and demonstrations around it.
  • Interoperability. Real general usefulness depends on models dealing with applications, formats, permissions, operating systems and broken workflows, not merely answering prompts.
  • Access concentration. The important question may become less “Who owns intelligence?” and more “Who controls access to the strongest deployed intelligence?”
  • Regulation by capability. The EU’s general-purpose AI obligations already distinguish models by risk rather than waiting for somebody to declare AGI. Enforcement with fines began on 2 August 2026 for applicable providers.
  • Labour distribution. Productivity gains mean very little politically if the owner receives the saving while the worker receives the redundancy letter.
  • Hosted versus distributed models. A provider can restrict its own service far more easily than society can recall every model, weight, derivative or local deployment.
  • Pricing. “Abundant intelligence” deserves a second reading when the strongest capability remains metered by subscription, usage allowance or API bill.

The Hermit take

AGI should survive contact with ordinary reality, not only the benchmark harness.
Until then, the score can be superhuman while the product remains gloriously mortal.

Keep or toss

Keep / Toss

Keep the astonishing progress and serious testing.
Toss the habit of treating every benchmark jump as a certificate announcing a new species.

Sources

  • OpenAI Charter: https://openai.com/charter/
  • OpenAI GPT-6 Astra: https://openai.com/index/gpt-6-astra/
  • Stanford AI Index 2026 technical performance: https://hai.stanford.edu/ai-index/2026-ai-index-report/technical-performance
  • METR task-completion time horizons: https://metr.org/time-horizons/
  • Google DeepMind Levels of AGI: https://deepmind.google/research/publications/66938/
  • European Commission general-purpose AI obligations: https://digital-strategy.ec.europa.eu/en/factpages/general-purpose-ai-obligations-under-ai-act
  • European Commission GPAI enforcement guidance: https://digital-strategy.ec.europa.eu/en/faqs/questions-and-answers-code-practice-general-purpose-ai
  • ILO Generative AI and jobs, 2025 update: https://www.ilo.org/publications/generative-ai-and-jobs-2025-update

Disclaimer: This article is satire and commentary. AGI, future access models and long-term labour outcomes remain uncertain. Current benchmark results measure specific evaluation conditions, not universal intelligence.


Satire and commentary. Opinion pieces for discussion. Sources sit with the article. Nothing here is legal, medical, financial or professional advice.

Leave a Reply


One roast at a time

No spam. No motivational soup. Just the latest receipt when it is ready.

JOIN OUR NEWSLETTER
One roast at a time. No spam. No motivational soup.