Lede
Every new AI model arrives smarter than the last, then somehow leaves me checking whether it remembered to put the bloody door on the house.
Hermit Off Script
The difference between these supposedly smarter models is starting to look suspiciously like the difference between the companies and people behind them. Each one seems smarter than the other until I draw the line, stop looking at benchmarks and put it to work on a real project. Then I discover the same dumb smart model underneath. Yes, it’s a little smarter. Yes, it can complete more tasks as an agent or assistant. And yes, it can save me hours. But somewhere inside a complicated project there is still that unforeseen glitch in the Matrix that leaves me repairing something afterwards, or checking the stupidest little detail that probably wouldn’t have gone wrong if I had simply done the task myself. That’s the ridiculous bargain: it saves enough time that I don’t mind repairing the damage, even when I’m redoing something that shouldn’t need redoing. What annoys me more is that every repair, retry and “don’t change that, repair only this” consumes more usage. Sometimes I joke that perhaps wasting our tokens is the actual business model. That’s satire, not evidence that companies deliberately make models stupid to extract more money from us. But the economic incentive deserves watching when companies control the model, the limits, the subscription and how much intelligence you receive for it. My experience with Astra is exactly why I’m suspicious of the gap between launch-day intelligence and everyday intelligence. When I first used it, it felt genuinely smart. I could feel the difference in real work. More recently, some similar tasks have required more corrections and more wasted usage. I can’t prove that means the model was deliberately nerfed. It could be limits, routing, settings, updates, context or something else behind the curtain. But as a paying user I don’t experience the architecture diagram. I experience whether the bloody thing still does the job. Then I wanted to test Claude Sonnet 5.5. For some tasks it may genuinely be better, particularly in areas such as coding and agentic work, but trying to compare visual work reminded me how meaningless the title “smartest model” can become when products have completely different capabilities. Claude does not currently provide the same kind of native raster image generation as ChatGPT Images, so treating the visual results as a straight intelligence comparison would be unfair. And this is precisely why the endless creator videos declaring another “winner” make me laugh. Maybe one model is better at coding, another at 3D work, another at research, another at images. But for ordinary everyday creative work, these systems are still capable of mistakes that a mediocre human designer or creator would be embarrassed to make. A useful superintelligence, whatever letters we eventually settle on, shouldn’t merely score higher. I should be able to give it work and stop checking whether it broke something stupid. And if AI really becomes ordinary infrastructure, I also think the subscription model has to mature. Imagine Netflix charging a subscription and then charging according to how many scenes you watched. People would eventually get bored and leave. Maybe AI inference makes unlimited frontier access impossible today, but eventually somebody may realise that a genuinely smart model with a simple monthly price and practical freedom to use it could be more attractive than another benchmark crown. Maybe that company won’t even be one of today’s leaders. Maybe Ilya Sutskever and Safe Superintelligence will eventually surprise everybody the way ChatGPT once did. That’s imagination, not a prediction. But I know what I would rather buy: less artificial genius on the announcement page and more dependable intelligence when I’m actually working.
And then I watch Emad…
And then I watch Emad Mostaque for the Nth time and again I admire his vision and economic forecasts, but I still don’t think everything will happen as quickly as he and others, especially Elon Musk, sometimes envision. AI, SI, AGI, ASI, or whatever name they invent tomorrow, can become unbelievably capable and still achieve very little economically if people reject using it. In the end, the entire revolution depends on adoption. People have to use these systems. Companies have to adopt them. Somebody has to subscribe or pay through whatever new method companies invent to move money from our pockets into theirs. So I keep coming back to one simple question: what happens if tomorrow people simply don’t want to use AI? Or what happens if they use it, but nowhere near as much as the forecasts assume? The same applies to companies. We talk as though the moment an AI can replace a worker, the worker disappears. Real companies don’t work like a benchmark. Replacing workers and established processes takes time. There are regulations, costs, decisions, responsibilities, existing systems and numerous other factors between “AI can theoretically do this job” and actually rebuilding a business around it. That’s why my own forecast isn’t five to ten years for the sort of enormous structural change being discussed. I think ten to twenty years is more plausible, mildly and gradually, even if we have advanced robots and superintelligence somewhere inside that period. That’s my opinion, not a fact or a prediction I can prove. But there is one possibility that could make me completely wrong: perhaps the transformation builds slowly and then happens almost at once. Current AI models are becoming useful across domains because computers already provide an environment they can operate in. If robots eventually reach that same level in the physical world – genuinely usable, dependable and affordable enough to operate across ordinary human environments – then digital intelligence and physical automation could suddenly meet. Ten or twenty years of gradual adoption could compress into a much shorter period of enormous change. Maybe Emad and Elon will then look conservative. But until that happens, technological capability isn’t the same thing as adoption. A superintelligence nobody wants to pay for is still looking for customers.
What does not make sense
- Every model generation is marketed through intelligence improvements, while users still spend time checking surprisingly basic mistakes in real projects.
- An agent can complete more of a task without necessarily becoming dependable enough to leave unsupervised.
- The same mistake that creates extra human work can also consume more of a limited usage allowance when the user asks the model to repair it.
- A user’s experience of a model can change without the user having enough information to determine whether the cause is the model itself, routing, limits, settings, context or product changes.
- Calling one model “the smartest” ignores specialisation. Coding, research, visual creation, reasoning and agentic execution are different jobs.
- Technical capability does not automatically produce economic adoption.
- A company being able to replace part of a job with AI does not mean it can or will reorganise its workforce and processes immediately.
- Forecasts of rapid AI transformation depend not only on intelligence improving, but on people, businesses and institutions actually adopting the technology.
- Robots could change that equation dramatically if they eventually become as practical in physical environments as AI models are becoming inside computers.
Sense check / The numbers
- Anthropic says Claude Sonnet 5.5 scores 70.6 per cent on Terminal-Bench 4.0, compared with 10.3 per cent for Sonnet 5. That is an enormous measured improvement on that particular agentic coding test, but it does not establish equivalent improvement across every creative or everyday workflow. [Anthropic]
- METR reported that leading agents could handle software tasks with measured horizons beyond 2 full-time working days in its early-2026 evaluations. Yet in another judgement evaluation, the best internal Anthropic models reached only 59 per cent while a METR researcher scored about 90 per cent on a subset. Capability and dependable judgement are still different problems. [METR]
- OpenAI currently documents a 200-message weekly GPT-6 Pro allowance on its Pro $200 plan, alongside separate GPT-5.6 Sol Pro allowances and combined limits. It also says that reaching certain limits can trigger a switch to another available model. Whatever one thinks of the economics, “subscription” does not necessarily mean unrestricted access to every frontier model. [OpenAI]
- OpenAI released ChatGPT Images 2.5 on 8 September 2026, describing improvements in detail, editing and speed. Anthropic’s current help documentation, by contrast, says Claude does not natively produce images. That’s a product-capability difference, not evidence that one underlying language model is universally more intelligent. [OpenAI] [Anthropic]
- Safe Superintelligence says it is pursuing 1 goal and 1 product – safe superintelligence – rather than a conventional sequence of consumer products. That makes SSI interesting to watch, but it tells us nothing yet about whether a future product would be unlimited, affordable or even offered through a consumer subscription. [SSI]
- Emad Mostaque’s August 2026 interview put the effective economic life expectancy of many remote or digital jobs at about 2 years, while describing physical automation as following through cheaper humanoid robotics. That is Mostaque’s forecast, not an observed timetable for economy-wide adoption. [Peter McCormack Show]
The sketch
Scene 1: The new genius
An AI robot labelled “NEW MODEL” stands on a stage before an applauding audience. The backdrop promotes it as faster, smarter, more capable and agentic, alongside higher benchmarks and a brighter future.
Dialogue:
AI: “Much smarter.”
Audience: “Amazing!”
Audience: “Revolutionary!”
Audience: “Finally!”
Scene 2: The repair
The same AI robot is being repaired by a human working on its wheel. Beside them, a usage meter shows generating, fixing, retrying and “almost there”, with only 12 prompts remaining. A toolbox reads “PROMPT REPAIR RETRY REPEAT”.
Dialogue:
Human: “Fix only the wheel.”
AI: “Certainly.”
Meter: “12 prompts remaining.”
Scene 3: Superintelligence
A human holding a checklist stands beneath a huge tower labelled “SI”. The checklist includes details, files, formatting and links, while the supposedly superintelligent system still sends the final checking back to the human.
Dialogue:
SI: “I am superintelligent.”
Human: “Did you check everything?”
SI: “You should verify.”

What to watch, not the show
- Whether models become more reliable on long-running real projects, not merely more capable on benchmarks.
- Whether users can understand when model behaviour changes and why.
- How subscription limits and inference costs develop as AI becomes an everyday working tool.
- Whether competition eventually produces simpler or effectively unlimited access rather than increasingly complicated usage allowances.
- The gap between technical capability and actual adoption by businesses.
- Regulation, liability, integration costs and existing processes that can slow replacement of established human work.
- Consumer willingness to continue paying for multiple AI subscriptions.
- Whether general-purpose robots become dependable and affordable enough to bring AI’s digital capabilities into ordinary physical work.
- The possibility that gradual AI adoption eventually reaches a threshold where digital intelligence and robotics cause much faster structural change.
The Hermit take
The smartest model isn’t the one that wins another benchmark. It’s the one that finishes my work without creating another job called “checking the AI”.
Keep or toss
Keep / Toss
Keep the intelligence, agents, automation and ridiculous amount of time they can already save.
Toss the model coronations, artificial scarcity theatre and assumption that capability automatically means adoption.
Sources
- OpenAI GPT-5.6 and GPT-6 Pro usage limits: https://help.openai.com/en/articles/20001354-gpt-56-in-chatgpt/
- OpenAI ChatGPT Images documentation: https://help.openai.com/en/articles/11084440-images-in-chatgpt
- OpenAI ChatGPT release notes: https://help.openai.com/en/articles/6825453-chatgpt-release-notes
- Anthropic Claude Sonnet 5.5 announcement: https://www.anthropic.com/claude-sonnet-5-5
- Anthropic Claude image-generation help: https://support.anthropic.com/en/articles/9002504-can-claude-produce-images
- Anthropic Claude Artifacts documentation: https://support.anthropic.com/en/articles/9487310-what-are-artifacts-and-how-do-i-use-them
- METR Frontier Risk Report: https://metr.org/frontier-risk-report
- METR Task-Completion Time Horizons: https://metr.org/time-horizons/
- Safe Superintelligence Inc.: https://ssi.inc/
- Peter McCormack Show – Emad Mostaque, “Your Job’s Economic Life Expectancy Ends In 2 Years”: https://shows.acast.com/the-peter-mccormack-show/episodes/205-emad-mostaque-your-jobs-economic-life-expectancy-ends-in


Leave a Reply