Lede
AI companies compete to make each answer cheaper while users keep paying, in money and time, to correct the answers they already paid for.
Hermit Off Script
The endless problem with AI is the one-time prompt that somehow becomes ten prompts, twenty corrections and a small personal investigation into why I bothered in the first place. I can understand using more tokens when I’ve asked something vague, forgotten an important detail or changed my mind after seeing the result. That’s fair. I cannot expect a model to read my mind, although given the industry’s ambitions, somebody is probably preparing a subscription tier for that too. But what happens when I’ve given a clear prompt, described what I want and the model creates something with obvious problems? An image with unnatural hands, missing fingers, strange proportions or a speech bubble pointing towards the wrong character. Or written content full of invented facts, imaginary sources and confident nonsense. Now I have to spend more tokens explaining what was already explained, correcting what shouldn’t have been wrong and hoping the next attempt doesn’t ruin something that was perfectly fine before. Sometimes it works after several attempts. Sometimes it becomes a waste loop, and eventually I give up. The original task remains unfinished, but congratulations, we’ve successfully generated a conversation about generating it. What annoys me is the economics. A supposedly cheap model might cost very little per attempt, but if I need to run it ten times to get one acceptable result, how cheap was it? An expensive model might cost an arm and a leg, but if it delivers exactly what I wanted on the first attempt, it could actually be cheaper overall. And that is before counting my time, the subscription I already pay for, or the frustration of inspecting every little detail because a correction to one hand might suddenly give the character a third foot. Even with a flat subscription, repeated failures consume usage allowances and my patience. I don’t expect mathematical perfection in every creative task. Some things depend on taste, and sometimes the model should ask a sensible question before starting. But I do expect clear instructions to be followed, facts to be checked and obvious errors to be caught before the work reaches me. Why should the customer become the unpaid quality-control department? My suspicion is that the future winner won’t necessarily be the company with the largest model, the most impressive benchmark or the cheapest individual token. It will be whoever creates an AI that understands a clear request, checks its own work, corrects its own mistakes and delivers something genuinely usable the first time, at a price ordinary people can afford. Perhaps OpenAI. Perhaps Anthropic, Google, a Chinese competitor or some smaller company nobody is taking seriously yet. I don’t know who will get there first, and truly perfect output may never be possible. But if one company gets reliably close while keeping the price affordable, I suspect people will abandon less dependable competitors in large numbers.
Maybe that is the next real revolution. But here’s the absurdity: before AI has even mastered delivering one correct result from one clear prompt, the industry is already promising agents that can complete entire jobs autonomously. If a single image or paragraph still needs several rounds of corrections, what happens when an agent is responsible for ten tasks, makes mistakes along the way and confidently announces that everything is finished? Who checks the work of something that was supposedly created to eliminate the need for constant checking?
P.S. And this is exactly what bothers me about the new era of AI agents. I give an agent a detailed brief, clear instructions, quality-control requirements and sometimes an entire checklist explaining what must be checked before anything is delivered or published. Yet the agent still misses obvious mistakes. And when I point them out, what do I hear? “I missed that.” “I overlooked it.” “I forgot to check.” Brilliant. We’ve apparently spent billions developing artificial intelligence capable of reproducing the most irritating excuses of a distracted human employee. But why do I need an autonomous agent if I still have to supervise every step, inspect every result, catch every mistake and remind it to follow instructions it already has? I understand that mistakes happen. Humans make them too. But wasn’t that one of the reasons for creating agents in the first place? To automate repetitive work, follow established procedures and perform the boring quality checks that humans sometimes forget? If I’ve already explained what needs to be checked, why am I checking whether the agent remembered to check it? And if I have to inspect every task before trusting the result, how much autonomy have I actually gained? At least a human employee can genuinely forget something. An AI agent following a written checklist shouldn’t need to improvise an excuse for skipping it. Perhaps the next milestone in artificial intelligence shouldn’t be making agents sound more human. It should be making them reliable enough that humans can finally stop doing their work twice. Because for now, the promised artificial intelligence is producing a rather familiar result: artificial overtime.
What does not make sense
- The customer pays for the model’s mistakes. An unclear prompt is the user’s responsibility. An ignored instruction or invented fact is a product failure. The billing system makes remarkably little distinction.
- Cheap per token can mean expensive per result. The cheapest generation isn’t necessarily the cheapest completed job.
- Image corrections become a lottery. Fix the hand, damage the face. Fix the face, change the clothing. Fix the clothing, discover the hand has returned to its previous employment.
- Models keep failing small tasks while claiming increasingly impressive general abilities. A benchmark trophy doesn’t repair a broken image, an incorrect spreadsheet or an invented citation.
- Subscription tiers sell access, not guaranteed successful outcomes. Monthly billing may be predictable, but the amount of usable work produced isn’t.
- Users perform quality control without compensation. They supply the specifications, identify defects, request corrections and sometimes verify the same requirement repeatedly.
- The industry promotes intelligence more loudly than reliability. Following ten instructions correctly is less glamorous than solving an obscure mathematical problem. Unfortunately, customers frequently need the ten instructions.
- Perfect first-time delivery is an unrealistic universal promise. Creative preferences can be subjective and specifications can conflict. But that doesn’t excuse failures against clear, measurable requirements.
Sense check / The numbers
- About 66 per cent task success. Stanford’s 2026 AI Index reports that AI agents reached roughly 66 per cent success on OSWorld, a benchmark testing real computer tasks. That still leaves approximately one in three attempts unsuccessful under those test conditions. It doesn’t represent every AI workflow, but it makes the gap between impressive capability and dependable completion measurable. [Stanford AI Index 2026]
- 50.1 per cent accuracy on analogue clocks. The same report highlights that the top model tested read analogue clocks correctly just 50.1 per cent of the time. Extraordinary mathematical reasoning doesn’t guarantee competence on every seemingly straightforward visual task. [Stanford AI Index 2026]
- More than 280-fold reduction in inference cost. Stanford’s 2025 AI Index found that the price of running a model performing at roughly GPT-3.5’s MMLU level fell from $20 per million tokens in November 2022 to about $0.07 by October 2024. The technology became dramatically cheaper at that measured capability level. That says nothing about how many attempts a particular customer needs to obtain an acceptable result. [Stanford AI Index 2025]
- A 100-fold difference in published token rates. OpenAI’s standard API pricing lists GPT-6 Astra at $5 per million input tokens and $25 per million output tokens, against GPT-6 Luna at $0.05 and $0.25 respectively. That is a 100-fold price difference at both rates, but it doesn’t establish a 100-fold difference in reliability. Subscriptions follow different charging arrangements. [OpenAI API Pricing]
- The Modern Hermit – Editorial takeaway: The real cost of AI isn’t what you pay for one attempt. It’s what you spend getting a usable result, including failed generations, corrections and your own time.
The sketch
Scene 1: The bargain of the century
A customer silhouette hands a detailed image brief to a smiling AI machine beneath a large sign reading “CHEAP”. A small price tag hangs beside the output slot.
Dialogue:
Machine: “Cheapest AI in town!”
Customer: “One correct image, please.”
Scene 2: The correction economy
The same customer sits behind a growing stack of rejected images showing distorted hands and misplaced speech bubbles. A large counter reads “ATTEMPT 10” while the AI machine prints another receipt.
Dialogue:
Customer: “The hand is still wrong.”
Machine: “Try again!”
Customer: “Again?”
Scene 3: The actual winner
The customer stands beside a second machine displaying one clean, completed image and a single receipt. The first machine stands behind a mountain of rejected pictures and bills.
Dialogue:
Old machine: “Cheapest per token!”
New machine: “Correct first time.”
Customer: “Now that’s cheaper.”

What to watch, not the show
- Outcome-based pricing. Whether AI providers begin measuring completed, accepted work rather than concentrating on tokens, messages and generation counts.
- First-pass accuracy. How frequently a model follows all stated requirements without prompting the user to repair the result.
- Self-verification. Whether models reliably check calculations, references, visual details and a code review checklist before returning their work.
- Correction guarantees. Whether companies absorb the cost of clearly defective generations or offer fair replacement attempts, subject to abuse prevention.
- Hidden labour. The time customers spend checking hallucinations, correcting generated images and testing supposedly finished software.
- Subscription economics. Whether flat-rate access remains affordable as providers invest more computing resources in verification and internal retries.
- Competition beyond benchmarks. Whether future releases can build real things that work under ordinary conditions, rather than merely perform well in demonstrations.
- Market concentration. Whether substantially better reliability concentrates users around one provider, or whether privacy, specialist capabilities, pricing and open models preserve meaningful competition.
- The technical cost of reliability. Additional checking requires computation. The winning provider must reduce failure costs without making the checking itself unaffordable.
The Hermit take
A cheap answer becomes expensive when the customer has to finish it.
The next AI crown belongs to whoever delivers the work, not whoever apologises most fluently.
Keep or toss
Keep / Toss
Keep the ambition of affordable, highly reliable AI that gets the job right on the first attempt.
Toss the idea that cheap tokens automatically mean cheap work, and that customers should finance an endless repair cycle.
Sources
- Stanford University, AI Index Report 2026, technical performance and agent benchmarks: https://hai.stanford.edu/ai-index/2026-ai-index-report
- Stanford University, AI Index Report 2025, research and development, inference costs: https://hai.stanford.edu/ai-index/2025-ai-index-report/research-and-development
- OpenAI, current API model pricing and token rates: https://developers.openai.com/api/docs/pricing
- OpenAI, SimpleQA factuality benchmark and hallucination measurement: https://openai.com/index/introducing-simpleqa/
Disclaimer: This article contains satire, opinion, hypothetical pricing examples and speculation about future AI competition. Benchmark findings describe particular tests, not universal model performance. No company has been established as capable of delivering perfect results from every prompt.


Leave a Reply