Lede
The future worth testing is AI that is intelligent before training, not a product whose birthday is conveniently scheduled after its education.
Hermit Off Script
The future of AI intelligence, in my view, is a machine that is smart from creation, without having been trained on anything. Call it a model, an agent or whatever name eventually wins the marketing meeting. I care about what it can work out. The intelligence would already be there; the knowledge could arrive when needed. Give it information, or let it search for the information it needs, and it figures out the problem. It learns the subject, understands how the pieces fit together and does something useful with them. Finding the right page would not be enough. I would want it to solve something unfamiliar, then cope when the conditions change, rather than present yesterday’s answer with fresh confidence. I am not asking for a machine that somehow knows every fact before encountering the world. My idea is that it would not need prior training to possess the intelligence with which it encounters that world. That is the distinction that interests me. As a future possibility, imagine it learning whatever a problem requires in an instant, then moving straight to the next task. The difference between that machine and the smartest human on Earth would be the speed at which it could learn and the capacity to remember and use what it learnt. I imagine that capacity as limitless: learning, remembering, doing and solving without the ceiling that a human lifetime puts on expertise. A human might spend years becoming the expert. The machine becomes the expert when the problem arrives.
What does not make sense
- Reducing trained models to storage cupboards. Learning patterns and applying them to new tasks is more than retrieving memorised passages. Research on language models includes demonstrations of reasoning and adaptation, not merely recall. Training can produce useful capabilities; that does not settle whether a different design could possess intelligence without it.
- Pointing to a model trained to learn quickly as though it answers the question of intelligence before training. Faster education is interesting. It still involves education.
- Treating “not trained” as shorthand for “not designed”. The serious question is how a design could supply general reasoning before learning begins. Calling the proposed machine empty hardware does not answer that question.
- Equating access to information with understanding. Fetching the instructions is a different test from using them to solve an unfamiliar problem and checking that the solution works. A library membership is not an engineering qualification.
- Turning “instant” and “limitless” into engineering measurements. They belong to the vision, not the evidence. Any finite device has finite storage, and faster thinking cannot supply evidence that does not exist.
Sense check / The numbers
- Meta’s Llama 3.1 announcement on 23 July 2024 described a model with 405 billion parameters, trained on over 15 trillion tokens using more than 16,000 H100 graphics processors. Parameters are the model’s numerical settings; tokens are the pieces of text it processes. Those figures describe the scale of its preparation, not a machine arriving without training. [Meta]
- OpenAI’s 2020 GPT-3 research described a model with 175 billion parameters that could perform tasks from instructions or examples without task-specific weight updates. It had already undergone pretraining. This demonstrates an important distinction: adapting without additional training is different from having had no training at all. [OpenAI]
- In its 2018 AlphaZero report, Google DeepMind described training runs taking approximately 9 hours for chess, 12 hours for shogi and 13 days for Go. The system started with the game rules and learnt through millions of self-play games. That is impressive learning speed, but the practice was part of the achievement, not an inconvenient detail to leave out of the announcement. [Google DeepMind]
- The 2017 Model-Agnostic Meta-Learning paper investigated training models across tasks so that they could adapt to new tasks using little data and a small number of further training steps. It covered classification, regression and reinforcement learning. The starting point was deliberately trained for adaptability: learning how to learn, rather than possessing that ability without prior training. [PMLR]
- In his 2019 paper On the Measure of Intelligence, Chollet proposed evaluating intelligence through the efficiency of acquiring new skills, while accounting for prior knowledge and experience. He argued that performance on a particular task alone can conceal how much preparation produced it. That gives this discussion a useful research connection, although it neither requires nor proves intelligence without training. [Chollet]
The sketch
Scene 1: The convenient birthday
An executive unveils a machine beneath a birthday banner while a conveyor behind the stage is still feeding it training examples. An engineer points towards the conveyor.
Dialogue:
Executive: “Intelligent from birth.”
Engineer: “Before or after training?”
Scene 2: The actual hypothesis
In an imagined laboratory, a newly activated machine faces an unfamiliar mechanical puzzle. A researcher places a single sheet of rules beside the loose components.
Dialogue:
Researcher: “No prior training. Try this.”
Machine: “I’ll work it out.”
Scene 3: The speed gap
The same machine demonstrates the completed mechanism beside a human expert holding years of study notebooks. The executive steps between them with a subscription contract.
Dialogue:
Expert: “Years of study.”
Machine: “What’s the next problem?”
Executive: “Expertise is now rented.”

What to watch, not the show
- Tests that reveal the starting point. Record prior training, built-in knowledge, supplied information, practice and computing resources. Otherwise, the comparison begins wherever the brochure finds convenient.
- Whether the machine can retain what it has learnt, correct a mistaken conclusion and verify completed work. A fast first answer should not receive a lifetime exemption from checking.
- The full cost of learning and acting. Moving work from before launch to after switch-on could move the bill rather than remove it. Count the searching, experiments, energy and human supervision.
- Who gets access to the strongest capabilities, who owns the discoveries and whether users can take their accumulated knowledge elsewhere. A machine with extraordinary capacity would still be a poor bargain under a deliberately restrictive contract.
- Whether faster expertise gives people more time and opportunity, or merely gives employers another reason to demand more output. The machine’s learning speed and the owner’s generosity are separate questions.
The Hermit take
The ambition is intelligence present before training.
The proof would be what it can learn and solve, not how confidently it introduces itself.
Keep or toss
Keep / Toss
Keep the possibility of intelligence from creation.
Toss the habit of presenting adjacent achievements as proof that the whole vision has already been realised.
Disclaimer: This is satire and speculative commentary. Intelligence without prior training, instant learning and limitless capacity are proposed possibilities, not capabilities established by the research cited here. The comic depicts an imagined scenario.
Sources
- Meta, Llama 3.1 model architecture and training scale: https://ai.meta.com/blog/meta-llama-3-1/
- Meta research paper, The Llama 3 Herd of Models: https://arxiv.org/abs/2407.21783
- OpenAI, Language Models Are Few-Shot Learners: https://openai.com/index/language-models-are-few-shot-learners/
- Google DeepMind, AlphaZero training and evaluation: https://deepmind.google/blog/alphazero-shedding-new-light-on-chess-shogi-and-go/
- PMLR, Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks: https://proceedings.mlr.press/v70/finn17a.html
- Chollet, On the Measure of Intelligence: https://arxiv.org/abs/1911.01547


Leave a Reply