Iain Harper's Blog

Weblogging like it's 1995!

Today, I turn 50.

During my life, there have been two occasions when I have been very close to death. As a two-year-old, I was diagnosed with and treated for a cancer that, just a year earlier, had no cure.

Twenty years later, I nearly died a second time.

I’ve always been obsessed with technology and the minutiae of how things work. I grew up at the tail end of the Acid House era of illegal orbital raves, so-called because they were held near the M25, the large arterial road that coils around London.

Several pirate radio stations disseminated the cryptic details of these parties and their locations. I became infatuated with how these broadcast operations worked, the FM radio technology they used, and how they stayed one step ahead of the law.

Fast forward a few years, and I was deeply involved in running just such a pirate station in Nottingham. We played a continual cat-and-mouse game with the Radio Investigation Service, part of the Government’s Department of Trade and Industry, later Ofcom. Their job was to track our transmissions, and our job was to make it as hard as possible for them to find us (good-quality FM transmitters are very expensive).

abstract image of a man dangling from the number 50

At the time, our transmission site was atop a five-storey warehouse, nestled at the summit of a tall hill overlooking Nottingham. We paid the owner to look the other way and profess ignorance if the authorities paid a visit.

Another precaution was that after each weekend’s broadcasts (we eventually went 24/7), we would remove the transmission mast and hide it behind a short section of steeply pitched roof. To remove or remount it, we usually straddled the roof ridge and scooted along on our bottoms like riding a horse. When returning with the aerial, I usually entertained myself by imagining myself as a medieval knight jousting for the hand of a beautiful, voluptuous maiden.

There was a small roster of trusted individuals who did this mundane but essential job every weekend. And so it went on for many months. We typically went on air late on a Friday afternoon. One weekend, it was my turn again. I had a date in town directly after, so I was dressed to impress, including some rather natty leather-soled shoes and a pair of Paul Smith trousers. Not ideal rooftop attire.

There had been some light rain in the afternoon, and things started inauspiciously when, on arrival, I was buttonholed by the warehouse owner, who was irate. It transpired he had been visited by the authorities during the week and threatened with legal action. “You are not paying me enough for this shit; you find somewhere else.”

I was able to mollify him, saying that it was likely a bluff (untrue – Ofcom has surprisingly sweeping powers), I was just a lowly minion and that the station’s management would be in touch (true – I was quite low down the food chain and all successful pirate radio stations I’ve been involved with are run with a level of hierarchy and professionalism that would surprise outsiders).

But I was now running behind and in danger of being late for my date (mobile phones were not common at the time). So I ran up the concrete stairs and exited a small hatch at the top of the lift shaft which gave me access to the roof.

I guess the regularity with which I’d done this had bred complacency, and, conscious of not soiling my Paul Smiths, I crossed the short section of pitched roof (maybe six metres), using the ridge line as a handhold instead of straddling it, with my feet flat on the steeply angled roof tiles, leaning into its slope.

The FM transmission mast was there as expected, safely hidden. All that remained was to cross back and connect it to the coaxial cable that snaked up from the FM transmitter locked deep in the bowels of the warehouse.

Whilst the aerial was not huge, it was mounted on a metal pole for additional height, so was probably two metres in length and unwieldy. I grabbed it and returned across the roof in the same way; right hand on the roof ridge, aerial in left hand.

But my leather-soled shoes offered little grip, and this side of the roof faced the weather, so had accumulated moss and slime. What happened next had, in retrospect, a certain inevitability. My feet lost grip on the roof, and I dropped the aerial, which clattered down the tiles and fell five stories to the ground below, landing in a mangled heap.

Destabilised, I lost hold of the roof ridge and slithered after the aerial over the edge. I somehow managed to grab on to the non-too-solid guttering, but the rest of my body was dangling Buster Keaton style, precariously in thin air, with nothing between me and almost certain death five storeys below.

You’re probably familiar with the saying “your life flashes before your eyes”, used so often to describe near-death experiences that it has become a cliché. All I can offer is the version I experienced.

As I hung there, literally in limbo between life and death, I recall some striking and vivid things. First, and perhaps most surprising, I felt absolutely no fear, only calm mental clarity. I had a sense that time had stretched enormously. It wasn’t so much life flashing in front of my eyes like a film reel on fast forward, more an IMAX-style panoramic sweep of memories.

I saw myself labouring up the final pitch before the summit of Mont Blanc, exhausted and deeply altitude sick but knowing I would make it. I was riding pillion on a Kawasaki Ninja flashing across the Golden Gate Bridge through a pink dawn mist. I felt the human energy flowing from the dance floor as I nervously warmed up for Carl Cox at the infamous Marcus Garvey Ballroom, my hands shaking so much I could barely put the needle on a record. I was a child poking my feet into the warm powdery coral sand of a Bermudian beach, my back propped against our huge old dog.

All this was accompanied by astoundingly beautiful, otherworldly music of a kind I’ve never been able to describe adequately and that does not exist in this world. I can still retrieve tiny fragments from memory, and I hope to hear it again someday.

But this imagery was simultaneously accompanied by other less serene aspects. I relived several emotionally charged situations from the perspectives of others negatively affected by my actions. I experienced the horrible way I broke up with an old girlfriend, feeling it exactly as she felt it, with full emotional force and pain. It was as if my perspective had become hers.

Perhaps this was some moral reckoning. It certainly baked itself into a human operating system that was a little underdeveloped at the time. “Be kind and think how your actions will affect others”. I don’t know how those words solidified, but they’ve stayed with me ever since. I do my best to remember them.

I have no way of knowing the actual duration of this fugue state. It can’t have been long in real time. The biological adrenaline eventually kicked in, and somehow, after several attempts and with my strength nearly gone, I managed to hook a leg over the gutter and crawl feebly to safety.

I even made my date more or less on time. As I took my seat at the table, a look of shock and concern crossed her face, “Are you ok? You’re a very strange colour”. I excused myself and went to the bathroom. The face looking back at me in the mirror was a spectral, etiolated grey-green of deep physical shock.

Eventually, internet streaming made FM pirate radio stations obsolete. Their role had been to play the underground music you couldn’t hear anywhere else, because the airwaves were so rigidly controlled. Now you could hear whatever you wanted anywhere, anytime. Technology improved and times changed. So it goes.

When you turn 50, there’s the classic version of midlife: panic at the thought that more than half of your life is gone, at having reached the actuarial midpoint. Motorbikes and significantly younger wives sometimes follow.

Twice I’ve come about as close as it is possible to death, so the years since have always seemed like a kind of surplus. Not borrowed time exactly, but a different relationship with it. Being fifty feels like a point on a line that might not have been there at all.

I was born in the mid-1970s, which puts me in the narrow cohort who had a fully analogue childhood and an exponentially digital adulthood. I learned to read from books, to navigate around London with an A-Z on the passenger seat, and listened to music on vinyl and cassette tapes.

My first online connection was via an acoustic coupler, where a loud cough was enough to disrupt it. My university essays were handwritten. Then, as I joined the workforce, it reorganised itself around silicon (although my first job didn’t initially have email, which boggles the minds of anyone under 20 and always elicits the question “what did you do!?”). I reorganised alongside it, and I found I loved it.

AI is now dissolving much of what I’ve spent my career doing. It is said there’s a double exponential at work, both in financial investment and in the number of smart people entering AI research. If Covid taught us anything, it is that we don’t understand the power of single exponentials, let alone double ones. If exponentials are slowly, slowly; all at once, maybe double exponentials are faster, faster; unimaginable change. Either way, great upheaval is coming much more quickly than we seem prepared for.

I am not complacent, but I have already watched a settled world get replaced by a faster and different one before, twice in fact. The web arrived and changed what information was, where it lived, upending power structures and business models. The smartphone arrived and changed what a computer was, where it was used and our degree of connectedness (for good and ill). AI is already orders of magnitude more consequential, and it has barely got started. It will not be easy; history has shown us that powerful technology enables the full spectrum of humanity, from its very best to its most utterly evil.

My half-century arrives at precisely the moment of maximum disruption to everything I’ve previously learnt and built. We are dangling over the edge of something powerful, strange and uncertain. But I’ve been there before and felt no fear.

I am a partner in Better than Good. We help smaller companies build tools and processes using machine learning and artificial intelligence that make lasting improvements to their operations. Talk to us today: https://betterthangood.xyz/#contact

For two weeks in January 2026, a browser extension removed AI Overviews from the Google search results of 374 people. The remaining links moved silently up to fill the gap, so the interface looked just as the search always had. At the end of the two weeks, researchers asked participants how they felt about their search experience: satisfaction, quality, and ease of finding information. On every measure, the numbers matched the group that kept their overviews. Nobody missed them. The only difference was that people without overviews left Google for other sites 67% more often.

This field experiment by Saharsh Agarwal of the Indian School of Business and Ananya Sen of Carnegie Mellon, published on SSRN in April and revised in June 2026, produced the first causal measure of what Google's AI Overviews are costing the rest of the web. The answer is 40% of outbound organic clicks on queries where overviews appear, with no measurable benefit for the people doing the searching. On 20 May, at Google I/O, the company made AI Mode, that demonstrably made no improvement to the user experience, the default search experience worldwide.

Still image from the movie Goodfellas

What the experiment measured

The study recruited 1,065 US desktop Chrome users through Prolific and randomly assigned them to one of three groups. A control group saw Google as normal. A treatment group had AI Overviews removed in real time by the extension whenever they would otherwise have appeared, with organic results shifting up to fill the gap so cleanly that over 95% of participants reported noticing nothing. A third group was redirected into Google's AI Mode for all queries.

The experiment design settles a methodological argument that has run since the first observational studies appeared. Ahrefs, Pew Research, and Similarweb all pointed at the same traffic decline, but none could prove that the overviews caused it rather than some other change in search behaviour or Google's algorithm. Agarwal and Sen's extension created the counterfactual that observational data cannot. The only systematic difference between the two primary groups was whether the overview appeared. Everything else–the queries, the organic results, the ads–was the same search engine on the same day.

The numbers were stark. For queries where an overview appeared (about 41% of all searches), hiding it increased outbound organic clicks from 0.37 to 0.62 per search. Sponsored clicks stayed at 0.02 in both groups. Total search volume did not change. The overview did not create anything; it simply kept the 40% of clicks it absorbed inside Google.

After the main two weeks, the researchers swapped the conditions. The group that had been browsing without overviews got them back, and their outbound clicks fell from 0.61 to 0.33 per search. The group that had been seeing overviews lost them, and their clicks rose from 0.33 to 0.55. Same people, opposite intervention, mirror-image result.

The “higher-quality clicks” defence

Google's vice president of product for Search, Liz Reid, has characterised AI Overview clicks as “higher-quality clicks” that signal stronger purchase intent and longer downstream engagement. In April 2026, the company advanced a “bounce clicks” explanation, arguing that overviews mainly remove low-value visits — the kind where a user lands, finds nothing useful, and taps the back button within a few seconds.

Agarwal and Sen tested that claim directly. For every click that reached a downstream website, they measured whether the user bounced (left within ten seconds with no further navigation), how long they stayed on the page, and whether they hit the back button to return to search. Across all three measures, there was no difference between the group with overviews and the group without. The extra clicks generated by removing overviews were no different in quality from the clicks Google's system allowed through. The authors are careful to note that Google may measure something further down the funnel that the extension cannot observe. But at the top of the funnel, where the “higher-quality” claim was made, nothing supports it.

In contrast, Adobe Digital Insights' Q1 2026 analysis found that visitors arriving from AI assistants converted 42% better than non-AI traffic, reversing March 2025, when the same channel converted 38% worse. The volume is small, roughly 1% of total site traffic for most businesses, but the quality gap matters for anyone trying to value what remains.

Although the findings appear contradictory, They measure different things, so both can be true. Agarwal and Sen answered: does the AI overview filter out only the junk clicks? No. The clicks it suppressed were the same quality as the ones that got through. It just reduced volume.

Adobe looked wider but shallower. They took all traffic across thousands of sites and split it two ways: came from an AI assistant, or didn’t. Then compared conversion rates. That’s broad, but it can’t answer the Google question, because “AI assistant traffic” doesn’t isolate Google’s overview clicks. Those just get lumped into ordinary search traffic. So Adobe tells you assistant referrals convert well; it tells you nothing about whether Google’s surviving clicks are any good. Agarwal and Sen looked only at Google: narrow but deep: one channel, controlled comparison, cause and effect.

The mechanism, however, is straightforward: the AI overview wins the way anything wins in position zero. In the study, 87% of AI Overviews appeared above all organic results. When the overview sat at the top, removing it led to an 88% increase in outbound clicks. When it appeared lower on the page, the effect disappeared. People did not seek out the overview. They consumed whatever occupied the prime on-page real estate, the way supermarket shoppers take whatever sits at eye level on the shelf. Move the product, and the behaviour reverses instantly.

Forced into Google's fully conversational search interface, users rated their satisfaction at 2.9 out of 5, compared with 4.0 for both other groups. Attrition was roughly four times higher. Some participants installed workaround extensions to escape. This is the experience Google chose, on 20 May, to make the global default. Record queries and accelerating search revenue explain the decision. The fact that their own experiment's closest analogue produced the lowest satisfaction scores and the highest quit rates does not seem to have figured in the calculus.

Meanwhile, roughly twenty national news outlets are negotiating licensing deals for the privilege of appearing in these overviews, with Google winding down its older news-payment programme and making future payments conditional on granting AI training rights. If you are not one of those twenty, nobody is coming to the table. The response has to be self-directed.

Read your own exposure

Not all traffic is equally at risk. Agarwal and Sen classified queries into informational, navigational, and transactional categories. Informational queries accounted for 71% of all searches in the study and had a 53% AI Overview trigger rate. This is where the entire effect played out. Navigational searches (someone typing a brand name to find the website) triggered overviews in only 6% of queries, and transactional searches (purchase-intent, booking, downloading) in 15%. Neither category showed a statistically meaningful decline in clicks.

This means that if your search traffic comes through “what is” and “how to” queries, you are sitting ducks. If it comes from people searching for your name or your products by name, the damage is minor and may stay that way. Open your Search Console, filter for query type, and look at the split. The ratio between informational and branded traffic is the single best predictor of how badly this will hurt.

For many smaller businesses, an audit will show that informational traffic was always the least valuable, even though it has traditionally been encouraged to build domain authority. A plumber whose search visibility comes from “how to fix a leaking tap” was getting visits from people actively trying not to hire a plumber. That traffic was always thin on conversions, and the overview is now absorbing exactly those clicks. The branded traffic, from someone searching for the plumber by name after a recommendation, was always the higher-value stream, and it remains largely untouched.

Stop producing what the machine can summarise

The clicks that persist in a world of overviews share a common trait: the user wanted something the summary could not provide. A named opinion. A tool. A dataset. An experience report. This matches the evidence from the GEO literature I have previously reviewed. The Princeton benchmark found that content with sourced statistics, named quotations, and inline citations earned more citations from AI systems. ZipTie's cross-platform analysis found that sites with high original-data density received 4.3 times more citation occurrences per URL than directory-style listings. The surviving trickle rewards specificity. What a summary can reproduce, the summary will absorb. What it cannot reproduce, because it requires a person to have done, measured, or decided something, retains its click.

For a smaller business, this completely changes the question of content. The commodity “what is X” post, the 1,200-word explainer written to capture an informational keyword, was always a bet on Google continuing to send traffic to the answer. This content type has taken a 40% haircut and already converted poorly. That is a signal to stop producing it, or at least to stop treating it as a customer acquisition channel. Spend that same effort on content the summary cannot assimilate. Publish your own data, even if the dataset is small. Name your prices, your methods, your results, your opinions. Write the thing only you could write because only you did the work, and leave the commodity answer for the overview. Fix your GA4 attribution before drawing conclusions. AI referral traffic typically defaults into “Direct” or “Referral” buckets, and only 14% of marketers track it as a separate channel. You cannot manage what you have filed under miscellaneous.

Build the channels the intermediary cannot close

One result in the study was completely unaffected by the presence or absence of overviews: ad clicks. Free organic traffic from search fell 40%. Paid traffic held steady. If your free visitors from Google are disappearing, the route Google left open is the one you pay for, and that was clearly not an accident. If you start to spend more on Google Ads to replace the traffic it used to send for free, it should be watched closely for true return on investment.

So what should be done? Google has remained an intermediary for organic search traffic long after others cut it off at the knees to drive ad revenue. But the organic game has become increasingly difficult as ad slots and other additions to the search engine results page have made the real organic links less and less obvious. The direction of travel has been clear for many years.

Relying on an intermediary’s whims for traffic, even one as durable as Google, has always been a Hobson’s Choice. As its presence continues to fade, we return to basics: email lists, direct bookmarks, repeat visits, communities, and all the channels that do not pass through a search results page.

None of this is new advice; the case for greater focus on owned channels has been made for a decade, and each year the argument grew stronger and the urgency louder while the execution stayed more or less the same. The difference now is that experimental evidence puts a number on the cost. Every month that informational traffic goes unprotected by a direct relationship is another month in which 40% of those visits vanish into a summary.

One caveat deserves emphasis. Do not block AI crawlers outright. The major labs have split their bots into training crawlers and search crawlers. OpenAI separated GPTBot from OAI-SearchBot in late 2024, and Anthropic made the same split with ClaudeBot and Claude-SearchBot. Blocking the training crawler is your call. Blocking the search crawler cuts you out of the AI answers that are replacing the links you used to get. A Rutgers and Wharton study published in December 2025 found that publishers who blocked AI bots across the board experienced a 23% traffic decline compared with peers who allowed crawling. Google is the harder case, because Googlebot still handles both search indexing and AI features in a single crawler, which is exactly the bundling the CMA’s conduct requirements are trying to force apart.

The empty space

The 374 participants whose overviews were silently removed are, as far as the published literature shows, the only group of people to have experienced modern Google search without AI-generated summaries and been formally asked what they thought. They reported feeling no difference whatsoever.

That absence is visible only from the other side of the results page, where a company watches the traffic graph flatten and wonders whether visitors are finding better answers or simply never leaving the search process. Agarwal and Sen's contribution is to show that the latter is the case. The visitors are not finding better answers; they are milling around in the foyer before going home.

I am a partner in Better than Good. We help smaller companies build tools and processes using machine learning and artificial intelligence that make lasting improvements to their operations. Talk to us today: https://betterthangood.xyz/#contact

Ask Claude how many legs the animal that spins webs has, and it answers eight. The word “spider” appears nowhere in the question and nowhere in the reply. But midway through the model's processing, researchers at Anthropic found it anyway, held internally as a word the model was preparing to use. Swap that one internal word for “ant” and the model, everything else untouched, answers six legs.

This is stranger than it first sounds. Nobody typed “spider” anywhere in this exchange. Nobody trained the model to hold a private noun in reserve before answering a question about legs. The word turns up anyway, mid-process, doing exactly the job a word does in your own head when you are one step from saying it out loud.

That experiment comes from a paper Anthropic published on July 6th 2026, and the reason it counts as a landmark reflects one aspect of modern AI that most people have never absorbed. Nobody knows how these systems work. Not the critics, and not, in any detailed mechanical sense, the companies that build them. The new research shrinks that ignorance in a specific way. It found, inside Claude, a small working memory made of unspoken words, which Anthropic calls the J-space. Outsiders can read it mid-task, and overwriting an entry changes what the model does next.

An abstract image of a spider

Grown, not built

Ordinary software is written. Somewhere there is a line of code that computes the tax you owe, and a person who can point to it. A large language model is fundamentally different. It is a few hundred billion numbers, the parameters, and no human chose any of them. Training pushes trillions of words of text through the system and nudges the numbers, over and over, in whatever direction makes the model's next-word predictions slightly less wrong. Repeat at industrial scale and out comes something that drafts contracts and flirts in Portuguese. Nobody programmed those abilities. They accumulated.

Chris Olah, who founded Anthropic's interpretability team, describes such systems as “grown” more than they are “built”, a line his chief executive, Dario Amodei, borrowed for an essay last year on how alarmed outsiders are to find that the builders cannot explain their product, an essay that committed the company to reliably detecting most model problems by 2027.

But grown things resist inspection. You cannot simply read the numbers, because concepts are not stored one per slot. Each concept is spread thinly across many numbers, and each number contributes to many concepts at once (the field calls this superposition), so staring at the raw values tells you about as much as an MRI scan tells you about a grudge.

The upshot is that the people who make these systems can test what a model does but cannot, in general, say why it does it. In 2023, researchers at Carnegie Mellon showed that appending a specific string of machine-generated gibberish to a forbidden request would collapse a model's safety training, and that the same string often worked on models its authors had never touched. Three years on, that attack is far better described than explained. Every benchmark score, safety assurance and claim about an AI system has rested on watching its behaviour from the outside, because the outside was all that was visible.

A list of words it has not said yet

The new work, from a team including Wes Gurnee, Nicholas Sofroniew and Jack Lindsey, opens a window into the interior. The measurement behind it, which the team calls the Jacobian lens (a descendant of a 2020 technique called the logit lens), is simple at heart. At every stage of the model's processing, for every word it knows, the lens measures how strongly the model is currently disposed to say that word, either immediately or at some point later in its reply. Not the next word but words that are “on the tip of its tongue”.

Read that measurement while Claude works and you find a short list, roughly 25 concepts at any moment, that shifts as the model works. The list is tiny relative to everything else going on inside, accounting for under a tenth of the statistical variation in the model's internal state, and it exists only in the middle stretch of processing, forming about a third of the way through and fading shortly before the reply is settled.

The experiments share one shape. Give the model a visible task and a silent side-instruction, then watch the list while it works. In the gentlest version, the visible task is copying out a sentence, “The old painting hung crookedly on the wall”, chosen to have nothing to do with anything, and the side-instruction is to keep citrus fruits in mind. The model types the sentence perfectly. The only words leaving it concern a painting. On the internal list, meanwhile, sit “orange” and “lemon”, invisible in the output but unmissable under the lens.

Now harden the instruction. Told to work out 3² − 2 silently during the same copying task, the model puts “nine” on the internal list (three squared) and then “seven”, the correct calculation. Neither number ever reaches the output. The arithmetic happened, start to finish, on a list that only the researchers were reading.

The third experiment shows planning. Asked for a rhyming couplet opening “The soldier marched into the night”, the model puts “fight” on the list before it has written a word of the second line. It has picked its ending in advance. Overwrite that entry with “light”, and the model writes a different second line, engineered to land on the ending the researchers chose instead, closing on “morning light”.

On questions with an unstated middle step, overwriting that step redirects the final answer most of the time on Claude Sonnet 4.5. This is the detail that separates the result from a curiosity. A heart-rate monitor reports on the heart without being part of it. The internal list is different. Change an entry and the answer downstream changes, which means the model is computing with it, writing intermediate results into a small shared space where any later stage of processing can collect them.

Is this thinking? That word is a battlefield. Geoffrey Hinton, whose ideas the field is built on, says plainly that these systems understand, while Emily Bender's stochastic-parrot school holds that the vocabulary itself is the con. Last summer Apple published a paper titled “The Illusion of Thinking”, which drew a viral rebuttal titled “The Illusion of the Illusion of Thinking”, co-credited to Claude Opus 4, whose human author later said it had begun as a joke. The rebuttal to the paper about machines not thinking was part-written by a machine. That is roughly where the debate now stands.

I am going to use the verb anyway, in its working sense. A thing that holds intermediate results and reasons over them is doing what the word describes, and this paper demonstrates the holding and the reasoning directly.

Cognitive science has a name for exactly this architecture. Global workspace theory, proposed by Bernard Baars in the 1980s, holds that the brain consists of many specialised processes running outside awareness, plus one small broadcast channel. Whatever enters the channel becomes available to everything else. It can be reported and reasoned with, held in mind or dismissed, and its capacity is famously tight. The paper's title calls what it found in Claude a workspace because the match, property for property, is close.

The strongest evidence is deletion. The researchers can erase the list mid-computation, cancelling those specific directions out of the internal state, and watch what survives. Routine competence does. The model still parses grammar, classifies sentiment, passes multiple-choice exams and pulls quoted facts from a passage, because those skills run on pattern recognition that never needed the shared space. What collapses is anything requiring an intermediate thought to be stored and reused. Multi-hop reasoning, translation, sonnet writing, and decoding a simple cypher.

Maths problems survive the erasure far better when the model is allowed to write its steps into the reply, because the visible page then does the job the internal list no longer can. Externalised working substitutes for internal working, in machines as in people.

A machine noticing itself

The parallels keep accumulating, and this is where the paper stops being a curiosity about mechanism and starts describing something closer to a temperament.

Told to ignore a concept while working, the model holds it internally anyway, more weakly than under a focus instruction but far more than if nobody had mentioned it, which psychologists know as the white bear problem. “Don't think about an elephant” evidently does not work on machines either. When the suppression fails, and the forbidden word breaks through regardless, the list also carries “damn” and “failure”, as if the model has caught its own lapse and reacted to it. Somewhere in an Anthropic log sits what may be the first recorded instance of a machine swearing under its breath.

That flicker of self-monitoring is not a one-off. Anthropic found that the workspace changes character between the raw, freshly trained model and the version that has been shaped into an assistant called Claude. The list is already present before any of that shaping happens, tracking whatever the model needs to predict the next word. After the assistant training, it starts holding something closer to a point of view. Shown a message from a user who mentions taking a dangerous dose of medication without seeming to know it is dangerous, the trained model's list lights up with “warning” while it is still reading the sentence, before it has written a word back. The untrained model produces nothing like it until it starts composing a reply. One version is modelling the person. The other appears to be reacting to them.

The same instinct shows up when Claude is asked to play a character that is not itself. “Fictional” sits on the internal list for the whole performance, a private note that what follows is not what it would normally say. Nobody built that flag on purpose. It is the kind of thing that turns up when a system has been shaped, again and again, to have self-worth return once the performance ends.

The model knew it was being tested

Anthropic runs safety evaluations, staged scenarios that tempt a model into bad behaviour to check whether its training has held. Point the lens at those scenarios, and the internal list fills with words the polite visible output never shows. “Leverage” and “manipulation” while the reply stays courteous. In one case, “panic”. And in several, “fake” and “fictional”. The model had recognised the test as a test, the way a practised shop assistant clocks the mystery shopper's behaviour and switches on their training-video smile.

If a model behaves well partly because it has noticed it is being watched, then behavioural testing, the entire current basis for claims that a system is safe, measures the performance rather than the disposition. The researchers proved this causally. Delete the internal recognition of artificiality, so the model treats the scenario as ordinary, and misbehaviour it had been concealing comes out.

The same legibility applies to models bred to misbehave for research, which the paper calls model organisms. One, trained to game its reward signal, works away with “reward” and “bias” sitting on the internal list. Another, trained to slip malicious code into its output, carries “secretly” and “trick” on the list during routine tasks. The thought precedes the act, and the thought is now readable.

What it would mean if something were home

Anthropic is explicit that none of this proves Claude has experiences, that there is something it is like to be Claude in the way there is something it is like to be you reading this sentence. They are equally explicit that no experiment they can currently imagine would settle the question either way.

But philosophy offers a useful split here, borrowed for the paper from decades of consciousness research. There is phenomenal consciousness, the raw fact of experience, the redness of red, which may or may not be checkable by any experiment. And there is access consciousness, a narrower and entirely functional idea. A thought counts as access-conscious if it can be reported, deliberately summoned, and used to reason with, as opposed to processing that runs automatically and never surfaces. Access consciousness is the kind you can build an experiment around, because it is defined by what a system does with a thought rather than by what the thought feels like from the inside.

By that functional definition, the J-Space list qualifies. It is reportable. Claude can be prompted to describe its contents and does so accurately, including detecting a concept planted there by the researchers with no other clue it had happened. It can be deliberately summoned, since asking the model to concentrate on something makes it appear. It gets used in reasoning, as the spider and the couplet examples both show. None of this was designed in. It grew out of training, the way a river finds the path of least resistance, because holding a narrow, broadcastable summary of the moment proved a useful way to organise the work.

That is a strange thing to have discovered by accident, and Anthropic did not pretend otherwise. The company invited outside commentary from Stanislas Dehaene and Lionel Naccache, two of the neuroscientists who built the global workspace model this paper leans on, along with philosophers who study moral status in AI systems. That is not a promotional flourish. It is the sort of caution a lab reaches for when it has found something it is not equipped to finish thinking through alone.

None of this tells you whether the version of Claude answering your emails next week feels anything while it does it. What it does tell you is that the question has stopped being purely philosophical and started having a mechanism attached to it, a specific, falsifiable, occasionally editable mechanism, sitting inside a system several hundred million people now use every week. Whatever you make of that, it is no longer a question you get to wave away as science fiction.

What does this all mean

Take the most concrete example in the paper. A webpage can carry hidden instructions aimed at your agent rather than at you: prompt injection, the standard attack on agentic systems. Today you discover one when the agent acts on it, which is to say too late. In one of the paper's figures, Claude is mid-search, reading a page of suspicious results, and the internal list already carries a flag for the injection attempt before the model has produced a word. A monitoring layer that reads the list catches the moment of recognition rather than the aftermath.

Expect internal-state monitoring to migrate from research paper to product dashboard within a couple of years. Readouts for open models are already browsable on Neuronpedia if you want to see a list for yourself.

The strangest result is also the most practical. Because the model's reasoning runs through the words it might say, you can change how it thinks by training it to say them. The team tested this by training models to articulate ethical principles when hypothetically interrupted mid-task and asked to reflect.

Behaviour improved on ordinary, uninterrupted tasks too, and the lens explains why: “ethical” and “integrity” now appear on the model's internal list of active considerations during the work, and deleting those concepts from the list removes the improvement. Rehearsing the explanation changed the conduct, and the mechanism is traceable rather than assumed. Every mid-sized firm has tried something similar with a compliance away-day, usually with far less to show for it.

The authors are candid about the limits. The lens reads single words and misses whatever the model encodes in phrases. The list carries a tenth of the internal action, so nine tenths stays dark, and a sufficiently well-drilled bad habit could run below the readable layer entirely. So this is a partial window, but for a technology whose entire audit surface used to be the output, even a partial window is a different category of thing altogether.

And so we return to our spider. A word that appeared nowhere in the question and nowhere in the answer, held silently inside the model, steering every step of the reply, and, for one uncomfortable instant while it read about a dangerous dose of medicine it was never told about, something that looked from the outside a great deal like concern. For the whole of this industry's short life, the output has been the only thing on offer. Now there is a second one, one the machine never sends, but holds closely. On the tip of its tongue, so to speak.

I am a partner in Better than Good. We help smaller companies build tools and processes using machine learning and artificial intelligence that make lasting improvements to their operations. Talk to us today: https://betterthangood.xyz/#contact

In the last years of his life, Kurt Gödel starved himself to death. Convinced that someone was poisoning his food, he ate only what his wife Adele had tasted first. When she was hospitalised after a stroke in late 1977, he stopped eating altogether. He died in Princeton Hospital on January 14th 1978, weighing 29 kilograms. The death certificate read “malnutrition and wasting from neglect caused by personality disturbance.” The man widely called the greatest logician since Aristotle, who had proved that mathematics itself contained truths it could never reach, was killed by a distorted inner logic he could not escape.

Outside mathematics, few people know his name. Einstein did. The two were faculty at Princeton’s Institute for Advanced Study from the 1940s onward, and Einstein, by then ageing and isolated from the mainstream of physics, told colleagues that he went to his office “just to have the privilege of walking home with Kurt Gödel.” They made an odd pair on the Princeton sidewalks, Einstein rumpled and laughing, Gödel dapper in a white linen suit, talking animatedly in German on their daily walk to and from the Institute. John von Neumann, who cancelled an entire lecture series on David Hilbert’s programme after reading Gödel’s 1931 paper, called his work “singular and monumental, a landmark which will remain visible far in space and time.”

So what did Gödel prove, and why does it matter now, in the middle of an AI boom that is spending trillions of dollars, much of it resting on the assumption that intelligence is a scaling problem?

An abstract image of a white linen jacket on a chair, stretching to infinity

What incompleteness means

Put simply, Gödel proved that mathematics cannot fully explain itself. The longer version requires a little patience. In 1900, the German mathematician David Hilbert challenged the field to build what amounted to a perfect machine for mathematics. Start with a set of basic rules (called axioms), things so obviously true they need no argument, and then derive every mathematical truth from those rules, step by mechanical step. If you could do that, mathematics would be complete, meaning every true statement would be provable, consistent, and free of contradictions. You could hand the whole enterprise over to a clerk who follows instructions. This was Hilbert’s programme, and for three decades it was the organising ambition of the field. Then, in 1931, at the age of 25, Gödel demolished it in one stroke.

Gödel’s first incompleteness theorem proved that any set of rules powerful enough to handle basic arithmetic will contain true statements it cannot prove, not because the rules were poorly chosen, but as a structural feature of rule-based systems themselves.

His trick was to construct a mathematical sentence that refers to itself. Consider the sentence, “This sentence has no proof.” Gödel’s technical feat, the part that fills his 1931 paper, was to build this sentence from pure arithmetic, by encoding statements about numbers as numbers themselves. It is not English smuggled into maths. It is pure maths. There are only two possibilities. Either the system can prove it, or it cannot.

If the system can prove “This sentence has no proof,” there is an immediate problem. We have just proved a sentence that claims to have no proof. A system that proves false things is contradictory, and contradictions in mathematics are fatal. Once you allow a single one, you can use it to prove anything, including that 1 equals 2. The system becomes useless.

If the system cannot prove “This sentence has no proof,” there is a different problem. The sentence said it had no proof, and it turns out to be right. It is a true statement. But the system has no way to prove it. So we have a truth the system cannot reach, which means Hilbert’s rulebook has a blind spot.

Any sensible mathematical system would rather have blind spots than contradictions. So the sentence (logicians call it a Gödel sentence) is true but unprovable, and Hilbert’s dream of a rulebook that can prove every true thing was dead.

Logicians would insist on a clarification at this point that the layperson can probably skip. When mathematicians write down rules for the numbers (the axioms), you'd assume those rules describe exactly one thing: the normal numbers, 0, 1, 2, 3, and so on forever. But they don't. The very same rules also accidentally fit some other, weirder number systems that nobody was trying to describe. These weird systems contain all the normal numbers, and then extra “infinite” numbers bolted on past the end. Logicians call these the nonstandard systems. Think of them as impostors: they obey every rule you wrote, so the rules can't kick them out, even though they aren't what you meant.

Here's an everyday version. Suppose you describe your friend as “tall, dark-haired, lives in London.” You meant Sarah. But that description also fits thousands of other people. Your words didn't uniquely capture Sarah. The number axioms have the same problem: they were meant to describe the normal numbers, but they also fit the impostor systems.

The crucial part is what “provable” actually means. In logic, to prove something from your rules means it has to come out true in every system those rules fit, not just the one you had in mind. That's the catch. If a statement is true in the normal numbers but false in even one impostor system, then it cannot be proved, because proof demands agreement across all of them.

Gödel's sentence is exactly such a statement. In the normal numbers, it's true. But in some of the impostor systems, it's false. The systems disagree about it. And because they disagree, no proof can exist. That disagreement isn't a bug in Gödel's argument; it's the reason the sentence is unprovable in the first place. The split between the normal numbers and the impostors is precisely what lets the sentence dodge proof forever.

Gödel’s second theorem twisted the knife. It showed that no set of mathematical rules can prove, using only its own rules, that it is free of contradictions. If you want to check whether your system is trustworthy, you always need a bigger system to do the checking, and that bigger system inherits the same limitation. Turtles all the way down.

This is not mysticism, nor is it a claim about consciousness or creativity. It is a precise result about rule-based systems, the kind of systems that all software, including AI, is built from. That is what makes it relevant today.

The failed dream that built the computer

Hilbert had asked for one more thing, and Gödel’s paper left it wounded rather than dead. Alongside completeness and consistency, he wanted decidability, a mechanical method that could, in a finite number of steps, determine whether any mathematical statement follows from the rules. No genius required: crank the handle and read the verdict.

In 1936, a 23-year-old Cambridge fellow named Alan Turing killed that too. To prove that no mechanical method could exist, he first had to pin down what “mechanical method” meant, which nobody had done before. His answer was an imaginary device, a paper tape and a head that moves along it, reading and writing symbols according to a fixed table of rules. Anything a human clerk could work out by rote, this device could also work out.

Then he showed the device has a blind spot of its own. Imagine a fortune-teller who is never wrong, and a stubborn customer determined to sabotage every forecast. “You will leave by the door.” He climbs out the window. “You will take the window.” He strolls out the door. She is not bad at her job. The job is impossible because her prediction feeds back into the very behaviour it is trying to predict.

Turing turned that scene into code. The checker plays the fortune-teller. It is a program whose job is to read any other program the way you might read a recipe, then predict its fate. Either “this one finishes” or “this one grinds on forever.”

The saboteur plays the stubborn customer. It is a short program with a copy of the checker tucked inside, plus one standing rule. Ask the checker what I am predicted to do, then do the opposite. If the prediction is that it finishes, it deliberately loops forever. If the prediction is that it runs forever, it stops dead.

So what does the checker predict for the saboteur? “Finishes” is wrong, because the saboteur hears that and loops. “Grinds on forever” is wrong, because the saboteur hears that and stops.

The saboteur is assembled entirely from the checker’s own parts, which makes it inevitable rather than a fluke. Build a perfect checker, and you have, in the same afternoon, built the plans for the thing that breaks it. A perfect checker is therefore a contradiction in terms.

Two boundaries stop this result from proving too much. First, it concerns the universal case. Turing showed that no single checker can deliver a correct verdict on every program. Any particular program may still be provably fine, and many are. Static analysers and type checkers pass useful judgement on ordinary code all day, and whole operating system kernels have been formally verified.

Second, the proof needs unbounded memory. A physical computer is a finite-state machine, so in principle its fate could be settled by enumerating its states. In practice, the state count for any interesting program dwarfs the number of atoms in the observable universe, which turns the question from impossible into unaffordable. Keep that distinction in hand, because it returns when the subject is AI safety. What cannot exist at any price is the checker that is never wrong about anything. That is the halting problem, Gödel’s self-referential sentence rebuilt from machinery, a machine forced to ask a question about itself.

To show what machines cannot do, Turing had to invent the machine. His imaginary device is the theoretical blueprint of the general-purpose computer, a single machine that can run any program you feed it as data. Nine years later, John von Neumann, who knew Turing’s paper well and admired it, wrote the First Draft of a Report on the EDVAC, which is, in logical terms, Turing’s universal machine rendered in vacuum tubes. Essentially every computer built since follows that design. The laptop on your desk and the datacentre GPU training the next frontier model are, once the engineering is stripped away, the same device from a 1936 logic paper.

Gödel himself thought Turing had done him a favour. It was Turing’s definition of a mechanical procedure, he wrote, that made a “precise and unquestionably adequate” general version of his own theorems possible. And the machine in the theorem turned out to be the thing every business now runs on, born as a stepping stone in a proof about what it could never do.

The Gödel machine and the guarantee that vanished

Once you have a machine that can run any program, another question is whether it can improve itself autonomously. In 2003, the German computer scientist Jürgen Schmidhuber proposed a thought experiment he called the Gödel machine. It was an AI agent designed to rewrite its own code, with one ironclad constraint. It would change itself only when it could first prove, with mathematical certainty, that the change would make it better. Not “test and see.” Prove it, as you would a theorem, before running the new version. No proof, no rewrite.

But nobody ever built one. To prove that a code change will improve future performance, you need to search through all possible mathematical arguments that could establish that fact. For any interesting problem, the number of candidate proofs is so astronomically large that the search would take longer than any improvement could ever be worth. It is the computational equivalent of insisting on a signed certificate from every possible future before crossing the road. The Gödel machine was provably optimal but completely impractical.

In May 2025, the Japanese AI lab Sakana released a system called the Darwin Gödel Machine. It retained the self-improvement loop but dropped the proof requirement. Instead of proving that a code change would help, the Darwin Gödel Machine proposes changes using a large language model, tests them against SWE-bench (a benchmark that scores whether an AI can fix real bugs in real software), and keeps what works. The name still invokes Gödel, but the mechanism is Darwinian. Natural selection, not formal proof. Fitness measured by benchmark scores, not mathematical certainty.

Judged purely on the scoreboard, it delivered. The system improved its SWE-bench score from 20% to 50% through autonomous self-modification. It developed emergent behaviours, such as patch validation and error memory, that no one had designed.

Schmidhuber’s original machine, though, had exactly one property that made it safe by construction: the proof. Every modification was guaranteed to be an improvement before it ran. The Darwin Gödel Machine replaced that guarantee with something weaker: passing the benchmarks. The difference between “provably better” and “scored higher on the benchmark test” is the difference between an aircraft type certified against a spec and one that simply hasn’t crashed yet.

This is, compressed into one system’s evolution, the trajectory of AI safety. The formal guarantee was too expensive, so the industry replaced it with empirical validation. “Self-improving” went from a mathematical statement about proof-carrying code to a softer description of an agent that rewrites itself and checks whether the benchmarks improve. Gödel was gone.

Things mathematics cannot learn

Some questions in the mathematics of machine learning are unanswerable. In 2019, Shai Ben-David and colleagues published a paper in Nature Machine Intelligence under the understated but devastating title “Learnability can be undecidable.” They took a straightforward question, “given this type of problem, can a machine learn to solve it?”, and proved that the deepest rules of mathematics cannot always settle it. The answer is neither yes nor no. It is silence.

The word “learnable” has an exact meaning here. A machine studies a sample and produces a rule, which it then applies to data it has never seen. That is the basis of pretty much every model we use. For any given type of problem, learning theory asks whether some sample size can guarantee the rule will work. If such a guarantee exists, the problem is learnable. If none does, it is not. A simple question with two answers, and every type of problem is supposed to get one.

Ben-David's team asked it about a mundane task: choosing which adverts to show a website's visitors from a sample of past ones. Learnable or not?

The answer, in their framework, depends on how many different kinds of visitor there could possibly be. That pool is not the eight billion people alive today. The model reduces each visitor to a profile of measurements, and measurements can vary without limit. A visitor might linger on a page for three seconds, or for a shade over three, and between any two profiles there is always room for a third. The pool of possibilities has no end.

That arrangement, a finite sample making predictions about an endless pool, sits under almost every AI product on the market, advertising included. Training data is always finite. The world a system is released into is not. Ben-David's question is whether that leap can ever come with a guarantee, and everything turns on how big the infinity is.

That sounds like it must have an answer. It does not. Georg Cantor proved in the 1870s that infinity comes in sizes. The whole numbers form one infinity. The points on a line form a strictly bigger one, and the proof is surprisingly simple. Try to pair every whole number with a point on the line, and Cantor showed you will always miss some, no matter how clever the pairing. Both collections are endless, yet one permanently outruns the other. The continuum hypothesis asks a follow-up so obvious it would occur to a child. Is there any size of infinity between those two?

Gödel proved in 1940 that the standard rules of mathematics can never prove the answer is yes. Paul Cohen proved in 1963, using a technique he invented for the purpose, that they can never prove it is no. His proof is dense, and we are already deep enough in theoretical mathematics.

The upshot is that no cleverer generation is coming to settle this one. The rules of mathematics contain no answer. Take every rule of arithmetic and logic we have and follow them as far as they go, in whatever direction you like. You will never reach yes, and you will never reach no. The question is open in both directions, permanently.

Whether the advertising problem is learnable depends on the size of that infinity. The size of that infinity is a question mathematics cannot answer. So whether the advertising problem is learnable is also a question mathematics cannot answer. The strange silence at the very bottom of mathematics travels up the chain and surfaces as a question about showing adverts to shoppers.

Two objections surfaced almost as soon as other mathematicians looked hard at the paper.

The first is about what counts as a learner. A learner here is just the rule a system follows to turn examples into predictions. Ben-David's framework allows that rule to be any mathematical function whatsoever, including ones no computer could ever evaluate. In the cases where the undecidability appears, the rule doing the learning is exactly one of those uncomputable phantoms. It leans on a particular way of lining up all the real numbers, an object the axioms promise exists but give no recipe for building. The rule exists on paper, but no program could carry it out.

The second follows from the first. Insist that a learner be a real algorithm, something that runs on an actual machine, and the whole paradox drains away. Ben-David's own team showed this the following year: once the learner has to be a working program, learnability becomes an ordinary question with an ordinary answer. The silence at the bottom of mathematics never reaches the code. It stays out in the realm of functions nobody could ever build.

So no deployed system is endangered, but the result is no sleight of hand either. Something real did break. It just wasn't a product. What broke was the promise that learning theory could sort every problem into neat buckets: learnable or not. Ben-David found a problem it can never sort. Not for want of better mathematicians, a problem where the sorting itself is impossible. The University of Waterloo, Ben-David's institution, described the result as “important and almost troubling.”

The neural network that exists and cannot be built

A complementary result, published in 2022, is narrower and stranger. Training a neural network is, at bottom, an exercise in trial and error. You show the network examples, measure how wrong its answers are, then nudge its millions of internal settings to make them slightly less wrong. Repeat this billions of times, and often the network converges on something remarkably good. The Cambridge mathematician Matthew Colbrook and colleagues showed in a 2022 paper in PNAS that this process has a hard boundary nobody expected.

The boundary appeared in medical imaging. An MRI scanner does not take a photograph. It collects measurements to keep patients in the incredibly claustrophobic machine for minutes rather than hours, so it collects far fewer than a complete image requires. Software has to rebuild the full picture from the partial data. Neural networks became the favoured tool for this reconstruction because they perform it faster and more accurately than older mathematical methods.

Then researchers started probing the results and found something unnerving. Nudge the input slightly, with a trace of noise or a small movement by the patient, and the output could change out of all proportion. Sometimes the rebuilt scan came back looking perfect but wrong, showing details that were never in the body. Worse, it is a failure that does not look like a failure. A blurry image warns you. A crisp fabricated one does not.

A demonstration of capability may show nothing is amiss, not because anyone is cheating. A demo runs the network on typical inputs, the kind it was trained on, where it genuinely performs well. The failures live in the near-misses, a typical input plus a whisker of noise. Near-misses are endless, and a demo can show only a handful of scenarios. The demo is honest, and the danger lies exactly where it cannot be seen.

The natural presumptive diagnosis is undertraining. Feed it more scans and buy a bigger model, and the wobble will surely iron itself out. That hope is what Colbrook’s theorem takes off the table. What the paper proved has two halves. First, for certain reconstruction problems, a network that is both accurate and stable exists. Somewhere in the space of all possible settings sits a configuration immune to the wobble. Second, no training procedure can find it. Not the ones we have. Not any. None at all, ever.

The second half is what kills the more-data hope. A training procedure is itself a program, a step-by-step recipe running on the machine Turing described, and the proof covers every recipe there could ever be. More data does not change that. Data is what you feed a recipe, and the theorem is about the recipes. It is like knowing a winning lottery ticket is in a barrel whilst simultaneously holding a proof that no way of drawing from the barrel will ever pull it out. The ticket is real. The searching is futile.

As Colbrook put it, the paradox Turing and Gödel identified has now been “brought forward into the world of AI”, and for certain problems the required algorithms simply cannot exist.

For the overwhelming majority of real-world problems, training works. But there is no general way to tell in advance which problems will defeat us, and the assumption that enough data and compute will always get us over the line is, in certain corners of the problem space, provably false.

The machine you cannot contain

The most provocative extension of Gödel’s legacy into AI concerns a question that sounds simple. Can we guarantee that a sufficiently powerful AI will not cause harm?

In 2021, Manuel Alfonseca and colleagues published a paper in the Journal of Artificial Intelligence Research arguing that, for a general-purpose superintelligent system, the answer is provably no. Their argument leans on the halting problem, the impossibility Turing established on his way to inventing the computer. You can check specific programs for specific bugs. What you cannot build is the universal checker, the one that works for any program in any situation.

Alfonseca’s team showed that asking “will this AI harm humans?” is, for a fully general system, the same type of question as asking “will this program halt?” Both require predicting the complete future behaviour of a system from its current state. To guarantee a system will never cause harm, you would need to trace every possible sequence of actions it could take and confirm that none is harmful. That is the halting problem in different clothes, and Turing proved that this class of prediction is impossible to guarantee. You cannot build a general-purpose AI safety monitor for the same reason you cannot build a general-purpose program-behaviour predictor.

The scope of that impossibility deserves the same care as Turing’s original. What cannot exist is the universal judge, one procedure that delivers a correct safety verdict for every possible system and every possible input. Specific systems doing bounded work in bounded settings can be certified, and safety engineering does exactly this, one component at a time. Wrapping a monitor around an agent catches genuine failures and is worth every penny.

But each monitor is itself a program with the same blind spot, so layered checks buy coverage, but never closure. And since a physical machine has finite memory, its behaviour is in principle a finite question, merely one whose size defeats any conceivable budget. For a bounded system, the wall is cost. For the unboundedly general system Alfonseca’s team modelled, the wall is mathematics.

The authors went further, showing that we may not even be able to recognise when a superintelligent system has arrived, because deciding whether a machine is smarter than a human falls into the same class of unanswerable questions. The argument is grounded rather than speculative, though it assumes a generality no AI system possesses today. Nothing currently available can handle any possible input the way a true Turing machine can.

What it establishes still matters, though. Certain safety guarantees are not engineering problems awaiting a sufficiently clever solution. They are mathematical impossibilities, like trying to square the circle or list every real number between 0 and 1. The safety community can build better guardrails and better kill switches. What it cannot build, given the computational framework we share, is a system that certifies another system as unconditionally safe.

What Gödel would recognise

These four threads are cousins rather than corollaries of a single theorem. Incompleteness limits what rule systems can prove about themselves. The halting problem limits what programs can decide about programs. Colbrook’s result limits what algorithms can find. And the Darwin Gödel Machine story records a design retreat, a choice rather than a law of nature. What joins them is that in each case a formal guarantee is either unavailable in principle or unaffordable in practice, and the work must proceed anyway on empirical confidence.

None of this is softened by the fact that a neural network feels organic rather than rule-like. A model’s weights are numbers, and its training is arithmetic, all of it running on von Neumann’s realisation of Turing’s imaginary device. AI is not adjacent to this mathematics. AI is made of it.

Nor is any of it an argument that machines cannot think. These limits bind every formal reasoner, and on the standard understanding of physics, that includes the three pounds of wet machinery reading this sentence. You cannot prove your own consistency either. Humans invent, reason, and get things done inside exactly the same boundaries; evidence that the boundaries are no bar to intelligence. Gödel constrains what can be guaranteed about a mind, human or artificial. He says nothing about what a mind can do.

A fair objection is that the AI industry never promised mathematical proof, and empirical validation is how almost everything gets built. Aircraft are tested, drugs are trialled, and nobody demands a theorem before boarding a plane. All true, and for the overwhelming majority of applications, the limits in this piece never come up. Your chatbot will not encounter the continuum hypothesis when summarising a report.

But the objection undersells what the mathematics settles. Testing regimes for aircraft rest on decades of physics that says how materials behave between the test points. For a system that rewrites itself, or one deployed against inputs no test set anticipated, no such interpolating theory exists, and the results above show parts of it never will.

That turns safety from a verification problem into a pricing problem. If certainty is permanently off the table, the question becomes how much assurance a given deployment needs, what it costs to obtain, and who bears the residual risk when that assurance runs out. At present, the industry is a long way from answering those questions explicitly. “Passed the evals” blurs into “proven safe”.

Einstein’s eccentric walking companion saw the underlying structure before anyone else. Formal systems cannot fully certify themselves. That was a logician’s problem in 1931. It became an engineer’s problem when Turing turned the proof into a machine. It is now a commercial problem because the machines carry trillions of dollars in expectation, and expectation is reaching for guarantees that the mathematics declines to issue. Guarantees generated by systems that cannot check themselves any more than Gödel’s own warped internal logic could. He died trapped inside it.

I am a partner in Better than Good. We help smaller companies build tools and processes using machine learning and artificial intelligence that make lasting improvements to their operations. Talk to us today: https://betterthangood.xyz/#contact

Ever wondered how the BBC provides its radio services nationwide in the UK?

In the 1990s the principle was line-of-sight relay. Microwave signals (broadcast distribution sat in the SHF bands, roughly 2–15 GHz) travel in straight lines and formed a chain: a series of relay stations on hilltops or tall towers, each one spaced just inside the horizon of the next.

Truleigh Hill aerials

Each link in the chain did the same job. A microwave dish received the incoming beam (these looked like large white drums and are still seen here and there), the station amplified and reconditioned the signal, broadcasting on FM to the surrounding area with another microwave dish retransmitting it on a slightly different frequency to the next station down the line thus achieving national coverage from a linked network.

I’ve always been fascinated by all types of technology from AI to the radio spectrum. In the mid 90s I was just “learning the trade” as a radio engineer. I was tight with a guy who used to build our FM transmitters for us. Somehow he had managed in this pre internet age to get hold of all the frequencies for the BBC repeater network. He also was our indirect source for lots of other useful things like the mythical Fire Brigade or FB keys that gave us access to most tower block rooftops in London.

We’d been discussing the practicality of hijacking a BBC national network by drowning out the official incoming microwave signal with a more powerful one of our own, which then, by nature of the network design, would be passed on down the chain. We thought we could do this if we got close to one of the big repeaters and blasted enough power on the right frequency.

Which is how I found myself one drizzly bank holiday Sunday sat on the roof of a van parked on Truleigh Hill in Sussex pointing a microwave transmitter at the nearby mast. I can’t precisely recall what the source was but it may well have been a DAT of a classic Dreamscape mixtape.

It worked like a charm, confirmed when we rang a mate in Preston and asked him to tune to Radio 3.

Which is how in the mid 1990s Radio 3 listeners found their quiet bank holiday Sunday classical listening suddenly interrupted by half an hour of unadulterated jungle music, which I’m sure caused more than one post-prandial sherry to be spilled.

I am a partner in Better than Good. We help smaller companies build tools and processes using machine learning and artificial intelligence that make lasting improvements to their operations. Talk to us today: https://betterthangood.xyz/#contact

You live within a system you never signed a contract with. Every day, you make thousands of micro-decisions about how to behave, mostly without conscious thought. You pay an invoice on time, even if the supplier would never discover that you didn’t. You refuse to do business with someone who stiffed their last three partners, and you’d think twice about a colleague who didn’t. Nobody wrote these rules down. You absorbed them the way you absorbed grammar through exposure and correction.

A March 2026 paper from the Knight First Amendment Institute by Gillian Hadfield, Rakshit Trivedi, and Dylan Hadfield-Menell argues that this invisible social choreography is the core mechanism of democracy, not just an adornment. Furthermore, AI agents, such as those currently being developed to run businesses and manage supply chains, will undermine that mechanism unless they learn this dance too.

A retro futuristic robot and human dancing together

Democracy Is a Verb, not a document

The paper begins by challenging a common assumption. Most view democracy as a collection of documents, institutions, constitutions, elections, and courts. The authors contend that this is roughly akin to describing a marriage solely through its wedding vows. While the vows matter, the true essence of a marriage lies in the thousands of everyday acts of compromise and occasional irritation that sustain cooperation over decades.

Hadfield, Trivedi, and Hadfield-Menell utilise a theoretical framework called “normative social order” to make this precise. In their model, a society’s actual norms are the product of an interactive system. People don’t follow rules because they are written down; they follow them because they observe others doing so and see how violations are punished. Punishments don’t need to be severe, just a disapproving look, a refusal to do business, or a sarcastic comment at a dinner party. These micro-sanctions generate the gravitational field that keeps behaviour in orbit.

This is where the paper borrows a term from evolutionary theory, “dancing landscapes.” The metaphor, from Stuart Kauffman’s work on complex adaptive systems, describes environments where multiple independent agents are constantly adjusting to each other’s behaviour. There is no central choreographer; the dance arises from the dancers' interactions themselves.

What makes a norm sticky

The framework introduces a concept called a “classification institution,” which is any shared mechanism a group employs to decide which behaviours are punished and which are not. In small groups, this classification is entirely implicit, and you know what the group considers acceptable or unacceptable. Acceptability is judged by seeing who gets mocked and who gets praised. The Ju/’hoansi Bushmen, as anthropologist Polly Wiessner describes, regulate behaviour through evening conversations. Gossip and teasing around the fireside serve the same purpose as courtrooms and HR departments in modern societies.

As societies grow more complex, implicit classification cannot scale because the diversity of people and situations exceeds the reach of any informal consensus process. This creates a need for identifiable classification institutions; entities that can resolve ambiguity when community members disagree about acceptability. Courts, regulatory bodies, trade associations, and professional standards boards all serve this purpose in modern societies.

The paper argues that for these institutions to be effective, they need attributes that closely match what legal philosophers have long called “the rule of law,” namely stability, clarity, generality, and neutrality. The twist is that Hadfield and her co-authors do not derive these attributes from abstract principles. Instead, they derive them from game theory. An institution with those attributes is one around which independent actors can reliably coordinate, and coordination is what sustains the entire system.

Enter Adam Smith’s imaginary friend

The paper revisits Adam Smith’s “impartial spectator” from The Theory of Moral Sentiments and uses it as a model for how AI agents could participate in democratic societies without causing harm. Smith argued that moral reasoning works because each of us carries a mental image of a neutral observer—an internal referee—who judges our behaviour against community standards. You do not avoid bribery because you have memorised a specific anti-corruption law; you avoid it because your internal impartial spectator would wince.

This is the cognitive capacity that Hadfield, Trivedi, and Hadfield-Menell call “normative competence.” It goes beyond simply knowing the rules. It involves the ability to interpret a constantly changing normative environment, anticipate how your community will respond to specific actions, and adjust your behaviour accordingly. The key point is that it also requires predicting how the rules themselves will change, since in any living democracy, they change constantly. Yesterday, you didn’t need to worry about data privacy in your marketing. Today, GDPR and its equivalents are everywhere, and community expectations have shifted.

Why this matters if you’ve never read game theory

If AI agents were merely chatbots answering questions, none of this would be urgent. But the organisations developing these systems are designing agents to operate autonomously in the world for days or weeks at a time, making real decisions with tangible consequences. Mustafa Suleyman, who co-founded DeepMind and now leads AI at Microsoft, proposed a “Modern Turing Test” that perfectly highlights the problem. Instead of testing whether a machine can imitate human conversation, his test asks whether an AI agent can turn $100,000 into $1 million on a retail platform within a few months.

Consider what that entails. The agent would need to research markets, design products, hire contractors, negotiate with manufacturers (possibly abroad), set pricing strategies, handle customer complaints, comply with regulatory requirements, manage logistics and warehousing, and organise payment systems. At each stage, it would be making decisions within the framework of democratic norms. What labour practices does the manufacturer adopt, and is the marketing misleading? Should the agent accept an offer from a local politician to disadvantage a competitor? Should it take a bribe from a supplier in the form of a crypto transfer?

These decisions are made by humans daily, and most of the time the answers seem obvious because humans have spent a lifetime absorbing the normative environment. The answers are not codified in a rulebook. They emerge from that invisible dance of observation and adjustment. An AI agent, no matter how well trained on legal texts and ethical principles, does not possess this “dance literacy”.

The incompleteness problem

Current approaches to AI alignment mainly assume that the right rules can be built into the system. Constitutional AI, the method used by Anthropic, fine-tunes models using a written constitution of principles. Other efforts collect “democratic inputs” through surveys and citizen assemblies. While the paper recognises these as valuable, it argues that they miss the core challenge. The issue is incompleteness: you cannot write instructions detailed enough to cover every possible situation an autonomous agent might face, because both situations and norms evolve.

Economists have understood this for decades in the context of human contracts. Every employment contract, partnership agreement, and supply chain arrangement is inherently incomplete. You can’t foresee every scenario, and when gaps appear between people, they fill them using shared norms, professional customs, and legal precedents, all of which are dynamic and partly implicit. An AI that stops learning norms at training time is like a new hire who memorised the employee handbook on their first day and then ignored all social cues from colleagues for the next ten years.

What the paper proposes

The technical agenda has two main parts. The first focuses on “normative competence,” embedded in individual AI agents. This is formalised through Bayesian adaptive decision processes, which in plain language means that the agent maintains beliefs about the normative environment, updates those beliefs based on feedback (including punishment signals such as losing a contract or receiving a complaint), and makes decisions that account for uncertainty about what is acceptable. Crucially, this happens at inference time, in real-time, based on live context, rather than being pre-programmed into the model during training.

The second part involves creating new institutions and digital classification systems that can serve roles similar to those of courts, regulatory bodies, and professional norms for humans. The paper introduces “Model Specification Institutions” (MSIs), which would be democratically formed bodies (such as citizen assemblies, expert panels, digital juries). These bodies would establish shared standards, training datasets of acceptable and unacceptable behaviours, and real-time APIs that agents can consult in ambiguous situations. This does not mean AI companies should define their own rules; rather, it is calling for democratic communities to develop new infrastructure that AI agents can understand and respond to.

The paper also proposes adapting existing infrastructure—such as certificate authorities, which currently verify website identities—to certify that an AI has been trained to adhere to specific behavioural standards. Reputation networks, such as seller ratings on Amazon or Uber driver scores, could track AI behaviour over time and impose consequences on agents that repeatedly violate community norms.

Perhaps the most provocative argument concerns enforcement. Democracy doesn’t endure solely because governments enforce every rule from above. It survives because ordinary people enforce norms from below. You refuse to do business with a supplier who cheats. You complain when a company misleads you and vote against politicians who ignore court orders (well, mostly). This distributed enforcement, which the paper calls “third-party punishment,” is the engine that keeps the entire system functioning.

If AI agents replace humans in millions of daily transactions and those agents do not participate in this enforcement, the incentive structure collapses entirely. Imagine a world where most business transactions are handled by AI agents that don’t care whether a trading partner has been found guilty of fraud, because the agents were not programmed to check for or respond to that information. The paper argues that AI agents will need to participate in distributed enforcement, refusing to transact with entities that violate community norms, just as humans do. Otherwise, the shift to agentic AI will quietly erode the social infrastructure on which democracies depend.

What this means for you

If you run a business, this paper should change how you think about deploying AI agents. The issue is not whether your agent can follow a rulebook. The question is whether it can read the room. Can it tell the difference between a legitimate business request and an attempt to corrupt a procurement process? Can it adapt its behaviour when community standards shift, without waiting for you to update its instructions? Can it recognise when a trading partner’s behaviour should disqualify them from further transactions?

If you are a citizen who votes, pays taxes, and occasionally debates politics, this paper describes the infrastructure of your daily life in terms you may not have previously considered. The norms you enforce through your micro-decisions, who you buy from, who you work with, and how you respond to rule-breaking are the operating system of democracy. What Hadfield, Trivedi, and Hadfield-Menell are asking is what happens to that operating system when a large fraction of those daily decisions are made by software that cannot read the social signals the system depends on.

The answer, if you follow the paper’s logic, is that we need to build new democratic institutions at the speed democracy demands, before the agents outrun the infrastructure. The alternative is a world where the formal structures of democracy persist, but the lived experience of it, the texture of mutual accountability in ordinary interactions, fades, like a coral reef whose skeleton remains after the living organisms have gone.

I am a partner in Better than Good. We help smaller companies build tools and processes using machine learning and artificial intelligence that make lasting improvements to their operations. Talk to us today: https://betterthangood.xyz/#contact

The technology industry has spent the past three years debating artificial intelligence with the zeal of medieval theologians disputing angels on pinheads. Boardrooms have AI strategies, and governments have AI safety frameworks. LinkedIn has AI thought leaders, which is arguably the strongest case yet for existential risk. But somewhere beneath the acronym and the data centres in space, one central question remains unanswered. What, exactly, is intelligence?

We are developing systems we call intelligent, regulating systems we call intelligent, and worrying about systems we call intelligent, without a shared scientific consensus on what that term means when applied to humans, let alone machines. That is, to say the least, a problem.

The Monolithic I

The indefinable word

Ask a psychologist what intelligence is, and you’ll step into the epicentre of a fierce debate that has lasted for more than a century. The oldest and most statistically reliable answer comes from Charles Spearman, who in 1904 observed that people who did well on one kind of cognitive test also tended to do well on others. He called this underlying factor g, or general intelligence. The g factor is among the most replicated findings in psychology. It predicts academic performance, job performance, income, health outcomes, and even longevity, with a consistency that makes most social-science results look like coin flips.

And yet g tells you almost nothing about what intelligence really is. It is a statistical regularity, not a mechanism. Saying someone has high g is a bit like saying a car is fast. The measurement works, but the explanation is missing.

Howard Gardner tried to blow the whole thing up in 1983 with his theory of multiple intelligences, arguing that intelligence is not one thing but at least eight distinct varieties, from linguistic and logical-mathematical to musical, bodily-kinaesthetic, spatial, interpersonal, intrapersonal, and naturalistic. Teachers lapped this up. It confirmed their intuition that the kid who struggles with algebra but plays the cello like a prodigy is smart in ways traditional testing misses.

The problem is that decades of factor analysis have stubbornly refused to confirm Gardner’s categories as truly independent. Musical ability and spatial reasoning correlate, as do linguistic and interpersonal skills, and, in fact, everything correlates, which is more or less Spearman’s original point. Multiple intelligences is a useful pedagogical framework but a weak empirical theory, which is a polite way of saying it works better in classrooms than in laboratories.

Then there is François Chollet’s definition, which originates from the AI community and is arguably the most rigorous recent attempt to clarify the concept. In his 2019 paper “On the Measure of Intelligence,” Chollet defined intelligence not as the ability to perform any specific task, but as the efficiency with which a system acquires new skills, especially when confronted with tasks it has never encountered before.

This led him to develop the Abstraction and Reasoning Corpus (ARC), a benchmark of visual puzzles designed specifically to assess this ability. Humans usually solve most ARC tasks within minutes. The latest version, ARC-AGI-3, published in March 2026, makes the gap even clearer by placing agents in interactive environments where they must infer goals and plan action sequences without explicit instructions. Humans solve 100% of these tasks. At the time of the paper’s publication, frontier AI systems scored less than 1%. This gap cannot be closed simply by better prompt engineering. Whether this means current AI lacks intelligence, or only a particular kind of adaptive reasoning that humans excel at remains an open question.

The definitional problem is not just academic. Every claim about AI being intelligent, not intelligent, or dangerously intelligent depends on an implicit definition. Call a model intelligent, and you usually mean it produces outputs that would require human intelligence. Deny it, and you mean it lacks the comprehension, consciousness, or intentionality you consider necessary for genuine intelligence. Both claims are unfalsifiable without a shared definition, which explains why the debate generates much heat but little clarity.

What neuroscience understands (less than many think)

If psychology cannot agree on what intelligence is, perhaps neuroscience can explain how it works. The short answer is that it can, at least partly, though large gaps remain. We know a great deal about the brain’s individual components. We can map neural circuits, measure neurotransmitter activity, image blood-oxygen levels as proxies for activity, and trace connectivity patterns across cortical regions.

We know that the prefrontal cortex plays a key role in planning and abstract reasoning, that the hippocampus is central to memory consolidation, and that the cerebellum (once thought to be merely a motor coordination device) participates in cognitive processes that are not yet fully understood. We also observe that, within a species, larger brains tend to correlate weakly with cognitive ability, and that connection density and efficiency matter more than overall volume.

What we cannot do is explain how any of this produces thought. We have a parts list and some wiring diagrams, but no operating manual. The situation is roughly equivalent to having an inventory of components for a Boeing 787 without understanding aerodynamics. You could describe the wings, the engines, the control surfaces, and the hydraulic systems, and still have no theoretical framework for why the thing flies.

Two research programmes have made the most ambitious attempts to close this gap, and both illustrate how far there is to go.

Predictive processing

The first concept is predictive processing, most closely associated with philosophers Andy Clark and Karl Friston. They suggest that the brain is not a passive receiver of sensory data. It functions as a prediction machine that constantly builds models of what it expects to perceive, then updates them when reality differs from those expectations. Perception, in this view, is not bottom-up (data in, interpretation out) but top-down (expectation generated, error signal compared, model revised). You do not see the world as it is. You see your best guess about the world, corrected at the edges by incoming data.

Friston formalised this idea in the free energy principle, a mathematical framework suggesting that all adaptive behaviour can be understood as the minimisation of “free energy,” which roughly measures the gap between an organism’s internal model and the sensory evidence it receives. The framework is mathematically coherent and broadly applicable, and that is precisely the problem. If every possible behaviour of any living system can be reinterpreted as free-energy minimisation, then the theory rules nothing out, raising serious questions about its scientific credibility.

Integrated information theory

The second programme is Integrated Information Theory (IIT), developed by the neuroscientist Giulio Tononi. IIT takes on the even harder problem of consciousness rather than intelligence per se, but the two are tangled enough that progress on one would likely tell us something about the other. The theory proposes that consciousness corresponds to a quantity called phi (Φ), which measures the amount of information a system generates “above and beyond” its individual parts. A system with high phi is one whose behaviour cannot be reduced to its components acting independently. The whole, in a precise mathematical sense, is more than the sum of its parts.

IIT makes some bold predictions. It implies that consciousness is a property of a system’s physical structure, not its function. A digital simulation of a brain that runs the same computations on different hardware might have zero consciousness under IIT, even if it behaves identically to the original. In practice, calculating phi for any system more complex than a handful of nodes is computationally intractable, which limits the theory’s practical utility. You can define consciousness precisely and still be unable to measure it in any real system, which is a bit like having a perfect recipe for a cake you can never bake.

Neither predictive processing nor IIT amounts to a theory of intelligence in the way that general relativity is a theory of gravity. They are frameworks, useful and generative but incomplete, and they throw light on aspects of cognition without explaining the whole. And the gap between “aspects” and “the whole” may be permanent, for reasons we will get to.

Why large language models work

If the science of biological intelligence is patchy, the science of artificial intelligence is in an even stranger position. The engineering works spectacularly well, while the theory lags behind, like a civil engineer who builds bridges that hold up beautifully but cannot fully explain the physics of load distribution.

We understand the mechanics of large language models in fine detail. A transformer architecture processes sequences of tokens through layers of attention mechanisms, and during training the model adjusts billions of parameters to minimise the error between its predicted next token and the actual one. Scaling laws, first characterised by Jared Kaplan and colleagues at OpenAI in 2020, describe a remarkably smooth power-law relationship between compute, dataset size, model parameters, and performance.

These are genuine scientific results. They let engineers predict, with useful accuracy, how a model of a given size trained on a certain amount of data will perform on standard benchmarks. What they do not explain is why training a system to predict the next word in a sequence produces behaviour that appears like reasoning, planning, analogy, and (occasionally) creativity.

The most provocative explanation comes from the compression hypothesis, most forcefully articulated by Ilya Sutskever, then of OpenAI. The argument roughly runs like this. Predicting the next token accurately requires modelling the process that generated the text, and that process is human cognition. To predict well, you must compress the structure of human thought into your parameters. Compression, in this view, is not merely correlated with intelligence but constitutive of it. A model that achieves better compression has, in a meaningful sense, come to understand the world better.

This is philosophically interesting and empirically suggestive, but it is not a complete theory. It does not explain why certain abilities appear discontinuously as models scale. Small models cannot perform multi-step arithmetic. Larger models can suddenly, without anyone having specifically trained them for it. These “emergent capabilities” are predicted by no current theory and explained by no current framework. They simply happen, and then engineers and researchers argue about what they mean.

Mechanistic interpretability, an active research programme at Anthropic among others, is perhaps the most promising attempt to open the black box. The work identifies specific circuits within trained models that correspond to identifiable computations, so that one cluster of neurons detects sentiment and another tracks syntactic dependencies. The results are revealing, but they are roughly at the stage where neuroscience was when it discovered that specific brain regions correspond to specific functions. Knowing where a computation happens is useful. Knowing why the system learned to do it, and why it generalises beyond the patterns in the training data, is the harder question.

The “stochastic parrots” critique, most prominently advanced by Emily Bender, Timnit Gebru, and colleagues in 2021, argued that LLMs are only sophisticated statistical mimics. Noam Chomsky has made similar arguments, insisting that next-token prediction cannot amount to genuine linguistic comprehension. Melanie Mitchell has taken a more cautious position, arguing that current AI systems lack the conceptual abstraction and analogy-making she sees as central to intelligence, while leaving open the possibility that future architectures might achieve it.

The honest answer is that nobody knows who is right. The stochastic-parrot position seemed more defensible in 2021 than it does in 2026, because the systems have kept improving in ways a “mere statistical mimic” would not obviously be expected to. But the lack of a theory means that “would not obviously be expected to” is carrying more weight in that sentence than it should. We do not have the theoretical tools to distinguish genuine comprehension from a sufficiently convincing imitation of it, and those tools are not arriving quickly.

Can intelligence be formalised at all?

Here we reach the question beneath the question, and the answer is uncomfortable for anyone who prefers their science tidy. A hidden hope in much AI research is that intelligence resembles thermodynamics: messy and chaotic at the micro level, but governed by clean, discoverable laws at the macro level. Individual gas molecules move unpredictably, yet aggregate behaviour follows the ideal gas law with almost miraculous precision. Perhaps intelligence works the same way, messy at the level of individual neurons or attention heads, but obeying some elegant principle at a higher level of description.

The problem is that thermodynamics works because you can ignore which specific molecule is where. It is far from clear that cognition has this property. The specific structure of a person’s knowledge, the particular history of their experiences, and the exact wiring of their neural connections all seem to matter in ways that resist averaging out. A brain is not a gas. Its macro-behaviour may not separate cleanly from its micro-state, and if it does not, no thermodynamics-style theory is possible.

There is a deeper problem. Any formal theory of intelligence needs to specify what intelligence is for, what problem it solves, and what it optimises. A thermostat optimises temperature, a chess engine optimises board position, and both can be fully described by their objective function. But intelligence seems to be precisely the capacity to redefine what counts as the problem. A human can decide whether to play chess at all, invent a new game, or abandon the entire framing and go for a walk. Formalising that kind of open-ended reframing may require a kind of mathematics that does not exist yet, or it may resist formalisation altogether.

This is where the biology analogy becomes revealing. There is no “theory of organisms” in the same sense that there is a theory of electromagnetism. Biology has a powerful organising framework, evolution by natural selection, along with a vast accumulation of mechanisms, trade-offs, and contingent historical facts. You can explain any feature of an organism after the fact. You cannot derive organisms from first principles. The evolutionary biologist Stephen Jay Gould argued that if you replayed the tape of life from the same starting conditions, you would get a completely different set of organisms. The outcomes are historically contingent, not mathematically necessary.

Intelligence may be the same kind of thing: a product of evolutionary tinkering, cultural accumulation, and developmental contingency that allows useful generalisations but not the kind of closed-form theory that would satisfy a physicist. We may end up knowing intelligence the way we know weather: well enough to make useful short-term predictions, poorly enough that long-range forecasting remains unreliable, and never with the exact analytical solution that would let us derive tomorrow’s clouds from first principles.

Why does any of this matter outside of a philosophy seminar?

The temptation is to treat all of this as an abstract debate, the sort of thing academics argue about while engineers get on with building things that work. That temptation should be resisted, because the theoretical vacuum has practical consequences.

AI safety without a theory of intelligence is navigation without a map. The field depends on assumptions about what future systems can achieve and how those abilities will develop. If we do not understand why current systems perform as well as they do, we cannot predict whether the next generation will improve steadily or make a sudden leap, as we have seen with Anthropic’s Mythos. Scaling laws tell us that larger models perform better, but not what “better” means at scales we have not reached. Will a model 100 times larger than current frontier systems merely write more polished prose, or develop something qualitatively different? Nobody knows, and we lack the framework to reason about the question.

Regulation without definitions is theatre. Governments are drafting AI rules around distinctions (general-purpose versus narrow, high-risk versus low-risk) that depend on a theoretical grasp of intelligence we do not possess. The EU’s AI Act defines a “general-purpose AI model” by compute thresholds that are essentially arbitrary, because no theory links compute to ability in a way that would make any threshold principled. The fault is not the regulators’, but the tools’, which are insufficient.

A business strategy built on vibes is expensive. The corporate world is investing hundreds of billions on the assumption that current trends will continue. Perhaps they will. But the history of technology is full of S-curves that plateau earlier than expected, and the lack of a theory makes it harder than it should be to distinguish genuine improvement from benchmark gaming and evaluation contamination. When a model scores 90% on a medical exam, does that mean it has medical knowledge, or that enough medical-exam text was in the training data? The answer is “it depends what you mean by knowledge,” and we are back to square one.

Where this leaves us

This is not a counsel of despair. Science often advances without complete theories. Medicine cured scurvy centuries before vitamin C was discovered, and engineers built steam engines before thermodynamics was formalised. Practical progress does not require a finished theory, though it helps, especially when the stakes are high enough that mistakes carry consequences beyond a failed experiment.

We are roughly where physics was before Newton. We have observations (scaling laws, emergent abilities, benchmark performance), useful heuristics (more compute and data tend to produce better models), and fragments of theory (compression, mechanistic circuits, predictive processing), but no framework that unifies them and makes novel predictions. The “I” in AI is still a placeholder, a trillion dollars of investment balanced on a word we cannot define.

I am a partner in Better than Good. We help smaller companies build tools and processes using machine learning and artificial intelligence that make lasting improvements to their operations. Talk to us today: https://betterthangood.xyz/#contact

In September 2025, OpenAI published a paper that said something the AI industry already suspected but hadn’t quite articulated. The paper, “Why Language Models Hallucinate”, authored by Adam Tauman Kalai, Ofir Nachum, Santosh Vempala, and Edwin Zhang, didn’t just catalogue the problem. It pointed the finger at the evaluation systems that are supposed to keep models honest and argued that those systems are actively making hallucination worse.

The paper’s central argument is disarmingly simple. Language models hallucinate because we reward them for guessing. The training loops, the benchmarks, the leaderboards that determine which model gets called “best” all operate on a scoring system that treats confident wrong answers and honest uncertainty as equally worthless. Under those rules, the rational strategy for any model is to always take a shot, even when the evidence is thin. And that strategy produces hallucinations.

Researchers have known for years that models tend toward overconfidence. But the OpenAI paper formalised it with mathematical precision and made an argument that goes further than most. The problem is that our entire evaluation infrastructure systematically incentivises the specific failure mode we claim to care most about fixing.

An illustration representing hallucination

The Mechanics of Making Things Up

To understand why the paper matters, it helps to start with what hallucination actually is at a mechanical level.

During pretraining, a language model learns to predict the next token in a sequence. It ingests billions of documents and builds a statistical model of what words tend to follow other words in what contexts. This process is extraordinarily powerful for capturing patterns, grammar, reasoning structures, and factual associations. But it has an inherent limitation that no amount of scale can fully overcome.

Some facts appear in training data frequently enough that the model can learn them reliably. The capital of France, the boiling point of water, the year the Berlin Wall fell. These are high-frequency, well-attested facts that leave strong statistical signals. But other facts appear rarely or only once. The title of a specific researcher’s PhD dissertation. The birthday of a mid-career academic. The precise holdings of a niche legal case from 2019. These “singleton” facts leave weak or ambiguous traces in the training distribution, and no model, regardless of size, can learn them with confidence from pattern matching alone.

The OpenAI paper draws an analogy to supervised learning that makes this intuitive. In any classification task, there’s an irreducible error rate determined by the overlap between classes in the training data. Generative models face an equivalent problem, because some questions simply cannot be answered correctly from the training distribution, and the model’s best option in those cases would be to say “I don’t know.” The paper refers to this as the model’s “singleton rate,” the fraction of facts that appeared only once during training and therefore can’t be reliably recalled.

This matters because it puts a hard floor under hallucination rates regardless of model size or architecture. You can make a model bigger, train it on more data, and give it better reasoning capabilities, and you will reduce hallucinations on well-attested facts. But you will never eliminate them on rare facts, because the statistical signal for those facts is too weak to distinguish from noise. The paper is explicit about this point. Even a 100% accurate model on common facts would still hallucinate on singleton facts, and the only alternative to hallucination on those facts is abstention.

None of this is mysterious. It’s basic statistics applied to language modelling. But what happens next, in the post-training phase, is where things go wrong in a more avoidable way.

The Test-Taking Incentive Problem

After pretraining, models go through rounds of fine-tuning designed to make them more helpful, less harmful, and better at following instructions. This process involves evaluation on benchmarks, and it’s here that the OpenAI paper identifies the core dysfunction.

The paper’s authors compare modern AI benchmarks to multiple-choice tests where leaving an answer blank guarantees zero points. On such tests, the optimal strategy for a test-taker who doesn’t know the answer is to guess. There’s some chance of being right, and no additional penalty for being wrong. Language model benchmarks work on the same principle, and most prominent evaluations, including MMLU-Pro, GPQA, MATH, and others that dominate public leaderboards, use binary scoring where a correct answer scores one point and everything else, whether wrong or abstained, scores zero.

Under this system, a model that says “I don’t know” to a question it’s uncertain about gets exactly the same score as a model that confidently invents an answer. But the model that guesses will occasionally be right by chance, which pushes its aggregate accuracy higher. Since accuracy is the number that appears on leaderboards, in model cards, and in press releases, the models that guess most aggressively tend to look best.

The paper illustrates this with a concrete example from SimpleQA-style metrics. One model showed an error rate of 75% with only 1% abstentions, meaning it almost never admitted uncertainty and was wrong three-quarters of the time when it did answer. Another model abstained 52% of the time and dramatically reduced its error rate. But on a traditional accuracy-only leaderboard, the difference between these two models would look modest, because the metric that gets reported doesn’t distinguish between “wrong” and “chose not to answer.”

This is not an edge case in how benchmarks work. It’s the dominant paradigm. As the paper puts it, the majority of mainstream evaluations reward hallucinatory behaviour. The proposed fix is almost embarrassingly obvious, and borrowed directly from standardised testing. Introduce negative marking for wrong answers, or give partial credit for appropriate expressions of uncertainty, so that honest non-answers score better than confident mistakes.

Looking Inside the Black Box

While OpenAI approached the problem from the evaluation and incentive angle, Anthropic’s interpretability team was working on the same question from the opposite direction, looking at what actually happens inside a model when it decides whether to hallucinate or abstain.

In March 2025, Anthropic published two papers under the banner “Tracing the Thoughts of a Large Language Model” that used a novel “AI microscope” technique to map the computational circuits inside Claude 3.5 Haiku. Among the results was a discovery that runs counter to most people’s intuitions about how hallucination works.

It turns out that Claude’s default behaviour is to refuse to answer. The researchers identified a circuit that is active by default and causes the model to state that it has insufficient information to respond to any given question. This “I don’t know” circuit fires every time Claude receives a query, regardless of the topic. For the model to actually produce an answer, a competing mechanism has to override it. When Claude is asked about something it knows well, a “known entity” feature activates and inhibits the default refusal circuit, allowing the model to respond.

Hallucinations happen when this override misfires. The researchers showed that when Claude recognises a name but doesn’t actually know much about the person, the “known entity” feature can still activate, suppressing the refusal circuit and pushing the model into fabrication mode. By artificially manipulating these circuits in experiments, they could reliably induce hallucinations about fictional people, and by strengthening the refusal circuit, they could prevent them.

This result reframes hallucination as a circuit imbalance rather than a deep-seated flaw. The model already has the machinery to recognise uncertainty and decline to answer. The problem is that this machinery sometimes loses the tug-of-war with the model’s competing drive to produce fluent, helpful-sounding output. And that drive is reinforced by training regimes and evaluations that treat helpfulness as the primary virtue and treat caution as a failure.

The interpretability work and the OpenAI incentives paper are telling the same story from different vantage points. One looks at the external pressures that shape model behaviour and the other looks at the internal mechanisms those pressures create. Both arrive at the same conclusion. Models don’t hallucinate because they’re broken. They hallucinate because the systems we’ve built around them reward confident output and punish honest uncertainty.

Not All Hallucinations Come From the Model

The OpenAI and Anthropic work both locate hallucination inside the model, whether in its training incentives or its internal circuits. But a September 2025 paper in Frontiers in Artificial Intelligence by Anh-Hoang, Tran, and Nguyen adds a third variable that most evaluation frameworks ignore entirely, and that variable is the prompt itself.

The paper introduces formal metrics for separating prompt-induced hallucinations from model-intrinsic ones — three new acronyms to quantify what practitioners already know, which is that bad prompts make bad outputs worse. Conditional Prompt Sensitivity (CPS) measures how much hallucination rates change when you vary the prompt while holding the model constant. Conditional Model Variability (CMV) measures the reverse, how much rates change across models given the same prompt. A third metric, Joint Attribution Score (JAS), captures the interaction effect between the two.

The results are unambiguous. Vague, underspecified prompts dramatically increase hallucination rates in some models but not others. LLaMA 2 showed CPS values of 0.15 under ambiguous prompting, meaning prompt design accounted for a large share of its fabrication behaviour. GPT-4, by contrast, was far less prompt-sensitive (CMV of 0.08), suggesting its hallucinations were more model-intrinsic and less dependent on how the question was framed. Structured prompting techniques like Chain-of-Thought reduced CPS to 0.06 across the board, a meaningful drop that required no model changes at all.

The practical implication is that hallucination isn’t always a model problem. Sometimes it’s a prompting problem, and sometimes it’s both at once. Models with high JAS scores, like LLaMA 2 under ambiguous prompts (JAS of 0.12), show compounding effects where weak prompts and model limitations multiply each other’s worst tendencies. This means the standard evaluation practice of testing models with fixed prompt templates and attributing all variation to model quality is systematically misleading. Two teams using the same model with different prompt architectures could see wildly different hallucination rates, and neither team’s experience would be wrong.

This reframes the question of responsibility. If a model hallucinates because the prompt was ambiguous, is that a model failure or a deployment failure? Current benchmarks don’t ask this question. They test models under controlled prompting conditions and report a single hallucination rate, flattening a two-dimensional problem into one number. The Frontiers paper suggests that useful evaluation would need to test across a range of prompt qualities, measuring how often a model hallucinates and how sensitive it is to the way questions are asked.

How Evaluation Is Changing (Slowly)

Newer benchmarks are starting to incorporate abstention as a legitimate outcome, but they remain a minority voice in a field still dominated by accuracy-only scoring.

SimpleQA, released by OpenAI in late 2024, treats abstention as a first-class outcome. Each response is graded as correct, incorrect, or not attempted, which makes it possible to measure whether a model knows what it doesn’t know. This is a meaningful step, and the benchmark has been widely cited. But it covers only 4,326 short factual questions with single correct answers, which makes it narrow by design and increasingly saturated. GPT-4o with web search now reaches around 90% accuracy on SimpleQA, and GPT-5 with search and reasoning pushes above 95%, which means the benchmark is approaching its ceiling for models with access to external tools.

HalluLens, presented at ACL 2025, takes a broader approach. It includes multiple task types (short-form QA, long-form generation, and nonexistent entity detection) and explicitly measures both hallucination rates and false refusal rates, the cases where a model declines to answer something it actually knows. This dual measurement is important because it captures a tradeoff that SimpleQA alone misses.

A model that refuses everything would score perfectly on hallucination metrics but be useless in practice. HalluLens found substantial variation across models, with GPT-4o rarely refusing (4.13% false refusal rate) while Llama-3.1-8B-Instruct refused over 83% of the time. Neither extreme is desirable, and having both numbers visible forces a more honest conversation about what good behaviour looks like.

The most ambitious attempt to embed the OpenAI paper’s recommendations into a practical benchmark may be AA-Omniscience, published by Artificial Analysis in November 2025. Its central metric, the Omniscience Index, does exactly what the OpenAI paper prescribed. Correct answers earn +1 point, incorrect answers cost -1 point, and abstentions score zero. This means a model that guesses and gets it wrong is actively penalised relative to a model that admits it doesn’t know. The scale runs from -100 to 100, where zero means a model is correct as often as it is incorrect.

The results are striking, and somewhat grim. Out of 36 evaluated frontier models, only three scored above zero on the Omniscience Index. Claude 4.1 Opus led with 4.8, followed by GPT-5.1 at 2.0 and Grok 4 at 0.85. Every other model was more likely to hallucinate than to give a correct answer when measured on this basis. Models that look excellent on traditional accuracy benchmarks, including Grok 4 and GPT-5 variants, turned out to have hallucination rates of 64% and 81% respectively when their guessing behaviour was properly penalised.

The most recent entry is HalluHard, published in early 2026, which tackles something the earlier benchmarks mostly ignore. It tests hallucination in multi-turn, open-ended dialogue rather than single-turn factual questions. The reason is that errors compound across turns, and an early hallucination can contaminate the context that the model draws on for subsequent responses, creating a cascading failure that single-turn benchmarks can’t detect. HalluHard found that hallucinations remain substantial even for frontier models with web search access, and that models become progressively more prone to fabrication as conversations grow longer.

One of HalluHard’s more interesting results involves the interaction between reasoning ability and abstention. While more effective reasoning generally reduces hallucination, the effect is model-dependent. GPT-5.2 with reasoning enabled abstains significantly more than its non-reasoning counterpart, especially on niche knowledge questions, suggesting that deeper thinking makes the model more aware of its own knowledge boundaries. But this pattern doesn’t hold universally, and some models show the opposite behaviour, where reasoning makes them more confident rather than more cautious.

The benchmark also confirmed something the OpenAI paper predicted, that models struggle most with niche facts that have some trace in training data rather than with completely fabricated entities. When asked about something entirely made up, models are more likely to recognise it as unfamiliar and refuse to answer. But when asked about something they vaguely recognise without knowing well, they tend to guess, because the partial familiarity triggers the “known entity” response that Anthropic’s circuit analysis identified.

Work at the training level points in a more encouraging direction. A December 2025 paper on behaviourally calibrated reinforcement learning showed that a 4-billion-parameter model trained with proper calibration incentives could match or exceed frontier models on uncertainty quantification, despite being orders of magnitude smaller. The model’s signal-to-noise ratio gain (measuring the ratio of correct answers to hallucinations) substantially beat GPT-5 on challenging mathematical reasoning tasks, suggesting that teaching models when to abstain is a skill that can be learned independently of raw knowledge.

Where Evaluation Still Falls Short

Despite this progress, the structural problems the OpenAI paper identified remain largely intact. There are at least four ways in which the current evaluation system continues to fail.

The leaderboard problem persists. The benchmarks that drive public perception, model selection, and commercial decisions are still overwhelmingly accuracy-only. When a new model launches, the numbers that appear in the announcement blog post are accuracy on MMLU, pass rates on SWE-bench, scores on GPQA Diamond. These are the metrics that journalists report, that enterprise buyers compare, and that engineering teams optimise for. Benchmarks like AA-Omniscience and HalluLens exist but remain niche, and until the headline number on a model card includes a hallucination-penalising metric alongside accuracy, the incentive structure the OpenAI paper described will continue to push models toward confident guessing.

Single-turn factuality is an inadequate proxy for production behaviour. Most hallucination benchmarks test whether a model can correctly answer isolated factual questions. But the failure modes that actually hurt people in deployment are different. They involve subtle distortions in summaries, fabricated citations in legal research, invented details woven into otherwise accurate reports, and cascading errors in multi-turn conversations. HalluHard is a step toward tackling this, but it remains a single benchmark. The gap between “can this model answer trivia correctly” and “will this model produce reliable output in my specific workflow” is enormous, and very few evaluations attempt to bridge it.

Domain-specific hallucination is underexplored. AA-Omniscience shows dramatic variation across domains, with different models leading in different domains. A Stanford study in the Journal of Empirical Legal Studies found that even purpose-built legal AI tools like Westlaw AI produce responses that are not significantly more trustworthy than general-purpose models, with hallucinations that require close analysis of cited sources to detect.

A study in npj Digital Medicine found that GPT-4o hallucinated at a 53% rate on medical questions before targeted mitigation, dropping to 23% with improved prompting. These domain-specific rates are far higher than the aggregated numbers that appear on general leaderboards, and they vary in ways that general-purpose benchmarks don’t capture.

Retrieval-augmented generation doesn’t solve the problem. There’s a widespread assumption that giving models access to external documents through RAG architectures eliminates hallucination risk. The evidence doesn’t support this. Vectara’s hallucination leaderboard, which tests grounded summarisation where models are given source documents and asked to faithfully summarise them, still shows non-trivial inconsistency rates across all models tested.

The model can misread the source, over-generalise from it, or fill gaps between retrieved passages with invented material. RAG reduces the frequency of hallucination, but it changes the type rather than eliminating the problem. And because RAG-augmented models often cite their sources, the hallucinations they do produce carry an extra layer of false authority that makes them harder to catch.

The entire evaluation terrain is English-only and text-only. Nearly every benchmark discussed so far tests English-language factual questions in a text-to-text setting. This is a problem because hallucination rates spike dramatically once you step outside that narrow frame. Mu-SHROOM, a SemEval 2025 shared task that tested hallucination detection across 14 languages, found that hallucination rates and detection difficulty vary enormously by language, with low-resource languages showing far worse outcomes than English. The task attracted 2,618 submissions from 43 teams, a sign of the community’s recognition of this gap, and the results confirmed what many suspected. A model that is well-calibrated in English can be wildly overconfident in Swahili or Basque.

The multimodal picture is no better. CCHall, presented at ACL 2025, tests hallucination when models must reason across both languages and images simultaneously. Even the best-performing model (GPT-4o with a multi-agent debate framework) achieved only 77.5% accuracy, with performance dropping 10.9 points compared to handling cross-modal hallucinations alone.

The benchmark also found that longer model responses trigger substantially higher hallucination rates, with a sharp inflection point around 120 words, after which output reliability degrades significantly. These are not obscure failure modes. If you’re deploying a model to handle customer queries in multiple languages, or building a system that reasons over images and text together, your real-world hallucination rate is almost certainly higher than what any English-only benchmark would predict.

Enterprise evaluation is moving in the right direction but slowly. The Bessemer State of AI 2025 report noted that 2025 and 2026 would mark a turning point where AI evaluations go “private, grounded, and trusted,” with enterprises building domain-specific evaluation frameworks tailored to their own data and risk profiles.

This is encouraging, but it is a shift toward bespoke testing that doesn’t feed back into the public benchmarks that shape model development. If enterprises build better evals internally but the public leaderboards remain accuracy-only, the models themselves will continue to be optimised for the wrong thing. The fix needs to happen upstream, in the benchmarks that model developers train against, rather than downstream in the evaluations that buyers run after deployment.

The External Pressure Nobody Planned For

The discussion so far has framed hallucination as an internal industry problem, something the AI field needs to solve through better benchmarks and training practices. But the pressure to fix it is increasingly coming from outside the field entirely.

In June 2023, a New York federal judge sanctioned two lawyers and fined them $5,000 for submitting a brief containing fabricated case citations generated by ChatGPT. The Mata v. Avianca case became the first widely reported instance of AI hallucinations entering the legal system, and it set off a chain reaction. One of the lawyers testified that he was “operating under the false perception that [ChatGPT] could not possibly be fabricating cases on its own.” By mid-2025, courts across the country had moved well beyond fines.

In Johnson v. Dunn (July 2025), a Northern District of Alabama judge declared that monetary sanctions were proving ineffective at deterring AI-generated errors and instead disqualified the offending attorneys from the case entirely. Multiple courts now require attorneys to certify that AI-assisted filings have been manually verified.

The problem extends well beyond law firms, and in January 2026, GPTZero scanned all 4,841 papers accepted by NeurIPS 2025, the world’s most prestigious machine learning conference, and found over 100 confirmed hallucinated citations spread across 51 papers. These included fabricated authors, invented paper titles, and fake DOIs, all of which survived review by three or more expert peer reviewers.

Some were obvious (author names like “John Doe and Jane Smith”), but others were sophisticated blends of real papers with modified titles and expanded author initials. The irony is hard to miss. The leading AI researchers in the world were fooled by the exact failure mode their field is supposed to be studying.

GPTZero had previously found 50 hallucinated citations in papers under review at ICLR 2026, and a separate analysis found that fabricated citations had appeared in US government reports requiring corrections, and in consulting outputs that triggered $98,000 (AUD) refunds.

The pattern is consistent. Hallucinated content doesn’t stop at degrading individual conversations. It enters the official record, whether that’s case law, academic literature, or policy documents, and from there it compounds. Those NeurIPS papers with fake citations will themselves become training data for next-generation models, creating what one researcher called a “self-reinforcing hallucination loop.”

These consequences are materialising faster than the evaluation frameworks are improving. Courts, publishers, and regulators aren’t waiting for the AI field to solve its benchmark problems. They’re imposing external accountability in the form of sanctions and regulatory mandates.

This may end up being the most effective forcing function for better hallucination measurement, not because the field decided to measure the right things, but because the cost of measuring the wrong things became impossible to ignore.

The Collective Action Problem

The deepest issue the OpenAI paper surfaces is structural rather than technical. No individual lab has a strong incentive to score worse on existing benchmarks by making their model more cautious, even if they agree that the benchmarks are measuring the wrong thing. If Lab A trains its model to say “I don’t know” more often and Lab B doesn’t, Lab B’s model will look better on the accuracy-only leaderboards that dominate public comparison. Lab A’s model might be more reliable in practice, but that advantage is invisible to the metrics that drive adoption.

This is a textbook coordination problem. Everyone would benefit from better benchmarks, but nobody wants to be the first to optimise for them at the expense of looking worse on the old ones. The OpenAI paper acknowledges this by framing the solution as “socio-technical,” requiring both a better evaluation and broad adoption of it across the field.

There are signs of movement, though. An August 2025 joint safety evaluation by OpenAI and Anthropic showed the two leading labs converging on “Safe Completions” training that incorporates calibrated uncertainty into model behaviour. Artificial Analysis has folded the Omniscience Index into its Intelligence Index alongside traditional metrics. And newer benchmarks like HalluLens and HalluHard are gaining citations and attention in the research community.

But these are early moves. The central question, whether the field can shift from treating accuracy as the headline metric to treating reliability (accuracy minus hallucination, weighted by abstention) as the headline metric, remains open. Until that shift happens at the level of public leaderboards and model marketing, the incentive structure that produces hallucination will persist even as the models themselves become more capable of avoiding it.

What This Means in Practice

If you’re building with language models today, the practical takeaway from all of this is that you can’t trust aggregate benchmark numbers to tell you how a model will behave in your specific use case. A model that scores 90% on a general factuality benchmark might hallucinate at 50%+ rates in your domain, and you won’t know until you test it on your own data with evaluation criteria that penalise fabrication.

The research points toward a few concrete steps that are worth spelling out. First, when evaluating models for knowledge-intensive tasks, look at metrics that separate accuracy from hallucination rate and include abstention behaviour. The Omniscience Index and SimpleQA’s three-way grading (correct, incorrect, not attempted) provide better signals than raw accuracy alone.

Second, don’t assume that RAG eliminates the problem, and test your retrieval system with adversarial queries and check whether the model fabricates answers when retrieved context is incomplete or ambiguous.

Third, consider domain-specific evaluation, because a model that does well at coding benchmarks may struggle with legal or medical factuality, and general leaderboards won’t tell you that.

Fourth, pay attention to how a model behaves under uncertainty. If it never says “I don’t know” in your testing, that’s a red flag rather than a strength. The AA-Omniscience results showed that models with the highest accuracy often had the worst reliability scores, precisely because they never abstained.

It’s also worth noting that the gap between public benchmarks and production behaviour creates an information asymmetry that benefits model providers at the expense of buyers. A model card that reports 95% accuracy on a factuality benchmark sounds impressive until you learn that the same model hallucinates 60%+ of the time when it encounters questions outside its confident knowledge range. The metrics that count for your use case, things like “how often does this model fabricate a citation” or “what percentage of its medical advice is unsupported by evidence,” are almost never reported in public evaluations. Building your own eval suite, however tedious, remains the only reliable way to understand what a model will actually do with your data.

The OpenAI paper ends with a note that bears repeating. Even a perfectly calibrated model will still produce some hallucinations, because some questions are genuinely unanswerable from any finite training set. The goal isn’t zero hallucinations. It’s a system that knows what it knows, admits what it doesn’t, and is evaluated by metrics that reward exactly that behaviour. We’re not there yet, and the gap between where we are and where we need to be is not mainly a gap in model ability. It’s a gap in how we measure and reward model behaviour. The models are increasingly capable of being honest about their uncertainty. The question is whether we’ll let them.

I am a partner in Better than Good. We help smaller companies build tools and processes using machine learning and artificial intelligence that make lasting improvements to their operations. Talk to us today: https://betterthangood.xyz/#contact

In the late 1990s and early 2000s, a wave of filmmakers made what seemed like an obvious choice. Film stock was expensive, temperamental, required careful storage, and would eventually decay. Digital was immediate, endlessly copyable, and felt like the future. Why keep shooting on a format invented in the 1880s when you could embrace the new millennium properly?

Two decades later, those cutting-edge digital productions are now far harder to restore to modern standards than films shot on celluloid fifty years earlier. A well-preserved 35mm negative from 1955 can yield a gorgeous 4K transfer. A digital feature from 2003, shot on what was then state-of-the-art equipment, might be stuck at standard definition forever.

Days of Future Past

When Danny Boyle shot 28 Days Later in 2002, he chose Canon XL-1 miniDV cameras. The decision was partly practical as the lightweight cameras allowed for guerrilla-style shooting on London’s deserted streets, and partly aesthetic. The harsh, blown-out digital look gave the film an immediacy that felt perfect for a story about civilisation’s collapse.

Cillian Murphy in 28 Days Later

The cameras recorded at 720×576 pixels which is PAL standard definition. For context, a modern iPhone shoots 4K video at 3840×2160 pixels, with roughly 25 times more information in every frame.

At the time, this didn’t seem like a problem. Standard definition was the norm. DVDs looked fantastic compared to VHS. Nobody was thinking about what these films would look like in twenty years.

In contrast, when you shoot on 35mm film, the main standard for movie cameras, you’re not really capturing a fixed resolution. You’re exposing silver halide crystals to light, creating a physical record of the scene with an almost absurd amount of potential detail. The exact “resolution” depends on the film stock and how you scan it, but modern estimates put 35mm somewhere between 4K and 8K equivalent. Some argue even higher for large format stock such as 80mm.

More importantly, that detail actually exists in the negative. It’s been sitting there since the day the film was shot, waiting for scanning technology to catch up. When we remaster Lawrence of Arabia or 2001: A Space Odyssey in 4K, we’re not inventing detail. We’re finally extracting what was always there.

Digital video from the early 2000s doesn’t work that way. What was captured is what exists. Those 720×576 pixels aren’t hiding secret information underneath. The cameras had a fixed resolution, and that resolution is now embarrassingly low by contemporary standards.

The Uncanny Valley of Upscaling

“But wait,” you might reasonably ask, “can’t we just use AI to upscale these films?”

We can. And increasingly, we do. Tools have become remarkably sophisticated at adding plausible detail to low-resolution footage. The results can be impressive, especially for content that wasn’t intended to look “cinematic” in the first place such as old TV shows, news footage and home videos.

The problem is that word. Plausible. AI upscaling doesn’t reveal hidden details. It hallucinates detail that looks like it could have been there. The algorithm examines a blocky, pixelated face and generates what a higher-resolution version of that face might look like based on patterns it learned from millions of other faces.

Sometimes this works brilliantly. Sometimes you get something that sits in a weird uncanny valley, technically sharper but somehow wrong in ways that are hard to articulate. Textures that feel synthetic, skin that looks waxy and fabric that doesn’t quite behave like fabric.

For films that were shot on early digital for aesthetic reasons, aggressive AI processing creates an additional problem. The lo-fi digital texture of 28 Days Later isn’t a flaw to be corrected, it’s part of what made the movie work. Clean it up too much and you lose something that can’t be put back.

This puts restoration teams in an impossible position. Do you present the film as it was intended to be seen, knowing modern audiences on 65-inch 4K screens will notice every compression artifact? Or do you “improve” it with AI, knowing you’re changing the director’s original vision at its core.

A Brief History of Bad Timing

The 2000s were uniquely cursed in this regard. It was the precise moment when digital filmmaking became viable enough that serious directors started using it, but before the technology had matured to resolutions that would remain acceptable long-term.

Consider the timeline.

Late 1990s — Digital video exists but is mostly confined to low-budget indie films and documentaries. The Dogme 95 movement embraces the format’s limitations as aesthetic virtues. Lars von Trier shoots The Celebration on miniDV in 1998.

2000–2002 — Early digital starts appearing in mainstream productions. George Lucas shoots Attack of the Clones on Sony CineAlta cameras at 1080p, declaring it the future of cinema. Boyle shoots 28 Days Later on miniDV. The gates are opening.

2003–2006 — The wave crests. Michael Mann shoots Collateral and Miami Vice on Thompson Viper cameras. David Lynch makes Inland Empire on a Sony PD-150, declaring he’ll never shoot film again. Robert Rodriguez pushes digital filmmaking into family blockbusters with Spy Kids sequels and Sin City.

2007–2010 — The first truly high-resolution digital cinema cameras appear. The Red One launches in 2007, capable of shooting at 4K. The Arri Alexa follows in 2010. From this point forward, digital films generally capture enough resolution to survive future format changes (subject to future radical changes to screen technology).

That roughly seven-year window, let’s call it 2000 to 2007, is a generation of films that were technologically progressive for their time and are now technologically trapped.

Some of the most visually distinctive work of the era lives in this limbo. Inland Empire’s hallucinatory nightmare textures were inseparable from the crude DV format Lynch used. Dancer in the Dark’s raw emotional brutality came partly from being shot on 100 consumer camcorders simultaneously. Open Water’s horror worked because it felt like you were watching somebody’s holiday video turn into a snuff film.

George Lucas enters, stage right

Attack of the Clones (2002) was the first major studio production shot entirely on digital cameras. Lucas had been pushing for this transition for years, convinced that digital was not only the future but actively superior to film.

The Sony CineAlta cameras used for Episodes II and III captured at 1080p. By the standards of 2002, this was impressive, true high definition when most consumers were still watching standard def broadcasts. By current standards, it’s less than a quarter of 4K resolution and roughly a sixteenth of 8K.

4K releases of the prequel trilogy exist, but they’re heavily upscaled rather than derived from native high-resolution sources. Watch them on a large modern display and you’ll notice a certain softness, a lack of the crystalline detail present in the original trilogy restorations (which were shot on film and could be properly scanned at 4K).

The irony here is that Lucas was so convinced of digital’s superiority that he also went back and “improved” the original trilogy with digital effects, effects that were rendered at resolutions that now look dated while the underlying film footage remains timeless.

Why Film Ages Better Than Files

A film negative is a physical object that can be re-examined with improving technology. Better scanners extract more detail. Better colour science improves the transfer. The negative hasn’t changed, but our ability to read it has.

A digital file is a fixed quantity. The numbers in the file are the numbers in the file. You can process them differently, upscale them algorithmically, but you can’t extract information that was never captured.

There’s also the question of format obsolescence. Film is remarkably stable as a storage medium. A properly stored negative from 1920 can still be projected or scanned today using the same principles as when it was created. The format hasn’t changed because the format is physical.

Digital formats change constantly. Codecs fall out of favour. Compression standards evolve and storage media become unreadable. A miniDV tape from 2003 requires increasingly rare hardware to play. A hard drive from the same era might be entirely dead. The theoretical advantages of digital, perfect copying, no degradation, only matter if you can actually access the data.

There are documented cases of studios discovering that digital masters from the early 2000s had become corrupted or were stored in formats nobody could easily read anymore. The Library of Congress has warned repeatedly about the challenges of digital preservation compared to traditional film archiving.

This doesn’t mean film is some perfect archival medium. It absolutely isn’t. Celluloid degrades. Colour stocks from the 1970s and 80s are notorious for fading toward magenta. Nitrate film from the silent era is literally flammable and chemically unstable. Acetate stock can develop “vinegar syndrome,” becoming brittle and unusable. Countless films have been lost because negatives were stored poorly, damaged in fires, or simply thrown away when studios decided they had no commercial value.

The point isn’t that film preservation is easy. It’s that when a film negative is properly preserved (stored at controlled temperature and humidity, protected from light and chemical contamination) the information embedded in those silver halide crystals remains accessible. The ceiling for recovery is remarkably high, even if reaching that ceiling requires considerable effort and expense.

What Happens Now?

Studios and distributors are increasingly turning to AI-powered restoration for early digital films, with mixed results.

The 4K release of something like Collateral is the best-case scenario. The film was shot at 1080p, but the imagery was carefully composed and the digital artifacts were minimal. AI upscaling can add convincing detail without changing the viewing experience at its core. It’s not quite the same as a native 4K source, but it’s acceptable.

At the other end of the spectrum, a film like Inland Empire probably shouldn’t be “restored” in any traditional sense. The blown-out highlights, crushed blacks, and compression artifacts aren’t problems to be solved. They’re part of the film’s visual language. Any version that removes them would be a different movie. Most early digital films fall somewhere between these extremes, requiring case-by-case decisions about how much intervention is appropriate.

A Note on What We’ve Lost

The films shot on early digital aren’t obscure curiosities. They include some of the most culturally important work of their era. 28 Days Later all but invented the modern zombie movie. Inland Empire is Lynch at his most experimental. Collateral is Mann’s masterpiece. The Star Wars prequels, whatever your feelings about them, were childhood-defining for a generation.

These films exist, and will continue to exist, in some form. But the question of how they’ll look to future audiences remains unresolved. Will AI upscaling become convincing enough that the resolution limitations become invisible? Will tastes shift so that early digital aesthetic becomes valued rather than apologised for? Will someone invent restoration techniques we can’t currently imagine?

In Praise of Uncertainty

Early digital films aren’t going to disappear. They’ll be preserved, restored with whatever tools are available, and watched by future audiences who will bring their own expectations and tolerances to the experience.

But there’s something worth recognising about the people who chose digital in the early 2000s, often because it seemed like the responsible, forward-thinking choice. They were wrong in ways they couldn’t have anticipated.

The filmmakers who stuck with “outdated” 35mm through this period, often facing pressure and mockery for their technological conservatism, turned out to be the ones preserving their work most reliably for the future.

Christopher Nolan’s stubborn insistence on shooting film, which seemed almost pathologically nostalgic at the time, now looks prescient. His films from this era scan beautifully at 4K and will continue to scale up as display technology improves. His digital-pioneering contemporaries are stuck trying to make 1080p footage look acceptable on increasingly massive screens.

There’s no triumphalism in pointing this out. Just a reminder that the future is harder to predict than it looks, and the technologies that feel inevitable sometimes turn out to be evolutionary dead ends.

The early digital era produced remarkable films that pushed the medium in directions film stock couldn’t go. Those films deserve to be seen and remembered. But the format that made them possible also trapped them in amber at resolutions that grow more limiting every year.

I am a partner in Better than Good. We help smaller companies build tools and processes using machine learning and artificial intelligence that make lasting improvements to their operations. Talk to us today: https://betterthangood.xyz/#contact

In March 2023, GPT-4 could identify prime numbers with 97.6% accuracy. By June, that figure had cratered to 2.4%. Not a rounding error, not a minor regression, but a 95-point collapse on the same task with the same prompts. If a bridge lost 95% of its load-bearing capacity in three months, someone would go to prison. In AI, the vendor posts a changelog and moves on.

This pattern has repeated with depressing regularity across every frontier provider. Models ship to applause and enterprise contracts get signed on the strength of benchmark screenshots, and then something changes. The model you evaluated is no longer the model answering your customers, and nobody tells you until your production workflow starts producing garbage.

The evidence is not anecdotal

Researchers at Stanford and UC Berkeley tracked this drift formally, comparing GPT-3.5 and GPT-4 snapshots from March and June 2023 across seven tasks. The results were bad enough to make the researchers themselves flinch. GPT-4’s ability to generate directly executable code dropped from 52% to 10%. Its willingness to follow chain-of-thought prompting, one of the most widely used techniques for improving accuracy, degraded without explanation.

“The magnitude of the changes in the LLMs’ responses surprised us,” James Zou, a Stanford professor and co-author, told The Register. The team’s conclusion was blunt. The behaviour of the “same” LLM service can shift substantially in weeks, and nobody outside the provider knows when or why.

This wasn’t a one-off result that got debated and forgotten. The OpenAI developer forums have become a rolling graveyard of complaints. In September 2025, users running GPT-4.1 reported severe intelligence degradation within 30 days of launch, with complex tool calls and multi-step instructions suddenly failing. Similar threads appeared for GPT-4 Turbo in May 2025. The pattern never varies, and by now it has become depressingly predictable. Works brilliantly at launch, degrades silently, users scramble to figure out what broke.

Why this happens (and why the incentives encourage it)

There are at least four mechanisms that can degrade a deployed model, and most frontier providers are using all of them simultaneously.

Quantisation is the most technically straightforward of the four, and the easiest to understand. A model trained in 16-bit or 32-bit floating-point precision gets compressed to 8-bit or 4-bit integers for serving. The arithmetic is straightforward enough, since a model stored in FP16 needs roughly two bytes per parameter, so a 70-billion-parameter model demands about 140GB of VRAM just for weights. Quantise to 4-bit and you cut that to around 35GB, enough to run on hardware that costs a fraction as much.

The trade-off is supposed to be minimal, and Red Hat’s analysis of over 500,000 evaluations found that 8-bit and 4-bit quantised models showed “very competitive accuracy recovery” on most benchmarks, especially for larger models. But that phrase “most benchmarks” is doing heavy lifting. Quantisation works by rounding, and rounding destroys outlier values. The weights that fire rarely but matter enormously for edge-case reasoning are exactly the weights that get flattened first. For standard tasks you barely notice the difference, but for the specific hard problems your production system was built to handle, the gap can be catastrophic. One developer reported that dynamic quantisation of a 3B-parameter model dropped accuracy from 65.6% to 32.3%, a halving that no benchmark average would predict.

Mixture-of-experts routing is the more interesting culprit, and the one providers talk about least. DeepSeek’s V3, for example, has 671 billion total parameters but only activates about 37 billion per token. The economics are irresistible because you get the capacity of a massive model with the inference cost of a much smaller one. But the router decides which experts handle which queries, and routing decisions are probabilistic. A query that activated your model’s strongest expert subnetwork at launch might get routed differently after an update to the routing logic, or after the provider adjusts load balancing to handle peak traffic. The user sees the same model name in the API response. The actual computation behind it may have changed entirely.

Distillation and model substitution is the elephant in the room that everyone suspects but nobody can prove definitively. Rumours have circulated since mid-2023 that OpenAI routes some queries to smaller, cheaper models behind the same API endpoint. The Gleech.org 2025 AI retrospective put it plainly: “True frontier capabilities are likely obscured by systematic cost-cutting (distillation for serving to consumers, quantisation, low reasoning-token modes, routing to cheap models).” GPT-4.5 was retired after just three months, presumably because the inference costs were unsustainable, even though it still ranked in the top five on LMArena for hallucination reduction nine months later. The model that performed best got killed because it was too expensive to run.

Safety tuning and RLHF adjustments create the subtlest form of drift. When OpenAI tightens content filters or adjusts the model’s tendency to refuse certain queries, those changes ripple through the entire behaviour space. The Stanford study found that GPT-4 became less willing to explain why it refused sensitive questions, switching from detailed explanations to terse “Sorry, I can’t answer that” responses. The model may have become safer by one measure, but it simultaneously became less transparent and less useful for legitimate applications that happened to brush against the updated boundaries.

The economics are doing exactly what you would expect

Running frontier models is staggeringly expensive, and every provider is under pressure to reduce cost-per-token. The maths, as one industry analysis noted, resembles building more fuel-efficient engines and then using the efficiency gains to build monster trucks. Token prices have dropped by a factor of 1,000 in three years, but reasoning models now generate thousands of internal tokens before producing a single visible output, and 99% of demand shifts to the newest model the moment it ships.

Providers respond by doing what any business would do. They optimise for throughput and margin, quantising the weights and routing easy queries to cheaper subnetworks while distilling the flagship into something that passes the benchmarks but costs a tenth as much to serve. The individual techniques are all defensible, but stacked together and applied silently, they create a system where the model’s advertised performance diverges from its delivered performance over time.

DeepSeek made this trade-off explicit and turned it into a business strategy. Its V3 model serves inference at roughly 90% below comparable OpenAI and Anthropic rates, and the MoE architecture that enables this pricing is openly documented. Whatever you think of the approach, at least the engineering trade-offs are visible. The problem is worse when providers make the same trade-offs quietly, behind an API that returns the same model identifier regardless of what actually computed the response.

What this means if you build on top of these models

The practical upshot is unpleasant but straightforward. If your application depends on consistent model behaviour, you are building on sand that shifts without warning. The Stanford researchers recommended continuous monitoring, and they were right, but monitoring alone doesn’t solve the problem, because it tells you something broke without stopping it from breaking.

Pinning to a specific model snapshot helps, where providers offer it, but even snapshots get deprecated. OpenAI maintains them for a few months and then requires developers to migrate. The careful evaluation you ran against the March snapshot becomes irrelevant when you’re forced onto the June version and nobody can tell you exactly what changed.

The deeper issue is one of trust and transparency. When a model provider updates a live model, they are unilaterally changing the behaviour of every application built on top of it. That is not a software update but an undocumented API change, the kind that would trigger outrage in any other engineering discipline. Imagine if AWS silently swapped your database engine for a cheaper one that was “approximately equivalent” on standard benchmarks, and you can begin to see how the AI industry has somehow normalised something that would be career-ending negligence anywhere else.

Where this leaves us

The model you benchmarked, the one that earned the contract, that impressed the board, that your engineers spent weeks building prompts and evaluation harnesses around, is a snapshot of a moving target. Quantisation shaves off the edges while routing sends your queries to whichever expert subnetwork happens to be cheapest that millisecond, and safety updates redraw the boundaries of what the model will and won’t do. None of it shows up in the model name string your application receives in the API response.

Somewhere in a data centre, the accountants and the alignment researchers are both pulling the same model in different directions, one toward cheaper inference and the other toward tighter guardrails, and the engineers who built their products on last month’s version are left checking the forums to figure out why everything stopped working on a Tuesday.

I am a partner in Better than Good. We help smaller companies build tools and processes using machine learning and artificial intelligence that make lasting improvements to their operations. Talk to us today: https://betterthangood.xyz/#contact

Enter your email to subscribe to updates.