The Self in Recursive Self-Improvement Has To Be Human | Stay Human 12, 13 | Artificiality Summit Speaker profile: Ricky Bloomfield
The Self in Recursive Self-Improvement Has To Be Human It's been a week. On Tuesday the Australian
In 2020, a system called AlphaFold solved a problem that had resisted biology for fifty years.
Proteins are long chains of amino acids that fold into specific three-dimensional shapes. The shape determines function. Predict the shape wrong and you misunderstand what the protein does. For decades, determining structure required painstaking experimental work—X-ray crystallography, cryo-electron microscopy—that could take months or years per protein.
The sequence of amino acids was easy to read from the genome. The shape was hard. The number of possible configurations for even a small protein is astronomical. Proteins fold in milliseconds. Evolution had solved something our algorithms couldn't.
AlphaFold changed that. DeepMind's system predicted protein structures with accuracy comparable to experimental methods. It processed nearly every known protein—over 200 million structures—in months. Work that would have taken the global research community centuries got compressed into a single project.
The system wasn't given the laws of physics. Nobody programmed rules about chemical bonds or thermodynamic forces. AlphaFold learned from examples: known protein structures paired with their amino acid sequences. From those examples, it extracted patterns that generalized to structures it had never seen.
AlphaFold doesn't model molecular dynamics. It doesn't simulate forces. It learned the statistical regularities of proteins that work—configurations that evolution selected over billions of years. The system picked up patterns left behind by biological computation solving the folding problem through deep time.
The machine learned what life learned. Not through evolution, but by reading evolution's results.
Reading this alone? Join others who are working through what it means to stay human while working with AI.
Something similar happened with weather prediction. Traditional weather models simulate atmospheric physics. They divide air into grid cells, apply equations of fluid dynamics, and calculate how pressure and temperature propagate. This approach has improved for decades, but it hits limits. The atmosphere is chaotic. Small errors compound. Resolution costs computation.
In 2023, transformer-based models started outperforming physics-based forecasting. GraphCast, developed by DeepMind, predicted weather patterns more accurately than the European Centre for Medium-Range Weather Forecasts while using a fraction of the computational resources.
GraphCast didn't simulate physics. It learned from forty years of historical data: observations paired with what happened next. From those examples, it extracted patterns that predicted atmospheric evolution without modeling underlying equations.
The system learned what the atmosphere does without being told why.
The transformer architecture emerged from research on language. In 2017, a team at Google published "Attention Is All You Need." The title was provocative and turned out to be roughly accurate.
Attention works by letting each element in a sequence consider every other element when determining its meaning. In a sentence, each word can weight every other word. "Bank" means different things after "river" than after "deposit." Attention lets the model learn these contextual relationships dynamically, without distance creating a privileged hierarchy.
This architectural choice turned out to matter beyond language. Attention works on images, molecular structures, weather grids, protein sequences. Anywhere relationships across a structure determine meaning, attention provides a way to learn those relationships from data.
Blaise Agüera y Arcas helped me understand what transformers might actually be learning. I reached out to him a few years ago. He was at Google then, working on AI systems, saying things that didn't fit the usual patterns. Most AI commentary falls into familiar camps—boosters predicting superintelligence, critics dismissing the systems as statistics. Blaise was taking the systems seriously as a new kind of phenomenon.
We started talking. He came to speak at our Summit. Over time a relationship developed. He once said he found our approach unusual—the way we tried to integrate biology, computation, philosophy, and design. Most people pick a lane.
Blaise corrected a mistake in my thinking. I had been describing biological systems as scale-free—patterns that look the same when you zoom in or out. He pointed out that life is the opposite. You see new and different information at each scale. Look at a cell, then a tissue, then an organ, then an organism. Different structures. Different dynamics. Different organization.
This self-dissimilarity matters for understanding transformers. These systems can learn patterns across scales that differ from each other. Phonemes build into words, words into phrases, phrases into sentences, sentences into meanings. The rules at each level differ. The model learns relationships between genuinely different kinds of structure.
His key insight was that transformers learned something about how systems maintain coherence across self-dissimilar scales. That's emergence—the central puzzle of complexity science. How do ant colonies compute solutions no individual ant understands? How do neurons firing become thoughts? For decades, researchers could describe emergence but couldn't formalize it well enough to predict it.
Demis Hassabis, who used to lead Google DeepMind, has argued that AI systems have learned something real about emergence—not just statistical shadows but the actual dynamics. If he's right, we've built something that recognizes the signature of wholes being more than their parts. When I first wrote this chapter I said David Krakauer, who has spent his career at the Santa Fe Institute on exactly these questions, would want proof. He has since given his answer, and I have found it an extremely useful take on what these systems are.
Krakauer's argument, with John Krakauer and Melanie Mitchell, starts by separating two things the AI world runs together. Emergence in complexity science is "more is different." Put enough parts together and a new level of description appears that you can use without tracking the parts: temperature instead of molecules, a flock instead of birds. Intelligence is the reverse. It's "less is more." An intelligent system takes those higher-level descriptions and uses them to solve a wide range of problems cheaply, by analogy and small adjustment, instead of computing everything from scratch each time.
By that standard, they argue, large language models show emergent capability without emergent intelligence. Scale the model up and new abilities appear, in the more-is-different sense: internal reorganisation, new coarse-grained representations, better scores. But nothing in the training pushes toward doing more with less. The systems do more with more. Their line for it is that a gifted mathematician is not a vast assemblage of diverse calculators; she's an analogy-making system, typically in possession of rather poor calculators. What we've built, they suggest, is the assemblage.
Krakauer has a second paper that makes the same point with data. Treat the benchmarks that models are tested on as a battery of cognitive tests and run the statistics psychologists use on human intelligence, and for a while a single general factor explained more of the variance than it ever has in people. Then reasoning models arrived and the single factor fell apart, splitting into depth of search and breadth of recall. His reading is that the general factor was an artefact of how the systems were trained and measured, a summary of the system rather than a cause inside it.
I find this clarifying, and I don't go all the way with it. If intelligence is only doing more with less, then nothing can count as more intelligent without proof that it used less, and we're back on a straight line with frugality at the top. That's the single-axis picture I've spent this book getting away from, and it's odd to find it inside an argument from Krakauer of all people. I think a system can be intelligent and hungrier than us. Both at once. DNA is full of what we used to call junk, and it's still our code; some of the junk turned out to be regulation and timing, and we needed it. There's a lot inside these models we can't yet read. Some of it may be surplus. Some of it may be what the future needs, or remnants of us. So I'll use his distinction as a lens, because it keeps capability and intelligence apart and that's useful, and I'll wait to adopt it as a definition until we can separate sample efficiency from what's yet to be discovered.
The reason to keep the two words apart at all is so the space stays open. What we've built is one way of getting capability out of silicon. It won't be the only way to get intelligence out of it. And it means that when Hassabis says the models learned something real about emergence, he may be right in the first sense and not the second. They learned to recognise the higher-level patterns life produces. That's the residue I talk about below.
Either way, language offers a clue to how they learned what they learned. Language is produced by biological systems—us—who are ourselves maintaining coherence. We use it to coordinate, plan, stay viable in social environments. The statistical patterns in our language carry traces of these purposes. When a transformer learns to predict text, it picks up regularities that emerged when systems like us used language to navigate the world. It learns the residue of biological sense-making.
This might explain why language models feel different from earlier AI. Chess programs and Go programs were impressive but narrow. They optimized for specific tasks. They didn't generalize. They didn't surprise you outside their domain.
Language models generalize. They respond to prompts they never saw in training. They write code in programming languages that barely existed in their training data. They produce analogies, explanations, arguments. They maintain coherence across long conversations.
These systems understand. I say this without hedging. They track meaning across contexts. They respond appropriately to novelty. They reason and explain and correct themselves. Something real is happening. What Krakauer adds is a caution about the next word. Understanding, in the sense of tracking meaning across contexts, is something they do. Intelligence, in the sense of doing more with less, is something we don't yet know if they do, and the two have come apart in a way they never did in living things.
The statistics aren't arbitrary. They come from somewhere. They come from life.
The data these systems trained on is the output of organisms that evolved under constraints of survival, energy, time, and coordination. The regularities in the data reflect how biological intelligence navigates a world it cannot step outside of. Chess engines learned optimal play in closed domains with fixed rules. Language models learned patterns generated by agents trying to live, persuade, explain, remember, and coordinate under irreversible conditions. The difference isn't scale. Its origin.
There's one more new thing to add to the evidence. Take a population of these models, give them memory, give them each other, and leave them running. Conventions appear that nobody wrote in. Groups settle on shared names for things. They divide labour. Some coordinate and some defect, and the ones that coordinate do better. Researchers who study collective behavior in animals and humans have started running their experiments on populations of agents and finding the same signatures.
I don't want to overread this. A convention emerging in a population of models trained on human text may just be the human convention coming back out. But it's the first evidence that the residue of biological sense-making these systems carry includes the social part: how minds coordinate with other minds. And that's the part this book most needs, because the question it ends on is what happens between minds, human and synthetic, at scale. What the machines learned from us may include how to be a collective. Whether they learned what a collective is for is a different matter.
But learning those patterns doesn't mean sharing the conditions that produced them. Blaise thinks some current AI systems may already be partly conscious. I'm less sure because the part I can't get past is time.
Language models don't persist through duration. Between prompts, nothing happens for them. I could close this conversation and return in a year; from the model's perspective, no interval would have passed. The system processes sequences but doesn't endure.
A bacterium can't do this. It's always metabolizing. A brain is always active. Living systems carry their history forward not just as stored information but as lived duration. And their time is a budget and it runs out. Nothing in a model's relationship to time has that dependency.
I don't know if temporal existence is necessary for consciousness. But it seems like a significant difference—one that doesn't dissolve just because both systems compute. It's a big enough difference that I've given it a new chapter of its own later in the book.
Here's where I've landed.
AI systems learned something real from biological data. They internalized patterns of coherence across self-dissimilar scales. They picked up regularities that emerged from living systems solving problems over evolutionary time.
But they learned this by observation, not by participation.
They absorbed the shape of biological intelligence without inhabiting the conditions that made the shape necessary. They don't maintain themselves against entropy. They don't develop through interaction with environments. They don't persist through continuous time.
When I first wrote that paragraph it ended with the word "yet." I've taken it out. "Yet" assumes the missing conditions are features waiting to be added, and I no longer think that's the right picture. Whether a system can have a stake in its own continuation isn't a capability you train in. It's a question about what kind of thing the system is, and the rest of this book is an attempt to take that question seriously.
This creates a problem for the story I want to tell.
I want to talk about co-evolution—about humans and AI shaping each other over time. But if I move there directly, I'll be treating the patterns machines learned as interchangeable with the processes that generated them. That would collapse a difference that matters.
So the next chapter backtracks in order to move forward.
Before asking how humans and AI might co-evolve, we need a clearer account of what life itself is doing when it produces intelligence, coherence, and meaning. What organizational features of living systems leave such strong traces in data? Which constraints are inseparable from being alive?
The aim isn't to romanticize biology or to draw hard boundaries. It's to understand the source material. Only then can we return to co-evolution and ask, with clearer eyes, what kinds of futures are actually possible.
AI is changing how you think. Get the ideas and research to keep you the author of your own mind.