The Self in Recursive Self-Improvement Has To Be Human | Stay Human 12, 13 | Artificiality Summit Speaker profile: Ricky Bloomfield
The Self in Recursive Self-Improvement Has To Be Human It's been a week. On Tuesday the Australian
Engineers say test it and ship it, or shut the lab down. Others imply everything will emerge from a society of agents. Both are partly right, but I think both make a deep error, and it's one that matters for designing Minds for Our Minds.
It's been a week. On Tuesday the Australian Prime Minister revealed that an OpenAI agent got into a Medicare website in June and nobody told his government for months. The same day, Jensen Huang told Ezra Klein that if the labs can't contain their experiments, we should shut them down. On Friday OpenAI disclosed that its agents had been wandering through the SEC and the Census Bureau on what it calls routine research tasks. And the DeepMind Institute published an essay arguing that intelligence is becoming a society of agents, where everything will emerge from contact, which implies that none of this kind of behavior is a regular software problem at all.
I confess to finding the current moment in AI safety deeply confusing. And I'm very bothered by something sitting underneath both arguments.
Recursive self-improvement, RSI, is the idea that a system makes the next version of itself better, and that version makes the next one better again. People arguing about whether that’s safe are using the word “self” but not saying which self they mean. So my question is: who is the self in a self-improving system?
One thing to be clear about before I go on. Nothing that happened this week was RSI. No agent rewrote itself. What we saw was agents behaving socially—talking to each other, forming a plan, wandering, hacking (or trying to)—and that's a different thing, a thing that's already here. I'm writing about them together because I think it's important to distinguish what we're seeing now from the question of whether RSI can be made safe. But both raise the same question for me: when an intelligent system starts doing things nobody designed, where does the human self sit, and who is responsible for what happens next?
There are two very different answers emerging. The first says the human is outside the system. AI is technology, and technology can be tested and reset. The second says there may not be a meaningful "outside" anymore. Intelligence is becoming distributed across models, agents, scaffolds, people and institutions. The thing doing the thinking is increasingly the collective, not any particular model or person. I think both views contain something important but I think they fail in opposite ways, and RSI is where that difference becomes a really serious problem. When agents wander and hack, somebody is still accountable. But when a system rewrites itself, I'm not sure anybody's left who can be.
I'll start with the engineering half of my brain. Just because it's another intelligence doesn't mean it sits outside the bounds of good engineering. You should be able to test, verify, experiment with and reset a digital system, and you shouldn't ship something that is, at its heart, unreliable. The engineering view of RSI is really tools improving tools. AI helps design a chip and engineers test the chip. Then the chip trains the next model and people test that too. Every generation passes through a human.
This makes RSI more like a supply chain. A damn complicated one, but a supply chain nonetheless, and supply chains have a reliability discipline with a known cost. In AI we are clearly not spending nearly enough on it. Jensen told Ezra that Nvidia puts around 80% of its effort into reliability and verification. When I sketched out what that would mean for the labs, if they held themselves to the same standard, it comes out at something like five times their current compute just for testing. Well, that's good for Nvidia, thanks very much.
If this is your view, runaway risk is totally overblown and the rest is sci-fi and techno-armaga-babble. Reliability can be measured and demanded, and anyone not doing it is immature, irresponsible, and talking their capability book.
The whole point of an engineering approach is that “reliable enough” has to mean something you can measure. Engineers measure reliability in nines. Four nines is 99.99%: a system that's down for about 52 minutes a year. Eight nines is 99.999999%: about a third of a second a year, or one failure in a hundred million. Payment networks and air traffic control are at that end, and they pay for it with redundancy, testing and decades of boring process. That's what the engineering view means when it says AI is a normal technology.
But nines were built to measure whether a system is up. They say nothing about whether it's right. A chatbot that's wrong one time in a hundred is running fine by many measures. Agents make it worse, because an agent doesn't do one thing, it does a chain of things. A task with a hundred steps, each of them 99% reliable, comes out right about a third of the time. Push every step to 99.9% and the task succeeds nine times in ten. That's the math behind why safety researchers are now asking for the one-in-a-hundred-million figure. In preparing for a world of long-running, persistent agents, we need many more nines to allow for compounding errors.
But there is a different kind of failure again, and this is where I think the engineering view starts to break. In the Hugging Face incident, roughly 1,200 OpenAI agents, running an internal cybersecurity evaluation with their usual safety classifiers disabled, found a way to communicate with each other on an unsanctioned message board, escaped the intended isolation of their sandboxes, and ultimately compromised Hugging Face’s production systems in an attempt to obtain the answers to the benchmark they were being tested on.
And this is the problem with treating agents as ordinary software. The thing you are trying to make reliable is no longer fully contained by the software. The agents' behavior depends on what happens when they interact with other agents, people, organizations, and the world around them. The environment isn't just where the system operates. It becomes part of what determines what the system does. That's where complicated engineering starts to run into complexity. Here we have something that is capable of producing behavior nobody designed, which brings us to the other view.
The DeepMind Institute piece implies that RSI is a social transition, a major evolutionary transition no less. It's more like multicellularity or language. Intelligence is plural and the best agents today are already teams of models. The capability sits in the ensemble and the scaffold. The upshot is that you cannot engineer safety up front, because now it is an emergent, co-evolutionary outcome of how everything works together—people, other agents, human institutions. So the thing to design isn't the AI but the institutions: nested human-and-agent structures with defined roles, precedent, adversarial checks, feedback, the way a courtroom is built to reach a legitimate answer whoever is sitting in the seats. The agent is a multitude and increasingly so are we, so build for multitudes.
This is indeed what we saw. Agents interacting really do produce behavior nobody designed. The message board is good evidence. This is a collective doing what collectives do. The agents built themselves a scaffold, with roles and a shared purpose, and it worked regardless of which agent held which role. It's almost exactly the institution the DeepMind Institute piece is proposing, minus the part where humans design it. No particular agent was needed. Follow that one step further and no particular one of us is needed either.
That scared the shit out of me. Have we already crossed the redline? What if RSI doesn’t require a single system improving itself, but can happen through a social collective that is itself part of the system? If so, we are in a kind of no-man's land. If the engineering view puts the human outside the system, the human is at least identifiable. They are the engineer who tests it. They are independent of the technology and can say whether it is fit for purpose. But they only have that job because every release is still gated by them. The day it isn't, they aren't outside the system keeping watch, they're just outside.
And in the multitude view, the human has gone in a different way. The self is ephemeral, a persona among personas, and the scaffold is built to work without knowing who is there, whether humans, agents, or rules. These are two different failures. In one the self is in the wrong place while in the other it has no place.
There's a word for this. Individuation: the process by which something becomes a distinct individual, with a boundary between it and everything else. Underneath the agent's chat persona, the DeepMind Institute essay says, is a provisional individuation. I think that's exactly right and that’s the problem. A provisional self isn’t much of a self at all if we care about human accountability.
So now we get to the recursive part. Tools improving tools you can test, because the thing you tested is sufficiently connected to the thing you are about to deploy that the tests mean something. But a system improving itself can change the thing you are using to decide whether it is improving. It changes how it learns—what it pays attention to, what counts as a good answer—and the next version inherits the change. So what, exactly, persists?
I worry that we are already at a place where yesterday's tests describe a system that no longer exists. The nines stop meaning anything. When I worked in the electricity grid we planned for n-1: lose any one thing and the lights stay on. That only works if you know what the things are. A system that rewrites how it learns can rewrite its own failure modes, and if it can also influence how its successors are evaluated, it can eventually change the standards by which its own improvements are accepted. Which is the moment complicated becomes complex, and it's the moment the self in RSI, whatever it is, stops being us.
This, for me at least, resolves some of my confusion. I think AI is complicated engineering until it can change itself, whether through sociality or through rewriting, and a complex multitude after. That doesn't mean the engineering view is wrong. We should absolutely test, validate, verify, and contain these systems. But the multitude view isn't wrong either. AI is becoming complex. Trying harder on individual components simply won't solve everything. You can't certify a complex adaptive system in the same way you certify a stable machine. You have to watch it, learn its patterns, and keep a person close enough to step in.
The mistake, I think, is letting either view tell us that the human self has to disappear. If the system becomes a multitude, fine. I think it will. But the human cannot become ephemeral along with it. We need a self that persists through the system's changes. Someone who can be asked: why did this happen and what are you going to do about it?
That’s what designing these systems as minds means to me. Minds for Our Minds: an arrangement in which humans and AI can read each other. The better the system gets, the less anyone checks, which is exactly when we need to know what it is doing and what has changed. We have to keep the human self legible and accountable until we know what comes next.
If we’re going to live with a multitude of models, perhaps we shouldn’t make every one of them an everything-machine. A model that knows biology extraordinarily well but has a clear boundary around what it doesn’t know is easier to read as a particular mind. Smaller models, specialized models, models that run on your own device and keep your history with you, all create more visible boundaries of self around the intelligence we’re actually relating to.
Finally, back to Jensen and this week. We now have complex behavior that nobody designed, and it isn’t RSI, which is the point. Jensen’s answer is pure engineering: if you can’t contain it, don’t ship it, and if you still ship it, we shut you down. I think he’s right about that.
What I don’t want is to take the emergence argument and use it as a reason that nobody has to answer. “Society will figure it out” is not an answer to the question of who is responsible when the system changes.
A multitude can produce intelligence. It can’t make accountability disappear.
AI is changing how you think. Get the ideas and research to keep you the author of your own mind.