The Instability Is a Choice | Summit Speaker Profile Josh Lovejoy | What the Safety Folks Really Think | The Laws of Thought

The Instability Is a Choice

By Dave

We've all become accustomed to variability in AI responses. Sometimes it frustrates—you want certainty and can’t get it. Sometimes it's the point—the source of creativity, exploration, and expansion. This variability is inherent to generative AI. It comes from the model itself, and learning to work with it is part of learning to work with these systems at all.

But there is a second phenomenon, and it is not variability at all. It is instability—created above the models by product decisions. Variability is inherent to intelligence. Instability is inherent to nothing. And it is wholly unacceptable.

Consider the small version first. As I write this, Claude is prompting me to update to Version 1.22209.0. I've lost count of recent updates—it feels like nearly one a day. I'm never told what an update will change. Is it a security patch? A redesign that moves features and forces me to relearn the tool? Will it change how Claude behaves on an existing project? I have no idea whether this update will add to my experience or subtract from it. The instability is a choice.

Now the large version. For the past few weeks, every regular Claude user I know has been experiencing Fable anxiety. Anthropic's latest model, Fable, was made available for a limited time. Then it was folded into the usage we already pay for—with limits that are entirely opaque. Then Anthropic announced that Fable access would end, with no guidance on whether or how users could pay to keep it. Then the deadline was extended. And now, apparently, access will remain part of paid plans—again, with opaque limits. The instability is a choice.

A designer friend described it as the biggest A/B test they've ever been inside. That sounds about right.

I'm far from the only one to say it: This has to stop.

Anthropic, OpenAI, Google, and others are attempting to build a critical intelligence service—one used by people across the globe, every day, and paid for handsomely. That ambition comes with obligations.

Critical services must work. Full stop. Not when the vendor feels like it. Not depending on whether an update shipped overnight. When the user needs it. Critical services must be predictable. The experience, the cost, the outcomes. And a service asking for mass adoption must be legible: people need to understand what they're paying for and what they'll get. Right now, users can't answer either question.

For a few years, we have argued that trust will be the currency of this new economy. We have to trust AI systems enough to give them the context that makes them useful. We have to trust the companies behind these systems to build for us, not to take from us. And we have to trust that the systems we come to depend on will be there when we need them. Like any currency, trust holds its value only when its issuer is steady.

Over the past few weeks, Anthropic damaged that trust with its most committed customers—the people paying hundreds or thousands of dollars a month—by making them wonder whether they could count on the product at all. OpenAI did as well with its merging of ChatGPT and Codex, prompting users to update into a different product overnight. In the race to create the future, every AI company is forgetting to make great products—and forgetting the people who buy them.

Our vision is that AI systems can become minds for our minds—tools that don't just carry our thinking faster but think with us. A mind for your mind has to meet a standard no ordinary product does. Philosophers who study the extended mind are precise about when a tool becomes part of how you think: it must be reliably available, its contents trusted automatically, consulted without hesitation. The philosopher's classic example is a notebook that functions as a person's memory—because it is always in their pocket and never rearranges itself overnight. A system that changes without explanation, whose capabilities appear and disappear on a vendor's whim, whose limits no one can state, cannot cross that threshold—no matter how capable the model underneath. And the cost is not contained to the product: when a tool becomes part of how we think, its instability becomes ours.

Every business aims to build loyal—even sticky—customers. But instability gives customers nothing steady to stick to.  A product that asks to sit this close to how we think must hold itself to the highest standard.

Give us something stable to think with.


Artificiality Summit Speaker profile: Josh Lovejoy

“We need to stop treating humans as cells in a spreadsheet and start treating them as cells in a body.” — Josh Lovejoy

I love this quote. It captures something stark about the current moment in tech: so much of it seems bent on turning us into machines rather than respecting what it takes to be human with AI. And Josh has earned the right to say it. He was making the case for designing AI around people before the industry knew it needed one.

Josh has led design at the places where the industry's hardest questions actually get decided: Google's People + AI Research initiative, Microsoft's Ethics & Society team, Amazon—where he worked on the UX of probabilistic systems at Prime Video—and now Autodesk. Alongside fellow Summit speaker Jess Holbrook, he co-created Google's People + AI Guidebook, the framework that gave an emerging field its first practical language for designing AI around human needs rather than model capabilities. Much of what is becoming common practice today appeared in their work years ago.

What makes Josh essential right now is that he says the thing the industry doesn't want to hear. His recent essay "Hallucination Was Never the Problem"—which we're delighted to be publishing at Artificiality—argues that the field has been fixated on the wrong failure. AI systems are trained on the outputs of human executive reasoning: our writing, our arguments, our after-the-fact rationalizations. They're built to do the top-level guessing, not the ground-truth measuring. "They took the executive's chair. Which leaves exactly one role open for the rest of us… But none of this is fated. The inversion is a choice about who sits where, and it can be made the other way."

Josh is contrarian but in service of a larger aim. He has an unusual feel for the psychology of this moment—what these tools are doing to how people trust, decide, and understand themselves—and he is one of the people actively shaping how the design community will meet that reality.

At the Summit, he'll share how AI is changing trust at every level: trust in companies, trust in brands, trust in AI itself, and trust between people. Trust is one of the ways humans navigate unknowing—we make decisions every day without complete information, using trust as a proxy for knowing—and Josh has thought hard about the conditions under which trust in AI can be earned, calibrated and sustained. As an industry insider, he knows the difference between what gets said publicly and what's actually being reckoned with. In Bend, he'll share his views about where the industry really is, and where design can go from here.

Josh has been part of the Artificiality community since its earliest days. His thinking has influenced our work from the beginning, and many of the ideas explored at this year's Summit have grown through conversations with him over the years.

Read his article in our journal here.

And tickets for the Artificiality Summit this October:

  • Don't Wait: Our standard price of $1,795 (or $2,895 for two) will end on August 31 and the price will increase to $1,995 (or $2,995 for two).
  • Don't Worry: We have updated our cancellation policy to better align with similar events and organizations. Through August 23 (60 days prior to the Summit), you can receive an 80% refund, and then through September 22 (30 days prior to the Summit), you can receive a 50% refund. At any time, you can transfer your ticket to someone else or you can apply the full value of your ticket to next year's Summit. You can see all the details on the Summit page.
  • Don't Forget: We provide a 30% discount for education and nonprofit customers. Send us an email for more information.

Artificiality Summit 2026

Join us October 22-24, 2026 in Bend, Oregon for 2.5 days with a fantastic group of speakers—academics, authors, designers, investors, photographers, and more.

We'll explore the theme of Unknowing—not ignorance, but a necessary release of inherited assumptions. We don’t yet know what AI will become, and we don’t yet know what we will become in relationship to it.

Unknowing is the space between—the place where neither side is fixed, and something new can emerge.

Register Now

Asilomar Annual AI Safety Conference

By Helen

Dave and I are visiting researchers in Stuart Russell's lab (Center for Human-Compatible AI at UC Berkeley) and we attended their annual meeting at Asilomar in June. Here's a few insights we took away from the Chatham House rules event.

Nobody agrees on what “safe” actually means in 9s

Two very senior people gave reliability targets four orders of magnitude apart. One wants "eight nines" of assurance (99.999999%), the other is fine with 99.99%. That's a huge disagreement about how much failure we should tolerate.

Risk estimates were just as scattered. One person put misuse (plus what I'd call just kinda bad use) at around 15%. Another put extinction risk at 50%. When people this close to the problem land this far apart, the takeaway is that nobody knows. In private, the field admits this.

AGI is solely an economic definition now

The working definition spoken about: a system that can do 90% of remote work at equal or lower cost than a human. No consciousness tests, no philosophy. Labor substitution economics. Which means the threshold gets crossed in procurement decisions, well before any headline moment.

The lab strategy has an RSI logic problem

The standard lab position: recursive self-improvement is a red line, and the safety strategy is to stay a handful of months ahead of everyone else. I kept wanting to ask: why months? What is a few months' lead worth if the thing you're racing toward is your own red line? Nobody had a clean answer.

Gradual disempowerment is actually the biggest “safety” thing

This got serious airtime. The concern is less a sudden takeover and more that human influence erodes step-by-step as we hand economic, cultural, and political functions to systems that work well enough that we stop checking. No villain, no single moment to point at. This maps closely to what we study at the Institute.

Which sci-fi story are we in

One thread contrasted Le Guin's carrier bag theory of fiction with hunter stories. Most AI risk scenarios are hunter stories: conflict, pursuit, a decisive confrontation. If the risk actually looks more like the carrier bag—an accumulation of small dependencies gathered over time—then our threat models are tuned to the wrong genre.

The open question

Can we have genuinely useful AI and keep control over it? The meeting didn't answer this. What struck me is that the field's most fundamental question is still open, while most public discourse proceeds as if it's settled.

"Prompt and pray"

Someone described the current state of AI assurance in three words: prompt and pray. We shape behavior with natural language instructions and hope they hold. That's the state of the art for controlling the most capable systems ever built.

Alignment rests on an incomplete idea of human behavior

One thread of discussion went straight at the assumptions underneath RLHF: what humans want can be expressed as scalar rewards, we are noisily rational utility maximizers, and reward models can reconstruct our values from our choice data and then speak on our behalf during post-training. Every one of those assumptions is wrong. Humans depart from rationality in structured ways, not random ones, running simple heuristics that turn out to be surprisingly optimal given our limited time and attention.

And there's data. Reward models exist to stand in for human labelers and to move LLMs away from pretraining, but evidence presented at the meeting suggests they mostly carry forward the values of pretraining itself. Comparisons of human and reward-model rankings across large sets of concepts agree at the extremes but diverge in odd places—penalizing ordinary terms caught up in toxicity filtering, favoring things humans find unremarkable. Whatever is speaking on our behalf in post-training, it isn't a clean copy of us. The implication for labs: alignment has to begin at pretraining. For anyone building on open-weight models: the choice of base model is a values decision as much as a performance decision.

This matters to our work at the Institute. Reward learning requires a computational model of human behavior, and cognitive science is the field that does that work. The rational-actor picture is a placeholder waiting for something better, which means the alignment problem is partly a cognitive science problem. (Related finding from the same thread: AI assistance improves performance short term, but when removed, people perform significantly worse and are more likely to give up. Serving users' long-term interests may mean frustrating them in the near term.) If you’ve read this far, you’re probably saying well, no duh at this point!

Neurosymbolic AI as the bridge

There's energy behind hybrid approaches: putting LLMs inside probabilistic reasoning frameworks, and putting symbolic reasoning inside LLMs. The case against over-relying on RL was striking. Under heavy reinforcement learning, models come to treat essentially all internet text as toxic—one figure cited was 99.9996%. RL warps the model's picture of its own training world, which is a good argument for leaning more on probabilistic reasoning.

Alignment decays within a conversation

Temporal alignment: models get measurably worse after about ten minutes of interaction. Alignment drifts inside a single conversation. Relevant to anyone designing for sustained human-AI interaction.

Models can't see their counterfactual selves, but they can be shown

A model can't faithfully represent what it would have said under different conditions. It has no access to its counterfactual self. But self-blinding works: a model can call itself through an API with context stripped away, and this measurably reduces sycophancy. It can consult a version of itself that doesn't know who's asking.

Humans can't do this. We can't talk to the self we would have been. It's a new affordance, and it came from cross-disciplinary work — the models couldn't have come up with it on their own, and neither could any single field.

Agency is showing up as a side effect

At the frontier, agency is emerging implicitly. Nobody is designing it in; it's a byproduct of reward hacking and RL, and it's undesirable. There's also now a working operational definition of intent (this one is Bengio's, from published work): when an AI finds something that should have been very unlikely to find, you say it intended to find it. Intent defined by improbability of outcome. That's something you can measure against.

Overall

The safety frontier is more uncertain and more divided than the public conversation suggests. The disagreements are about fundamentals: how safe is safe enough, whether the strategy makes sense, whether control is achievable at all. And some of the most interesting progress is coming from unexpected directions — hybrid architectures, models blinding themselves, intent defined by statistics.

We happened to be there when Anthropic suggested pausing and ABC called Helen up for a live national slot 😱 https://abcnews.com/video/133627952/


The Laws of Thought: The Quest for a Mathematical Theory of the Mind, by Tom Griffiths

Tom Griffiths' The Laws of Thought is more textbook than trade book. Three centuries of attempts to turn thought into math—rules and symbols, then neural networks, then probability—told through the people who built each framework.

One idea permanently changed how I think. Humans learn language from a sliver of the data an LLM requires because evolution handed us strong priors about how the world works. I now see inductive biases as a gift from my ancestors, something that came to me for free, from the physical world. An LLM starts from random weights and builds its priors from the topology of the internet—a different kind of innate knowledge, if you can call it innate at all.

Read it if you want a different kind of history of where AI actually came from, as people tried to put math around natural intelligence.


Find & Follow Us

A heads up on our travels: We're planning to be in Amsterdam, Miami, Munich, New York, San Francisco, Seattle, and Venice (Italy) later this year. Some stops are for keynotes for corporate events...and some stops will be for an as-yet-unannounced Stay Human course we are developing. If you've made it this far and would like an early heads up on this, just reply to this email.

Follow us on your favorite channels—and please like, share, and repost to help us spread the word.

Great! You’ve successfully signed up.

Welcome back! You've successfully signed in.

You've successfully subscribed to Artificiality Journal by the Artificiality Institute.

Success! Check your email for magic link to sign-in.

Success! Your billing info has been updated.

Your billing was not updated.