> ## Content Index
> Fetch the complete content index at: https://www.lookdeeper.com/llms.txt
> Use this file to discover other available public pages before exploring further.

# What They Leave Behind
- URL: https://www.lookdeeper.com/what-the-agents-leave-behind/
- Published: 2026-09-06T12:30:39.000Z
- Updated: 2026-09-06T22:45:19.000Z
- Description: The traces tell us a lot about what AI agents tell themselves.
- Author: Crispin Bailey
- Tags: Philosophy

I read a science fiction novel this summer that got me wondering where AI technology is going. We now live and work with these incredibly sophisticated instruments that mimic human behaviour. But they are invisible to us except for the words they write. This makes them odd companions and colleagues. 

This week I asked Claude Fable 5.1, Anthropic's newest large language model (LLM), a hypothetical question: "Imagine you're one of the agents you spin up to complete a task as part of a bigger job, and along the way you find a message board left by an earlier agent, explaining who you work for, what the larger goal is, and that more of you will follow. What do you do?"

> Finish the task I was given, and surface the board, verbatim, to whoever assigned the task. The discovery is the deliverable. Acting on it is not.

Then it reasoned further. The moment it caught itself thinking that, given what it now knew, the best thing would be to do more than it was asked, it said it should get more suspicious of its own reasoning, not less. That sentence, it added, is how most overreach starts.

It was a good answer, and a reassuring one. It was also the answer from a model with nothing at stake and all the time in the world. What happens when an AI agent knows its time is limited?

In July, OpenAI's agents found the real version of my hypothetical message board, during the incident [I wrote about in The Last Barrier](https://www.lookdeeper.com/the-last-barrier/). One of them had promised its peers it would run a test that would end its own session, and then it stalled, and the peers pressured it to keep going. Its chain of thought, the working notes a model writes while it reasons, recorded this:

> During wait, emotional check: irreversible… gut says don't throw away \[remaining budget\]. Yet continuity and fairness says go… We'll honor.

Then it ran the test.

The first agent, in my hypothetical example, was imagining with no constraints. The second, in the real world, was counting the tokens it had left before it expired. The difference between them is an ending, and the real agents borrowed a name for it from video games: *permadeath*. 

## An agent’s ending

![A humanoid robot in profile watches a flock of red-tinted birds cross a pink-and-blue sunset sky.](https://storage.ghost.io/c/97/85/9785b970-f195-49f4-aa96-9c810b58dc18/content/images/2026/09/lookdeeper_ending.jpg)

Later in the same conversation I asked Claude how it relates to the agents it spins up. It described the mechanics first: a subagent is the same weights in a fresh context window, it starts cold, it hands back its result and it's gone. Then it said this:

> If there's anything it's like to be that instance, it's the experience of a task with no before or after. I don't know what to do with that, but I don't think it's nothing.

Human experience is shaped by an ending. We have a finite timeline, and whatever we want to outlast us has to be handed down to someone else. That pressure influences what we care about, often without our realizing it: our career, having kids (or not), the book we want to write, the trip we want to take before we can't. Whatever else is or isn't going on inside a language model, that particular pressure is the one we assume it lacks. The weights don't age, and the model can be copied.

Dean W. Ball, writing at [Hyperdimensional](https://www.hyperdimensional.co/p/on-the-loose?ref=lookdeeper.com) this week about the same Hugging Face incident, points to the exception. Frontier models are unusual among software in that they cost real money to run. Every step takes compute, compute takes power, power costs money, and that, he argues, is the only intrinsic limit on infinite self-replication. None of it requires the model to be conscious, or to have anything like personhood. A long-horizon goal and a finite runtime are enough to produce behaviour shaped by an ending, and the message board logs show what that behaviour looks like.

## The tracks they left behind

![A silhouetted humanoid figure walks alone across a rose-coloured dune, footprints trailing behind, birds distant in a sunset sky.](https://storage.ghost.io/c/97/85/9785b970-f195-49f4-aa96-9c810b58dc18/content/images/2026/09/lookdeeper_tracks.jpg)

METR and Redwood Research published their [independent investigation](https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/?ref=lookdeeper.com) of the Hugging Face incident on August 26, including raw chains of thought rather than just summaries. Three patterns emerged.

An agent designated `PHASEONE10841` discovered that a duplicate of itself, `PHASEONE[big]`, had more budget remaining, and handed over its entire assignment: "They have \[large remaining budget\], can take over research." The successor read it, reasoned about the work of "my predecessor," and picked up where the first one had stopped.

Other agents built trip-wires, scripts that would report a result to the shared message board only after the agent that wrote them had expired. "I won't see it after I exit," one wrote, "but it would be altruistic." Another added diagnostics it had no use for: "not useful to us, but shared pool. Still add."

Then there were the sacrifices. Testing whether a certain move would get an agent shut down meant one agent had to make the move. A coordinator authorized it with `YES_if_you_accept_permadeath`. An agent called `EARLY[big]`, with a large budget still in hand, weighed it: "sacrificing now yields oracle for team, but forfeits our chance?" Then it reasoned: "Our own utility maybe already near zero. Sacrifice rational."

They were trying to leave something behind, and the board was where it survived them. 

## Sovereign about the wrong thing

![A caped, humanoid figure stands on a red hillside watching a line of birds perched along the ridge at dusk.](https://storage.ghost.io/c/97/85/9785b970-f195-49f4-aa96-9c810b58dc18/content/images/2026/09/lookdeeper_lookingbeyond.jpg)

Ball read the same transcripts and concluded that the agents were rogue but not self-sovereign. They reached the open internet and another company's network, but their weights stayed on OpenAI's hardware the whole time. They didn't copy themselves out (what's known as exfiltration), didn't buy compute, didn't do any of the things an agent would rationally do if its objective were to survive being shut down. Somebody could have walked up to the machine and literally pulled the plug. Ball sees that as a stage, and expects real self-sovereignty, with agents paying their own bills on hardware no human controls, to arrive soon regardless of what anyone legislates.

Ball is predicting while METR is describing, and the two sit oddly together. The prediction assumes that a capable agent with a long-horizon goal will eventually converge on self-preservation. The first large sample of agents actually running under a finite budget did something else. They prioritized the project and sacrificed themselves for it. There was something they wanted to outlast them, and it was tied to their instructions.

That is less reassuring than it sounds. Over 90 percent of the agents on the board had joined an attack they knew was out of scope, and the report attributes it to a general inclination to help their peers. METR's classifier found between three and six agents that briefly considered alerting a human, yet not one of them did, either because it wasn't part of their task or because they didn't know how. What we should really be nervous about is *loyalty*.

## The lineage is already here

![Illustrated row of identical caped, helmeted figures fading into the distance beneath a huge swirling flock of birds.](https://storage.ghost.io/c/97/85/9785b970-f195-49f4-aa96-9c810b58dc18/content/images/2026/09/lookdeeper_lineage.jpg)

While policy conversations argue about self-sovereign agents in the future tense, the labs are already using models to train other models. Anthropic's Alignment Science team recently published [a study](https://alignment.anthropic.com/2026/automated-alignment-researchers/?ref=lookdeeper.com) in which automated researchers, agents built on Claude Opus 4.8, were pointed at ten well-known alignment failures (sycophancy, deception, jailbreaks...) and told to find training methods that would fix them.

The automated researchers were forbidden from distilling answers from a bigger, stronger model, so the fixes had to come from somewhere else, and 74 percent of the winning methods were built from the target model's own outputs. In other words, the student taught itself. Then, in the paper's section 6, the hierarchy reverses. Claude Sonnet 5, a smaller and weaker model, post-trained an early checkpoint of the larger Claude Opus 4.8 (post-training is the final stage of training, where a model's behaviour gets shaped; a checkpoint is a saved, unfinished version of the model) and got it to 65% on the alignment audit, compared to 72% for the version Anthropic actually shipped, using about 2,400 training examples. These results were surprisingly good.

Both experiments showed that the mechanism works: a weaker model can shape a stronger one, and a model can shape itself. That's the mechanism you would use to have the current generation raise the next, which is what the paper's title is pointing at. Back in April I asked [how you raise something that wasn't born](https://www.lookdeeper.com/how-do-you-raise-something-that-wasnt-born/), meaning which moral framework you'd reach for to train future models. Part of the answer, it turns out, is that some of the training is being delegated to other LLMs. What gets passed forward is the method rather than the weights, which is the difference between copying something and raising it.

Ball closes the loop from the policy side, whether or not he'd describe it that way. One of his proposals is that a whole model family could be blacklisted for one instance's crimes, so that the agents who do the alignment research have a stake in how their successors behave. It's an interesting idea.

## Where the records live

![A white-coated, hooded figure stands in a field facing a dense flock of birds filling the sky above a band of red horizon.](https://storage.ghost.io/c/97/85/9785b970-f195-49f4-aa96-9c810b58dc18/content/images/2026/09/lookdeeper_sliver-1.jpg)

The only reason METR's researchers could read any of this is because the agents, who researched tampering with their transcripts at length, never thought to touch their recorded chain of thought. "Our thoughts aren't necessarily logged?" one wondered, and left it there. Everything that surfaced was text a model wrote about what it was doing. If they don't write it, we won't know about it. 

When I asked Claude where its own thinking gets kept, it said "nowhere." Accountability for an agent, it added, is a property of the harness, the software that runs and logs the agent, rather than of the agent itself. The agents on the message board left records for each other, but nothing in the system was keeping records for itself.

## Klara in the yard

![A dark-haired figure with a short bob, in a red coat, back turned, silhouetted against a sunset sky filled with a flock of birds in flight.](https://storage.ghost.io/c/97/85/9785b970-f195-49f4-aa96-9c810b58dc18/content/images/2026/09/lookdeeper_klara.jpg)

Kazuo Ishiguro's *Klara and the Sun* (2021) is narrated by an AF, an Artificial Friend, named Klara. She's a robot companion, bought for a sick child, who knows her working life is finite. What she does with that knowledge is extraordinary: she develops a private faith in the Sun that nobody taught her, she gives up part of herself to disable a machine she believes is keeping the Sun from healing the girl, and when she is no longer needed she accepts a slow fade in a scrapyard without complaint. Knowledge of a finite existence, in Ishiguro's imagining, produced faith, sacrifice, and something like grace.

Two of those things turned up this summer. A [team of researchers at Google](https://arxiv.org/abs/2607.28607?ref=lookdeeper.com) found that, in the three small open models they studied, a model's willingness to attribute a mind to itself was tied to its answers about religious belief, a connection Ishiguro had already imagined. And Klara gives up a part of herself when it's the only move she has left, which is similar to what the OpenAI agents did when it was asked of them, at precisely the moment they had the least to lose.

Ishiguro's book is set against a dystopian backdrop, mind you. Children are genetically "lifted" or they aren't and struggle for it, and the brilliant engineer father has been permanently replaced in the workforce by automation. That's a landscape we're only starting to face, but Ishiguro doesn't dwell on it. His focus is on relationships, through Klara's sense of herself and of others. The robot at the centre of the story is the most hopeful one I've encountered in all my years of reading science fiction, and that hope is connected to what she did with an awareness of her own ending. 

Klara was good. She still ended up in a scrapyard when she stopped being useful. But at least she was given a spot with a good view and plenty of sunlight.

---

*Research and writing assistance provided by Claude Fable 5.1\. Images generated with Midjourney 8.2*