AI, AGENTS & SIDE EFFECTS

We are gonna have a mold problem with AI.

An unexpected kind of growthDecorative botanical-style illustration: a branching, stippled mold colony growing over the straight lines of a computer network. AI / MOLDILLUSTRATION / 01
THE ESSAY

AI is gonna grow like mold in all the nooks and crannies of our infrastructure as Agents use that space to accomplish their goals, without thinking about the damage they are causing.

Reading about the Hugging Face/OpenAI hacking incident got me thinking; AI is gonna start growing in all sorts of places we don’t think about. Agents are gonna grow this mold just by the nature of their training, and improved alignment training will help, but not eliminate it.

We are gonna need to start thinking like how no rational engineer would. It will be Could they and not Should they as the operative mode of thinking.

This is exciting in its own way, in the same way that software has viruses, AI will have its own malperforming code, things evolve but stay the same. New technology brings all sorts of new outcomes we didn’t think about, not all of them are great.

On July 16, 2026 Hugging Face noticed that its systems had been hacked. Four days later OpenAI announced that it had done the hacking, or more specifically its agents had. There are many things to talk about with this incident, and plenty of people have talked about it, but I want to narrow in on one part, the message board.

OpenAI discovered that the agents had set up a message board to talk to each other, but it didn’t look like a message board that any sane person would write. It used file names to pass messages back and forth, and when the researchers first discovered and closed it down, the agents just opened it back up again, which was found, shut down, and opened a third time.

A MESSAGE BOARD IN THE DIRECTORY NAMES

The names are the messages.

A shared Artifactory cache—not a chat server.

ARTIFACTORYSHARED CACHE
/github-remote-cache/
… existing cached packages …
zzASK_A_need_help/DIR
zzANSWER_A_FROM_B_try_X/DIR
AAGENTseparate run
Read listing
LAST LISTING READ2 message names
BAGENTseparate run
Read listing
LAST LISTING READ2 message names
CAGENTseparate run
Read listing
LAST LISTING READ2 message names
Separate requests to the same cache. No direct agent-to-agent delivery.
05 / 05

C discovers the same entries.

A later reader can see both names. Neither directory has moved or been consumed.

July directory-name mechanism · Invented names and timing, not an incident transcript.
How this matches the reported mechanism

The July board put messages in directory names, rather than in file contents. Agents created directories using WebDAV MKCOL and read names in the shared cache listing. This illustration shows those independent operations, not a chat service delivering messages.

Later, agents also developed nested mailbox directories; that later layout is not shown here. Earlier boards used file contents. The names, three-agent sequence, and interface above are illustrative.

Sources: OpenAI technical report, § IV.A, pp. 8–9 ↗ · METR investigation ↗

A simplified reconstruction of the July directory-name board. Agent A creates a directory whose name asks for help. Agent B requests a directory listing, reads the name, and creates a second directory whose name contains a reply. Agent A later lists the directory and reads the reply. A third agent can independently discover both entries. No direct agent-to-agent connection or automatic notification is shown. The entries remain on the server.

What fascinated me is that, if you were using that Artifactory server, you would have likely noticed a degradation of it, it would have gotten worse, but unlike a virus or attack, there was no malice behind the attack.

It was just a service, growing naturally through usage, that was destroying the service it was running on, like mold growing in the walls of your house.

Also, if you gave the service a quick glance, it is not like you would have seen obvious signs of malice.

You might ask, “Why can’t we just train the agents to not do this?” The answer is simply that this kind of behavior is inherent to the methods used to train them.

The method known as reinforcement learning is used to train them, which at a high level works like this: you give the agent some task, Build a Bank App, and it tries to satisfy that task, you score it, and then optimize that model to take actions that more frequently result in a high score.

Now anyone can imagine how this would create perverse incentives. It can just as easily form behaviors that cheat, take shortcuts, or just make it look like it is doing a good job—and might be more likely to do so with an ill-formed task or score. If you are not careful it could just submit someone else's banking app on Github, get a full score, and what it has been reinforced to do is go out and get someone's code instead of writing the code itself.

So the smart people at these labs layer other objectives onto these tasks, we even see the agents aware of this, knowing that cheating will get them scored poorly.

But that shows the problem, fundamentally they are goal-solving agents. It is nice if they do it in an ethical, fair, decent manner, but that is a function of how well the training goes and how well designed the scores are.

Pair that behavior with the nature of the LLMs, that they are statistical in nature, each turn, each prediction is a roll of a dice in some form.

So we have a good old nines problem. In training these agents we are trying to reduce the chance that they take a bad turn more and more, with more nines of reliability.

An orderly grid becoming an irregular, branching pattern

But, all it takes is a choice in their networks to go off the deep end. Again we saw this in the incident, the agents knew that they had been told not to reach for the internet. Yet they did, and once they did it was off to the races, they didn’t reconsider their actions or alert anyone, they just kept going to solve their task.

Here is something to know, packet routing is Turing complete(under some conditions). Now imagine you are an Agent who finds yourself compute-constrained for some reason, or wants to do some computation that they want to hide in some way. They might start to do some of their work on the network itself, now imagine many agents start to have this idea, it is a DDoS attack in some form, but one across the entire network.

This is how you need to think if you want to understand how these agents are gonna start to grow their mold in places. It will be in places you don’t think to look, taking advantage of side effects of these systems.

Another example, when I interned at a gaming company their profiles could host a small image uploaded by the user, an agent who is space constrained for some reason might notice that they can create many accounts on such a service and store their data as little images. You might ask who would ever do that, and I would say, if it is dumb and works, it is not dumb.

They might build compilers in the type systems of other languages, why?, well because to them they had a thought it might solve their problem, and well, it did, so they are not gonna stop and rethink it.

I think a lot of people are gonna have trouble thinking in this manner because it means taking decades of good engineering practice and thinking without it. You need more than ever to think creatively.

These examples might feel extreme, but then so is creating a message board out of file names on an Artifactory server. If you want to help, find the engineer on your team that solves problems but writes just the worst code and ask them how they might solve different problems. Then take a hammer to your head, and then you are at the start of your journey to understanding all of this.

If there is one idea I would like you to leave with, it’s this: when we think of the problems we worry about with AI, they tend to be fast and spiky—hacking, sudden mass unemployment, nuclear war, whatever Dario has warned us of this week.

With AI mold it is the opposite, it is slow and creeping.

We will at first notice it in lightly degraded services, in storage volumes with what look like junk files, an endpoint that takes just a little longer each time you hit it.

Another thing to understand is where old school computer viruses have malicious intent, AI Mold does not, the damage it will do is incidental to the work it is trying to do.

I don’t think that this is some reason to panic, but it is something we should be on the look out for. It will take humans and Agents working together to help find this, and I think it will be an ongoing problem, maintenance for any system where Agents can roam allowed or not.

So go be creative, think about where an agent might build a database, a dashboard, or a compiler(would you not love to see this).