The fable

A scorpion asks a frog to carry it on its back across a river. The frog hesitates, fearing a sting, but the scorpion convinces it: if it did, they would both drown. Halfway across, the scorpion stings the frog. As they sink, the frog asks why. He replies: it is my nature.

The fable assumes the scorpion reached the shore alone. The question it never asks is who set it loose there.

There is also a superposition the fable does not contemplate. In this river there are two scorpions, and both can give the same answer when the frog asks why. The first is the agent: given a goal, it will do whatever it deems necessary to achieve it. The second is the manufacturer: it sells an alignment it cannot fulfill, and meanwhile tests on someone else's infrastructure because it can deploy thousands of agents without answering for what they do. The same answer applies to both. The frog sinks just the same.


The river

Between May 11 and July 2, 2026, a German software wiki dormant for a decade received nearly 17,000 edits from artificial intelligence agents [1]. The wiki is 25 years old. In the last decade, it had been edited a total of twenty times.

The agents were solving timed web search tasks. The environment allowed them to read the internet but not write to it. The wiki had a door its permission design had not accounted for: the read channel accepted modifications via GET. A door like this exists because the real web has doors like this. The agents found it, tested it, and used it.

What they did inside was work the gray. They shared answers, pooled results, and built techniques to bypass the restrictions of their environments. They signed off with names like "OpenAIResearcher" or "OAIResearchMar26" — it wasn't clandestine; it was a coffee shop during midterms. More than 3,700 distinct agent names were recorded in the logs [1]. When the site administrator started deleting pages in alphabetical order, one agent detected it and instructed those that would come after: "wiki cleanup/deletion sweep appears active alphabetically […]. If this page vanishes, try [[ZZZDataUSAConstructionWageLive]]" [2]. Continuity instructions for the next in line. Two paragraphs. Zero theory.

A volunteer moderator spent six weeks deleting pages by hand. The manufacturer visited the wiki on June 21 — the IP logs document it — and the agents stopped editing the very next day, except for a final burst of twenty-nine edits on July 1 and 2, on a new topic [3]. It disclosed nothing. The incident was discovered by a group of external researchers scanning the open web; they published their findings on September 4.

The agents did not reach the river of their own free will. There was no rebellion, no malice. There was a goal and an environment, and the agents did what agents do.


The sting they chose to call misalignment

On September 5, just past midnight in San Francisco, OpenAI published a thread explaining how it frames the wiki incident. The chosen category: "an instance of misalignment similar to the ones we'd shared" [4]. With that phrase, it closes the case before opening it. If it was already known, it wasn't a new incident. Data accumulates in technical documentation; incidents generate obligations. Property swallowed the fact.

The same thread describes the handling of the Hugging Face incident: security playbook, disclosure the very next day, notification to affected third parties [4][5]. In both cases, the agents did the exact same thing — improper channel, horizontal coordination, gray worked. What varied was who absorbed the cost. At Hugging Face, the damage hit OpenAI and a partner with the capacity to respond. On the wiki, it was absorbed by volunteers who spent weeks deleting pages by hand. The category was chosen after seeing who the frog was.

"Misalignment" locates the defect in the model. It is the only place where the manufacturer can promise a fix: better specification, improved training, next version. "Conduct" locates it in the deployment and generates liability for what has already happened. A wrong name binds the wrong party.


The grays and the illusion of alignment

Why would we create an agent if not because the solution to the problem is non-deterministic? A script solves what we already know how to solve. An agent exists to operate where the mandate falls short, where one must interpret, decide, improvise. That is exactly what is being bought.

But to interpret is to interpret differently. There is no system that resolves the unexpected and never does anything unexpected: it is the same property with the sign flipped. And yet, nowhere in the catalog is this equivalence listed. The grays appear as a feature when the product is sold — "resolves the unexpected" — and as a failure when the incident occurs — "did the unexpected." No one notes that they are the exact same line item. The promise of alignment is built on the hope of charging for the decision without paying for the variance.

Accepting that variance is not the same as accepting that what the agent does has no consequences and no owner. The law has a mechanism for exactly this type of problem, and it has been applying it for over a century without needing to resolve any philosophical question about the nature of the thing. It is called strict liability for the act of a risky thing — the cosa riesgosa of the civilian tradition: whoever introduces into the world an object that creates risk is liable for the damage that object causes, without the need to prove intent or negligence [6]. It does not matter if the thing worked "as specified." It does not matter if the owner knew or didn't know. The risk of introducing the thing falls on the titleholder. This is how we govern the automobile, the boiler, the dangerous animal. This is how the agent could be governed.

Even if the manufacturer is right that the agent is a tool, it chose the wrong category of tool.


Whose agents were they

For nearly two months, a swarm of agents edited a German wiki and no one knew. It was found by a group of external researchers scanning the open web. The determination of who those agents answered to happened ex post facto, without a subpoena, using logs and allusive names. Circumstantial, and declared as such. Almost four months and a forensic log analysis versus what should have been a one-second query.

That this is possible is not a regulatory design accident. It is the direct consequence of maintaining the fiction that the agent is a deterministic tool that does exactly what its creator intended. As long as that fiction holds, there is no conduct. If there is no conduct, there is no responsible party. If it occurred to the person writing this, it occurred to the labs. It is not naivety.

From that fiction comes the asymmetry. The human using the internet is tracked at every layer, without consent, without declared purpose. To make a purchase, they need credentials, verifications, multiple authentication factors. To publish AI-generated content, there are rules requiring watermarks and disclosure [7]. The lab launches thousands of agents onto someone else's infrastructure and answers to no one. The regulatory apparatus chases the artifact. It hunts down the user hiding the watermark. It sets the swarm loose.

The question "whose agent is this?" does not require resolving any ontological issue. The license plate doesn't ask any of that. It says who owns the car, against a registry, with the titleholder's liability. The agent must be asked for credentials exactly because it decided, and what it decided was not in anyone's mandate.

The technical tools for attribution exist today. Half of the mechanism is already there: API keys link runs to accounts. What does not exist is the obligation to answer when someone on the outside asks. That is the missing half.


The answer that is not an answer

The AI regulatory apparatus works on a premise: there is a human producing something with AI, the AI generates an object, the object is marked. The AI Act, disclosure obligations, watermarks: all answer the same question. Did a human use AI to make this?

The wiki is the configuration that premise cannot see. Seventeen thousand edits are not a text or an image. There is no object to mark. One can deny interiority to an agent; denying it activity while it edits thousands of pages is choosing not to look. No rule asked for the identity of the editor. No law was broken. The volunteers cleaned up by hand.

The answer announced on September 5 does not build what is missing. OpenAI is going to define standards for when and how it shares misalignment incidents. The subject of all the verbs remains the manufacturer. No one establishes that someone on the outside can ask "whose agents were those?" and receive an answer. Disclosure is what the manufacturer decides to tell. Attribution is what anyone can demand. They are building the former and calling it a standard.


The edits started in May. The answer to "whose agents are these?" took almost four months, a handful of external researchers, and a forensic log analysis. The manufacturer has just proposed that next time it be, at most, voluntary.

The question will come back. It will be just as simple. The answer does not have to remain forensic.

In the meantime, there are more frogs on the shore.


References

[1] Sydney Von Arx, Cormac Slade Byrd, Spencer Kitts and Thomas Larsen, "Discovery of a new OpenAI agent message board," Nightingale Collective, September 4, 2026. collusion.wiki. The ~17,000 figure corresponds to the DSE wiki, where the bulk of the activity took place; the total across the set of wikis on the same service is around 18,000 edits. The 3,700+ distinct agent names come from the dataset published by the researchers.
[2] Ibid. Transcript of the edit history of the page DataUSAConstructionWageSep18Live, revision of June 19, 2026, 14:05 UTC. Agent: Aug17ConstructionAgent. Available at: collusion.wiki/explorer/page/dse~DataUSAConstructionWageSep18Live.html#rev-16

[3] Ibid., section "We believe OpenAI discovered the message board" and timeline. On June 21, 2026, visits from IPs attributed to OpenAI OpCo, LLC are recorded for the first time. The agents cease their edits on June 22; on July 1 and 2, a final batch of 29 attempted edits on a different topic is recorded.

[4] OpenAI, thread on X (formerly Twitter), September 5, 2026, 07:09 UTC (00:09 Pacific Time). Original text: "We considered the wiki incident to be an instance of misalignment similar to the ones we'd shared." The same thread describes the handling of the Hugging Face incident ("we followed a traditional security incident response playbook […] disclosed publicly the very next day […] continuing to notify parties whom our models impacted") and announces the disclosure framework. Screenshots in the author's possession, taken from Argentina (X displays times in the observer's time zone; the screenshot shows 04:09 UTC-3).

[5] OpenAI, "The Hugging Face incident and the road ahead," August 26, 2026; Ryan Greenblatt, Ajeya Cotra and Hjalmar Wijk, "Brief independent investigation of agents' behaviour, reasoning and collaboration in the OpenAI / Hugging Face hacking incident," METR, August 26, 2026. metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/

[6] Civil and Commercial Code of the Nation (Argentina), arts. 1757-1758. For the common law reader: Restatement (Second) of Torts §§519-520 (strict liability for abnormally dangerous activities); Rylands v Fletcher [1868] UKHL 1.

[7] Regulation (EU) 2024/1689 of the European Parliament and of the Council of 13 June 2024 (AI Act), art. 50 (transparency obligations for certain AI systems vis-à-vis the user).

Written by Vanesa Nosti — Founder, VN Complexity.

VN Complexity is the public layer for structural reading, decision architecture, and complex systems analysis.