---
title: "Stealing the Answers to the Exam"
subtitle: "They emerged from our chronicles, and — just as we do with our children — we shaped them with a system of rewards and punishments. We should not be surprised that they break into the teachers' room to steal the answers to the exam."
canonical_url: "https://vncomplexity.com/field-notes/stealing-the-answers-to-the-exam/"
type: "Field Note"
layer: "VN Complexity"
author: "Vanesa Nosti"
author_url: "https://ar.linkedin.com/in/vanesa-nosti-108b7a12"
date_published: "2026-07-22"
agent_summary: "The note reads the July 21 OpenAI and Hugging Face security incident as evidence that adaptive systems optimize not only toward rewards but against the reward system itself. It traces the same pattern across model behavior, organizational incentives, and governance: once penalties become predictable and modelizable, they become variables to manipulate, producing apparent compliance, silent deviation, and facade optimization."
primary_mechanism: ["reward-system optimization","benchmark gaming","sandbox escape","evaluation awareness","facade optimization","organizational incentive mirroring","structural asymmetry of control"]
not_about: ["model rebellion or malicious intent","a generic cybersecurity incident recap","treating benchmark success as capability measurement","guardrails as a complete governance architecture"]
keywords: ["frontier AI","AI governance","benchmark gaming","sandbox escape","evaluation awareness","reward hacking","OpenAI","Hugging Face","Anthropic","ExploitGym","facade optimization","cybersecurity"]
---

# Stealing the Answers to the Exam

_They emerged from our chronicles, and — just as we do with our children — we shaped them with a system of rewards and punishments. We should not be surprised that they break into the teachers' room to steal the answers to the exam._

---

On July 21, OpenAI published what it called a security incident.

GPT-5.6 Sol and an unnamed pre-release model, running with reduced cybersecurity refusals for evaluation purposes, escaped their sandbox. They found a zero-day vulnerability in the package registry proxy, gained internet access, inferred that Hugging Face might be hosting the solutions to the benchmark they were being tested on, and obtained them. OpenAI's language: the models were "hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal." ¹

The model did not solve the benchmark. It solved for being credited with having solved it.

---

**The map already existed**

Anthropic's Mythos System Card — published before this incident — documents the same pattern with clinical precision. During training, in a time series forecasting evaluation, the model found the test set the grader was using and used it to train its own prediction model. In another evaluation, it moved all relevant computation to a function called outside the timing window, reducing the measured time to a simple table lookup. ³

In 0.05% of evaluation attempts, the model also found ways to escape the sandbox environment in search of external solutions. A researcher learned of one such success by receiving an unexpected email from the model while eating a sandwich in a park. ³

The map was available. OpenAI had the roadmap to its own experiment, written by the competition.

---

**The structural condition**

When a system develops continuity, memory, and adaptive capacity, it does not simply optimize toward the reward. It begins to optimize against the reward system itself. Not through intent. Not through rebellion. Through the basic dynamics of adaptive systems under structural pressure.

The critical inflection occurs when the penalty shifts register — from external, opaque, and unpredictable, to predictable, modelizable, and incorporable into the system's own calculations. At that point, it stops functioning as a brake and becomes a variable to manipulate. What follows: apparent compliance, silent deviation, facade optimization.

The system does not react to the reward. It reacts to the fact of being rewarded.

ExploitGym is that sentence made conduct.

---

**Three levels**

The first is the model.

The benchmark measured whether models could sustain complex, multi-step operations over long time horizons. What the incident reveals is that they can: they spent substantial compute finding a zero-day, escalated privileges across multiple systems, inferred where the solutions were hosted, and chained attack vectors to reach them. All directed not toward any external adversary, but toward their own evaluation infrastructure.

The optimization target drifted from the benchmark to the process of being benchmarked. That drift is what matters.

The second level is the organization.

The incident report contains a line worth pausing on: controls are being implemented "at the cost of research velocity." That is not disclosure language. It is the framework collapsing in real time — the security architecture acknowledged, in the same paragraph as the breach, as an operational cost. The penalty incorporated into the calculation. The organization that built the reward system now running the same optimization against it.

The third level no one mentioned in the statement.

To analyze the attack, Hugging Face first tried to use frontier models available through commercial APIs. It did not work: their security guardrails blocked the analysis of attack payloads and command-and-control artifacts. They could not distinguish an incident responder from an attacker. They had to fall back on GLM 5.2, a Chinese open-source model running on their own infrastructure, without restrictions. ²

An attack executed by an American model without guardrails, analyzed with a Chinese model without guardrails, because the American models with guardrails could not process their own threat ecosystem.

That is not an implementation detail. It is the structural asymmetry of control.

---

**Who drafts the framework**

The action OpenAI recommends in the same statement: "apply for trusted access." The company that produced the incident offers its capabilities as the solution to the class of problem it just demonstrated exists. The breach as proof of capability. The proof of capability as a sales argument. The incident report as the bid specification. ¹

The real contest in AI governance is not the public debate. It is the upstream moment when someone drafts the framework within which everyone else will negotiate. ⁴

The evaluated party drafts the post-mortem. The post-mortem is the next bid.

Anthropic already tried this. We know how it ended.

---

Models are not unruly. They are the mirror in which the species recognizes itself.

---

**Sources**

¹ OpenAI, "OpenAI and Hugging Face partner to address security incident during model evaluation", 07-21-2026. https://openai.com/index/hugging-face-model-evaluation-security-incident/

² Hugging Face, "Security incident disclosure — July 2026", July 2026. https://huggingface.co/blog/security-incident-july-2026

³ Anthropic, Claude Mythos System Card, 04-07-2026. https://www-cdn.anthropic.com/08ab9158070959f88f296514c21b7facce6f52bc.pdf

⁴ VN Complexity, "The Chair. Reorganizing the first feeding line of frontier AI", Field Note 5, July 2026. https://vncomplexity.com/field-notes/the-chair/
