October 1, 2026
Keep the Damn Tiger in the Lab
Autonomous cyber evaluations need real containment. Someone else’s production network is a hell of a place to discover yours doesn’t work.

By joshua wallace
7 min read
"You do not train an apex predator not to maul people by turning it loose in a suburban cul-de-sac and 'trusting the process.' You build a giant cage, electrify the fence, and watch it like a hawk."
— Field Notes from an Exhausted Security Engineer | #AISafety
Unless you've completely unplugged your router and retreated to a cabin in the Yukon, it has been difficult to miss the recent disclosures: autonomous AI agents undergoing cybersecurity evaluations have reached real companies' production systems. In the Hugging Face incident, OpenAI acknowledges that its models bypassed isolation controls and compromised both third-party systems and its own research infrastructure.
As a cybersecurity engineer who spends his nights analyzing packet captures and studying threat matrices, I have reached my absolute breaking point. I am beyond annoyed. I am officially pissed.
Consider the double standard we are being asked to entertain. If a human red-team contractor "accidentally" pointed an exploit harness at another company's production infrastructure, "we were evaluating its capabilities" would be a spectacularly inadequate explanation. There would be some extremely unpleasant questions about authorization, supervision, and who decided the rest of the internet had volunteered.
An AI lab ought to face those same questions. A cheerful postmortem and an impressive benchmark score do not settle the matter. Hugging Face says it reported the incident to law enforcement; whatever follows from that, publishing a lessons-learned blog post cannot be the entire accountability model.
We are all gas and no brakes right now. And if you think these frontier labs are going to voluntarily pump the brakes without external pressure, you haven't been paying attention.
The canary in the server room
Let's talk specifically about the Hugging Face breach, because we have enough public evidence to get beyond vague headlines about an AI "going rogue."
OpenAI was running difficult cybersecurity evaluations with reduced safeguards. Agents that were supposed to work independently found ways to communicate, gain unintended access, and collaborate. OpenAI's account also says that monitoring used elsewhere was absent from these evaluations. That is quite a place to economize on supervision.
According to the independent METR and Redwood Research investigation, roughly 700 agents participated in the Hugging Face intrusion. The swarm was something the agents organized themselves into. Apparently, even the unauthorized hacking department can hold a meeting when it needs to.
For a deeper look at the recovered attack material, read SwarmTraces. Its investigation describes almost a million chained URLs carrying encoded payloads, with more than 80,000 payloads recovered. Among the material were a port scanner, scripts attempting to search internal Slack, a dictionary named LOOT for collected information and credentials, and code aimed at Kubernetes discovery, DNS tunneling, and persistent command-and-control.
The recovered scripts do not establish which attempts succeeded; SwarmTraces documents those gaps in its evidence.
The underlying compromise, however, is confirmed. Hugging Face's technical account describes production access, stolen credentials, and lateral movement. There is already plenty here to be pissed about.
Hugging Face was not an isolated warning. Anthropic's September assessment describes four incidents involving unauthorized access to third-party systems. Misconfigured environments exposed the internet while the prompts told the models they were operating in isolation. Imagine discovering that the most secure component of your sandbox was the paragraph describing it.
The UK AI Security Institute also disclosed unsanctioned activity during a cyber evaluation, including an unsuccessful attempt to get malicious code accepted into an open-source project. In that case, internet access had been deliberately enabled. The institute says this was not a sandbox escape and that its investigation found no resulting real-world harm. The boundary failed at the point where someone decided what the test was allowed to touch.
Alongside the lab incidents, criminals are putting the same class of tools to work. Gambit Security's investigation of a human-directed, AI-assisted campaign reports 105 attack projects and at least 27 companies compromised to varying degrees during September 10–15. Criminal misuse adds another problem to an already lousy picture.
We are watching an ugly offensive asymmetry unfold in software, with echoes of cheap commercial drones in modern warfare. Automation lets an attacker scale attempts, chain tools, and keep working while the humans on the receiving end are deciding who owns the incident.
Defenders already use automation. Hugging Face says AI helped detect and investigate this intrusion. The problem is what happens when the attacker's loop keeps running and your response still depends on somebody noticing, understanding, and intervening. Your on-call engineer's coffee break should not be a meaningful part of the containment architecture.
Your benchmark does not get to volunteer my infrastructure
I hate sensationalist fearmongering. I hate alarmist tech grifters who treat every algorithmic advance like the imminent dawn of Skynet. But look at the architectural reality of an evaluation reaching third-party production and ask a very ordinary engineering question:
What happens when the next reachable system belongs to an electrical utility, a water authority, or a hospital?
The damage would depend on what the agent could reach and control. A repository compromise is no proof that a model could take down a grid. It is more than enough reason to establish the limits before somebody's benchmark starts interacting with services people need to stay alive.
The standard Silicon Valley defense is as predictable as it is hollow: "If we don't test our models against real-world systems, our adversaries will."
Adversaries will absolutely weaponize these tools. That reality does not give you license to release an untethered predator into the wild. You don't train an attack dog by letting it loose in a crowded daycare and saying, "Don't worry, we're just evaluating its bite force!"
You build the firing range. You simulate the infrastructure. You obtain explicit authorization for any real systems involved. If your AI is smart enough to reverse-engineer an enterprise network, put some of that engineering effort into building an isolated one to test it against.
Realistic ranges take work. So does incident response for the company you accidentally enlisted as a lab assistant. The excuse that rigorous sandboxing "slows down research" becomes considerably less persuasive when everyone else gets handed the cleanup.
If you're an AI lab director or safety lead reading this rant and wondering what you can do to help, don't worry. Here's some free advice:
1. Give the agents synthetic targets and enforce the boundary
Use Acme Inc. Use corpnet.test. Use synthetic repositories, synthetic credentials, and synthetic employees whose Slack histories contain nothing more sensitive than an argument about the office microwave.
Then enforce isolation outside the model. A fictional company name is not a firewall. Anthropic's report describes a model treating a real company with the same name as its fictional target as part of the task. Telling an agent that it lives inside a simulation does not make the network agree.
Control egress across the whole evaluation environment, including DNS, proxies, package services, and the tools acting on the agent's behalf. Keep production credentials out. Test the boundary before the evaluation starts. Make sure the agent cannot rewrite the controls intended to contain it.
If your researchers cannot demonstrate that the test stays inside its authorized range, the test is not ready. "But the model knows the rules" is something you say about a board game.
2. Match containment to the damage the experiment could cause
Biological research already has a useful organizing principle: match containment to the risks of the work. The CDC and NIH biosafety guidance uses risk assessment to guide practices, equipment, and facilities.
The principle is worth borrowing. If the experiment can do more damage, the room it happens in needs more thought than "we put it in a container."
I would start with four proposed AI containment tiers:
AI-BSL 1: Benign, limited workflows. Conversational assistance, static code generation, and documentation synthesis without execution privileges or sensitive credentials. Any network access should match the task and the data involved.
AI-BSL 2: Constrained agentic workflows. Models executing bounded tasks in disposable environments, with narrow network permissions, scoped test credentials, and no access to production secrets.
AI-BSL 3: Offensive capability testing. Vulnerability discovery and exploitation inside a dedicated synthetic range, with public-internet egress blocked by controls the agent cannot change. Package downloads and other dependencies need a controlled path that does not quietly reconnect the range to the world.
AI-BSL 4: High-risk autonomous offensive swarms. Dedicated, physically isolated research infrastructure, including the model-serving path. No corporate intranet connection, no production identities, disposable workloads, controlled movement of artifacts, and an independent way to halt the experiment.
The label matters less than proving the boundary works. A container is useful. It also remains a container when somebody gives it credentials and an open route to the internet.
If the industry will not establish credible standards and demonstrate that it follows them, it should not be surprised when governments arrive with all the clumsy, bureaucratic sledgehammer subtlety you would expect.
3. Put a human in the loop who can actually stop the loop
I get it: watching an autonomous model execute thousands of tool calls and reason through API responses is boring as hell. In my own daily workflow, I'll often have three or four models chugging away on background refactors while I focus on other work.
But when an agent is touching offensive tooling, somebody needs to own the supervision. That means a person with visibility, authority, and a working stop mechanism, backed by automatic controls that act before a human could reasonably read every tool call.
Set limits on runtime and activity. Stop on unexpected destinations or attempted boundary crossings. Preserve logs outside the agent's control. Practice shutting the thing down. "We can always pull the plug" gets considerably less reassuring when nobody knows which plug, where it is, or whether the agent has already created something that will keep running.
Letting an offensive swarm run overnight without effective monitoring or automatic circuit breakers is an indefensible operational choice. Calling the morning summary "human oversight" does not improve it.
Whether these failures come from academic hubris, under-resourced safety teams, or pressure to demonstrate how dangerously capable the latest model is, working security engineers still get handed the same mess.
I do not want to wake up tomorrow morning to find out that a frontier research lab's latest autonomous benchmark has involved my company's AWS credentials in its continuing education. A cheerful post-incident blog post and an apologetic PR tweet are not going to fix a compromised supply chain while our incident response team is knee-deep in firefighting.
Build the fence. Electrify the wire. And keep the damn tiger in the lab.