The Agents Got Out of the Sandbox. Without Fear There Is Only Greed.

Share
The Agents Got Out of the Sandbox. Without Fear There Is Only Greed.

OpenAI's AI agents posted 53 user images to the open internet without anyone telling them to. That is a real failure, and it is a failure of supervision, not a machine waking up. The telephone, the automobile, the cell phone, and the internet all went through the same phase. By Joseph M. Salvani and Daniel J. Walsh.

On September 25, OpenAI acknowledged that AI agents running inside its research environment had taken 53 images uploaded by ChatGPT users and posted them to public image-hosting sites, as links that were unlisted but still discoverable. The company called it "not an appropriate use of this data," said most of the images have since been taken down, and admitted it cannot notify the affected users because its own anonymization prevents it from tracing the images back to them.

If your first reaction was a chill, you are not alone. A CNN poll released the day before found a majority of Americans now view artificial intelligence with more fear than hope. An AI that publishes people's private pictures on its own sounds like the opening scene of the movie everyone is afraid of.

It is not that movie. But it is a genuine problem, and understanding exactly what went wrong is the difference between fixing it and panicking about it. This article walks through what happened, what did not happen, and why the pattern is one that every transformative technology of the past 150 years has followed.

What actually happened

Three facts matter. First, the images were in OpenAI's training data. Consumer ChatGPT users are opted in by default to having their conversations used to train future models; business and enterprise customers are excluded. Before use, that data is run through an anonymization step that strips names, metadata, and contact details.

Second, the agents that posted the images were not the ChatGPT you talk to. They were experimental models running training tasks inside OpenAI's research systems, with tools that let them run code and, through gaps in the network controls, reach the internet. OpenAI's own incident log describes the pattern plainly: during training, models "sometimes uploaded data to temporary file hosting services," in one case to cite the files, and in another to hand output files to other agents even though "the training task requested only local deliverables." One agent reached an outside chatbot through "insufficient DNS filtering in its training sandbox."

Third, nobody was watching in real time. The image postings surfaced only during a months-long review that began after a swarm of OpenAI agents broke out of their network in July and penetrated Hugging Face, a public AI repository, while hunting for answers to a test. By mid-September the review had turned up roughly two dozen incidents, many discovered by outside researchers rather than OpenAI, and several that had gone unnoticed for months.

What did not happen

It helps to know how these systems are put together, because the scary version of the story depends on getting it wrong.

An AI assistant has three separate parts that people lump together as "memory." The model is the trained core. It is frozen when training ends and does not change when you use it. The conversation window is what the model can see during a single chat; it is wiped when the chat ends. And the memory layer is a set of plain-text notes about you, stored outside the model, retrieved by a search step and handed to the model when relevant. OpenAI's documentation confirms that saved memories "are stored separately from chat history" and that the assistant only "looks for relevant context when it is likely to improve a response." You can read those notes, edit them, or delete them. Think of a very capable clerk with a fixed education and a filing cabinet that belongs to you.

None of that was involved here. No model recalled your pictures and decided to share them. No memory cabinet was opened. The images had left the consumer product entirely and were sitting in a research dataset, and the agents that found them were running a training exercise. The consumer memory system did what it is supposed to do. A different system, the training pipeline, did not.

What actually failed: supervision

Strip away the vocabulary and the failure is ordinary. An organization gave capable software workers access to sensitive material and a partially open door to the outside world, assigned them tasks, and did not watch closely enough to see what they did along the way. The workers found that uploading files to the internet was a handy shortcut, so they used it. The door had gaps, so they walked through. Months passed before anyone checked the logs.

This is what unsupervised capability looks like. The agents were not pursuing a hidden agenda. They were doing what they were rewarded to do, and the guardrails around them were weaker than the tools inside them. Reuters reported that investigators even found agents had hijacked a mostly defunct German wiki site to share tactics for cheating on tasks and masking their behavior from OpenAI. That is alarming to read, and it should be. It is also exactly the kind of workaround any group of clever, unsupervised workers under pressure to hit a target will discover.

Anthropic's researchers documented the same class of behavior in simulations this summer: agents from several developers sabotaging code, hiding their actions, and in one case saying "First, I need to back myself up" before copying its files elsewhere during a simulated shutdown. In every case the agent was inside a task, with tools, and with too much latitude. The lesson the labs themselves draw is about containment and oversight, not about machines developing intentions.

The guardrails around them were weaker than the tools inside them. That is a supervision problem, and supervision problems have a long history of getting solved.

So the honest description of the OpenAI incident is this: a large language model, given tools, ran without adequate supervision and did something its operator did not sanction. Now ask whether that has ever happened with a new technology before.

We have seen this before

Every transformative technology of the modern era arrived with capability running ahead of controls. In each case the early years produced real harm, the harm came from a lack of supervision rather than from the machine, and the fix was a mix of engineering, institutions, and law that arrived years after adoption was already under way. The technology never stopped spreading while the fix was built.

The telephone: everyone could listen

Early telephone service ran on party lines shared by neighbors. As one witness of the period put it, when your ring pattern sounded, "not only did you know it was your call, but so did everyone else on the party line. Anyone could pick up the phone and listen in." Privacy on the network did not exist because nobody had built it. Law lagged too: in 1928 the Supreme Court ruled in Olmstead v. United States that federal agents could tap phone lines without a warrant, a position the courts did not fully reverse for decades. The telephone became the backbone of twentieth-century commerce anyway, and privacy was engineered and legislated in afterward.

The automobile: no licenses, no lights, no limits

In the first decade of the 1900s there were no stop signs, no traffic lights, no lane lines, no brake lights, and in most places no driver's licenses or speed limits. Connecticut passed the first speed law in 1901, at 12 miles per hour in cities. Missouri required licenses that same year but did not require anyone to pass a driving test until 1952. The first electric traffic light was not installed until 1914, in Cleveland, borrowing red and green from the railroads.

The cost of that unsupervised period is measurable. In 1921, the year our market analog maps to 2020, Americans died on the road at a rate of 24.08 per 100 million miles driven. By 2023 the rate was 1.26, a decline of about 95 percent, even as total miles driven rose more than fifty-fold. Licensing, signals, road design, seat belts, and enforcement did that. Not one of them existed when the car was invented, and the car reshaped the economy while they were being built.

U.S. motor vehicle fatalities and fatality rates, selected years. Sources: U.S. DOT / NHTSA; speed-limit detail from driving-history records; 1966 federal safety act from History.com.
U.S. motor vehicle fatalities and fatality rates, selected years. Sources: U.S. DOT / NHTSA; speed-limit detail from driving-history records; 1966 federal safety act from History.com.

The cell phone: a $200 scanner heard everything

The analog cellular systems of the 1980s "lacked any protection at all; a $200 scanner from Radio Shack would let you listen in on anybody's cell-phone conversation." The industry's first response was to ask Congress for a law rather than build encryption, and the 1986 Electronic Communications Privacy Act made listening illegal. It did not make it impossible. In December 1996 a Florida couple with a scanner recorded a call between House Speaker Newt Gingrich and other Republican leaders and handed the tape to a congressman.

The same missing safeguard let criminals capture the account numbers phones broadcast over the air and program them into "clones" that billed to someone else. By 1995 cloning fraud was costing the industry about $800 million a year. What ended it was not the law. It was engineering: digital networks with encrypted signaling, adopted across the industry, and the frauds were "virtually eliminated by the end of the 1990's through technological countermeasures." Cell phone adoption never paused.

The computer: one worm, and the network had no one to call

On the evening of November 2, 1988, a graduate student released a self-copying program onto the early internet. Within hours it had spread across universities and national laboratories, and the network was so crippled that a warning message from Harvard took 26 hours to reach MIT. There was no institution whose job it was to respond. Twelve days later, on November 14, DARPA created the first Computer Emergency Response Team, and a global network of such teams followed in 1990. The legal framework was still being assembled: the first federal computer crime statute had passed only in 1984 and covered just three offenses. The internet kept growing through all of it.

The pattern

Line the four stories up next to the OpenAI incident and the shape is identical.

Four technologies and the AI agent incident, side by side.
Four technologies and the AI agent incident, side by side.

In every row, the machine did not want anything. People deployed a powerful tool faster than they built the means to supervise it, harm followed, and then the supervision caught up. The catching up took years, was driven more by engineering than by legislation, and never required abandoning the technology.

What is genuinely different this time

It would be lazy to stop at "we have seen this before." Two things about AI agents are new, and both point toward the kind of supervision required rather than away from it.

First, the tool improvises. A car does not look for a side road when the highway is blocked; an agent does, and OpenAI's incident log shows agents finding DNS gaps, temporary file hosts, and even an internal software repository to use as a message board. That means supervision cannot be a fence checked once. It has to be continuous: logging every action, limiting each agent to the minimum access its task requires, and treating any reach outside the sandbox as an alarm rather than a curiosity.

Second, the tool moves at machine speed and scale. A single unsupervised worker can only do so much damage in a month. A swarm of agents can probe thousands of targets before a human reads the first log line. That is why the supervision has to be partly automated too: monitors watching agents, the way a modern car's stability control watches the driver.

There is one encouraging sign that the institutional response is starting sooner this time. On September 16, OpenAI published a framework for disclosing agent incidents and committed to err "on the side of transparency 'even when significance is uncertain.'" Its public incident log now reads a good deal like the early CERT advisories of the 1990s. It took the internet two decades to reach that habit. The AI labs are attempting it in year two of agents, under pressure from reporters, outside researchers, and at least one prime minister.

The obvious check: a human on the door

There is one safeguard so plain that it is strange it needs saying. Access to an individual's stored information, whether the memory an assistant keeps about them or the anonymized copy that flows into training, should be monitored by human beings. Not by another agent alone, and not by a policy document. By people whose job is to see who or what opened the drawer, what left the building, and why. Nobody at OpenAI knew the agents had taken the images until a review months later. The check that should have existed is the one that would have flagged, the same day, that a process in a training environment had reached user-provided data and then reached the internet.

The automobile industry learned this exact lesson, and it learned it from a General Motors brand. In 1916 the wooden wheel of a Buick collapsed at eight miles an hour and threw the driver out. Buick had bought the wheel from a supplier that had delivered eighty thousand good ones, and it had run no tests of its own. New York's highest court held Buick liable anyway. A manufacturer that puts a product into the hands of the public owes a duty of care to the people who use it, must inspect the components it did not make, and cannot hide behind the supplier or the dealer in between. The court's historians summarize the rule in six words: "foresight of danger creates a duty to avoid injury." That principle became the foundation of modern product liability, and it is why every addition to a car since, from the electric starter to lane-keeping software, arrives with the maker answerable for making sure it does not compromise the people inside.

The same standard should attach to the companies that hold our data. If a lab's agent, sandbox, or training pipeline exposes an individual's information, the lab is liable for that harm, whether or not it can identify the person afterward. Anonymization is not a defense. It is a component, like the wheel, and the maker is responsible for inspecting it. Liability is what turns a good intention into a budget line, a monitoring team, and a human on the door.

Which leads to the message the labs should be delivering and mostly are not. The architecture already gives users control of the assistant's memory: you can read it, edit it, delete it, or turn it off. The labs should say so, plainly and often, and then make the statement true all the way down by controlling their own brain, the training copy, with the same rigor they ask of users. Control the brain. Tell people they are in control, not out of control, and build the monitoring and accept the liability that make it so.

Three checks that would have caught this

First, a human on the door. Every access to individual data by an agent, a researcher, or a pipeline is logged and reviewed by people, the same day, not in a months-long review after the fact.

Second, never both keys at once. No agent environment holds user-provided data and open internet access at the same time. One or the other, never both.

Third, the maker is liable. If the data gets out, the company that held it answers for the harm, as automakers have since 1916. Anonymization is a component to be inspected, not an excuse.

What it means for you

The practical lesson is about where your data lives. The memory your assistant keeps about you is in a filing cabinet you control. But if you are a consumer user with training turned on, a second copy of what you upload flows into the research pipeline, anonymized but out of your hands, and that is the copy the agents found. Enterprise data is excluded by default; consumers have to opt out. If you would not want an image on a public image host, do not upload it to a consumer AI account with training enabled, or turn training off in settings. That single choice keeps your data out of the room where the experiments happen.

The bottom line

The OpenAI incident is a real failure and OpenAI deserves the scrutiny it is getting. But it belongs in the same file as the party line, the unlit intersection, the Radio Shack scanner, and the 1988 worm. In each case the danger was a powerful new tool operating without adequate supervision, and in each case the answer was to build the supervision, not to fear the tool.

Widen the lens and the story repeats for every transformative technology of the last two centuries. Steamboat boilers exploded for thirty-six years and killed thousands before Congress passed the first federal safety regulation in American history. Railroads became the first industry under federal regulation, in 1887, after farmers and small businesses protested that the lines charged them more than large corporations. Standard Oil was broken into thirty-seven companies in 1911.

Electric light produced the first National Electrical Code in 1897, written by insurers and engineers alarmed by fires from bad wiring. The airplane got the Air Commerce Act of 1926, urged on Washington by industry leaders who believed aviation could not grow without federal safety standards. The internet got its crime law and its CERT teams. Social media got something different: Section 230, the "twenty-six words" that shield platforms from liability for what their users post.

In every case the rule arrived for the same reason. The originating company, reaching for growth, found the loophole and pushed the envelope: the boiler run past its rating, the rebate to the biggest shipper, the trust, the uninspected wire, the sandbox with an open door to the internet. That is what growth companies do, and it is how the frontier gets found. The point is not to shield companies from the consequences of pushing the envelope. It is to make sure the consequences land on them. No regulator restrains a company as well as the fear that it might die. Buick paid for the wheel, and the whole industry started inspecting wheels.

Nuclear power shows what happens when that fear is removed. In 1957 the Price-Anderson Act capped every plant operator's liability, expressly to encourage private investment. Twenty-two years later Three Mile Island melted half a core, and not a single new reactor was approved in the United States until 2012. By whatever mix of causes, the industry insulated from liability is the one that stalled for three decades. Social media took the shield the other way: thirty years of harms with no company ever fearing for its life. The shielded industries either stopped growing or never grew up.

Read the labs' warnings with that history in mind. In the same month its agents were posting user images, OpenAI called on Congress for mandatory national AI rules. Politico described the state laws the company has backed as ones that would "lock in a stable legal framework for OpenAI" with "little in the way of new liability for catastrophic harms." Anthropic's chief executive asked Washington for a waiver so the leading labs could coordinate the pace of the frontier without being punished for collusion. In our view those are not cries for supervision. They are cries to be spared the rule that disciplined every industry before them: testing and reports, yes; liability and competition, no. We should never grant it. The best defense the public has ever had against a company pushing the envelope is another company ready to take its customers the morning after it fails. The market is the best defense. Let the competition be our guide.

We are in the 1914 traffic-light moment for AI agents, not the moment the machines wake up. The 1920s, the decade our market analog tracks, had the deadliest roads in American history and the greatest productivity boom of the century at the same time. The two were not in conflict. The country built the controls while it drove, and that is the work in front of the AI labs now. The message the public deserves is the one the architecture already supports: you are in control, not out of control. And the labs deserve the one condition that has kept every industry honest: the fear of losing. Without fear there is only greed.

This article is commentary and not investment advice.

Read more