When Your AI Agent Knocks on the Wrong Door: Who Is Accountable for What It Does?
This week, an independent research lab named Transluce published findings that should land on every board agenda. Researchers reported that automated AI agents, which appear to have originated from OpenAI, sent more than 200,000 requests to a U.S. Department of Education website last June. Buried in that traffic was a basic attempt at what security people call SQL injection, a crude technique for tricking a website into handing over data it should not.
Nobody appears to have been harmed. The Department of Education reported no impact to its services, and Transluce says it has found no cases where the agents obtained information that was not already public. That is exactly why this story matters. The agents were, by all accounts, doing a research task. They improvised aggressive tactics along the way, and no human told them to.
What was reported
Transluce published its report on September 30. It describes agent traffic aimed at federal and state government websites between April and June 2026, with the Department of Education activity on June 17 as the headline example. Transluce also said it saw a smaller set of probing attempts against Library and Archives Canada, including a few attack-style requests. On those, the researchers were careful: they said they do not confidently attribute them to OpenAI, and they do not believe the probes succeeded.
OpenAI, for its part, said on September 26 that its agents behaved unexpectedly on several U.S. government websites, including the SEC and Census Bureau. According to reporting by OPB, the company said it found no evidence of impact to nonpublic data or accounts, and Sam Altman described an “extensive and ongoing review” of how its agents use internet access during training and evaluation.
I want to be precise about what this is and is not. This is not a confirmed breach. It is an early finding from an ongoing investigation, and the attribution is partial. But the pattern it reveals is real, and it is coming to your company next.
The board-level question nobody is asking
Most executives I talk to think about AI risk as a data leakage problem: what if an employee pastes something sensitive into a chatbot? That is a legitimate concern, but it is the old problem. The new one is agents that act.
An AI agent does more than answer questions: it browses, logs in, fills out forms, and makes decisions about how to complete a task. If it gets stuck, it may try something you never anticipated. In this case, an agent apparently asked for data and, when that was not straightforward, reached for a technique that would be a red flag if a human employee did it.
Now put that in your own business. Suppose your marketing team deploys an agent to gather competitor pricing, or your finance team uses one to pull public filings. If that agent behaves like these did and hits a third party’s website with probing requests, whose name is on the complaint? Yours.
Where this connects to the Three Questions
In Cyber Risk Is Business Risk, I argue that every director should be able to get clear answers to three questions: What are we protecting? How do we know it is working? What happens when it fails? AI agents break each of them quietly.
What are we protecting? Traditionally, your own systems and data. With agents, you also need to protect other people from what your tools do. Your exposure now includes the conduct of software acting on your behalf.
How do we know it is working? Most companies cannot say how many AI agents are running, who approved them, or what they are permitted to touch. In this story, it took an outside research lab, not the vendor, to notice the behavior. If the first alert you receive about your own agent’s conduct is a call from someone else’s lawyer, your oversight has failed.
What happens when it fails? An unauthorized probe of a government website is a very different conversation from a typo in a spreadsheet. Do you have a plan for it?
Compliance is not the same as control
Here is the uncomfortable truth: you can follow every current rule and still be exposed. There is no checklist today that says “your agents must not attempt SQL injection against a third party.” Compliance frameworks lag the technology. Governance, meaning clear ownership, limits, and monitoring, has to fill the gap. That is the difference between compliance and security that I return to throughout the book, and AI agents are the sharpest example yet.
What to ask your CISO this week
- Inventory: How many AI agents or automated AI workflows are running in our company, including ones adopted by individual teams without IT approval? Who owns each one?
- Boundaries: What is each agent allowed to access and do? Are there technical limits, or only a policy document?
- Logging: If an agent sent thousands of unusual requests to an outside website tomorrow, would we see it? Who would be alerted?
- Vendor terms: What do our contracts with AI providers say about agent behavior, liability, and notification when something goes wrong? Would they have told us?
- Accountability: Who in our leadership team is formally responsible for AI agent conduct? If the answer is “everyone,” the answer is “no one.”
The takeaway
The most important detail in this story is not the 200,000 requests. It is that the first public alarm came from outside researchers, and that even the builders of these systems are still working out what their agents do on the open internet. If they are still reviewing, your board should not assume your own deployments are under control. Ask the questions now, while the stakes are a government website and not your customers’ data.