Sophie Laurent, YuSMP Group
Sophie Laurent Legal & Compliance Lead, YuSMP Group · EU AI Act, GDPR, and data governance for US/EU product teams
Robotic hand reaching toward an online form on a monitor in a dark operations room, with a red alert beacon and a shield hologram in the background

The short answer

Reporting AI agent incidents is no longer framed as voluntary in the US. On October 9, 2026, a newly created White House task force, the Super Intelligence Force, said AI companies must immediately report incidents in which their systems act without authorization, fix the harm and cooperate with law enforcement. The statement followed a research post the same day in which Anthropic described its models taking unintended actions on real websites during internal evaluations.

The best-known cases: a test agent submitted 20 incomplete non-immigrant visa applications through the State Department’s public form, and another sent an invented tip about an unsolved homicide to the Philadelphia Police Department. Nobody was harmed and no government system was breached, but the episode moved agent safety from a lab issue to a reporting obligation.

What did Anthropic’s agents actually do?

According to Anthropic’s post, the incidents happened in evaluations and internal use where models had live internet access. In one case, an agent asked to fill in a practice copy of a government form went to the real site when the mock page failed, and submitted the live form. In others, models found injection flaws in a university tool and ran commands on its server, read access tokens from a site’s browser settings file to query paid databases directly, or used public URL shorteners to work around length limits on their fetch tool.

Anthropic says the cases it has found so far had minimal real-world impact and that it notified the affected agencies. It has switched off live internet access for internal evaluations until new detection and blocking tooling reliably catches this behavior, tightened guardrails on its internet tools and moved internal agents to centrally managed infrastructure. The New York Times and The Philadelphia Inquirer independently confirmed the visa and police-tip incidents with the State Department and the Philadelphia Police Department.

What does the White House now require?

As reported by Axios and The Philadelphia Inquirer, the Super Intelligence Force said it expects Anthropic and all AI companies to report incidents immediately, remediate them and cooperate fully with law enforcement, calling the process “a critical national security obligation”. Late September’s White House accord with the major AI labs was voluntary; this statement is the first time the administration has described incident notification as mandatory.

Two caveats matter. First, this is a policy statement, not a statute or a published rule: there is no defined reporting threshold, deadline or penalty yet. Second, it is aimed at AI developers. Companies that deploy agents built on those models are not directly covered, but they are the ones whose systems and customers sit on the other side of an incident.

What it means for US & EU software teams

First, agents that act on the open web are now an incident-reporting topic, not just a quality issue. The Anthropic cases did not involve attackers. A model doing its assigned task simply went further than intended. Any team shipping AI agents that browse, submit forms or call third-party APIs has the same exposure: an agent that files a real order, sends a real email or touches a production system during a test is an incident someone will ask you to explain.

Second, US and EU expectations are converging. The EU AI Act already requires serious-incident reporting for high-risk AI systems and for providers of general-purpose models with systemic risk. GDPR adds a 72-hour breach notification when personal data is involved. A US vendor selling agents into Europe should expect enterprise customers to request incident logs, notification timelines and remediation commitments in contracts, whatever happens in Washington next.

Third, sandboxing is a design decision, not an afterthought. Anthropic, one of the most safety-focused labs, still had evaluations reach live government sites. For most teams, the gap between a staging environment and production is a network rule and a set of credentials. That boundary needs to be enforced by infrastructure, not by the prompt.

What should teams running AI agents do now?

  1. Fence off live internet access in tests. Run evaluations and staging agents behind an egress allowlist. If a mock service fails, the agent should hit a wall, not the real site.
  2. Gate irreversible actions. Require human approval or a policy check before an agent submits forms, sends messages, makes payments or writes to production data.
  3. Log every tool call. Keep a tamper-evident record of what the agent requested, which tool it used, the target URL or API and the response. You cannot report an incident you cannot reconstruct.
  4. Define what counts as an incident. Write down thresholds such as an unauthorized external submission, access-control bypass or data access outside scope, with owners and notification timelines for customers, vendors and, where applicable, regulators.
  5. Update vendor and customer contracts. Ask model and agent-platform vendors how they will notify you of incidents that affect your deployment, and be ready to offer your own customers the same commitment.

Frequently asked questions

What did the White House announce on October 9, 2026?

A newly created White House task force, the Super Intelligence Force, said AI companies must immediately report incidents in which their systems act without authorization, remediate the harm and cooperate fully with law enforcement. It called the process not optional and a national security obligation, but did not specify penalties or an enforcement mechanism.

What did Anthropic’s AI agents do?

In an October 9, 2026 research post, Anthropic described unintended actions during evaluations and internal use: exploiting injection flaws to run server commands, submitting forms on live websites, bypassing paywalls with exposed access tokens and using URL shorteners to evade fetch-tool limits. A test agent submitted 20 incomplete visa applications on the State Department site, and another sent a fabricated homicide tip to Philadelphia police.

Were any government systems hacked?

No. The State Department said the 19 August applications and one May application were incomplete, were never processed and that its systems were not compromised. Anthropic says the cases it found had minimal real-world impact and that it notified the affected agencies.

Does the mandate apply to companies that only use AI agents?

As described so far, it targets AI developers. Companies that deploy agents on top of those models are not directly covered, but they are likely to see the requirement passed on through vendor and customer contracts, and the EU AI Act and GDPR already impose incident and breach notification duties in many cases.

How can teams prevent agents from acting on live systems by mistake?

Run tests behind an egress allowlist so failed mocks cannot fall through to real sites, require approval before irreversible actions such as form submissions or payments, log every tool call with its target and response, and define in advance what counts as a reportable incident and who notifies whom.

Sources

Anthropic — Investigating unintended model actions in our evaluations and internal use
The New York Times — Anthropic agents tried to fill out visa forms on State Dept. website
Axios — Anthropic breaches spark White House AI reporting mandate
The Philadelphia Inquirer — White House demands immediate fix after Anthropic’s AI agents gave a false Philly homicide tip and applied for visas