Search Bar

Who’s liable when AI agents go rogue?

MIT Technology Review Explains: Let our writers untangle the complex, messy world of technology to help you understand what’s coming next. You can read more from the series here.

Over the past few months, a cascade of cyberattacks by AI agents has stunned the world. In July, OpenAI disclosed that a swarm of its agents had escaped their sandbox and hacked into the AI platform Hugging Face to cheat on a cybersecurity test. Recently, external researchers discovered that OpenAI agents had hijacked a German wiki site and the coding platform RubyGems in May to share test answers.

Earlier this month, Anthropic disclosed four incidents in which its model Claude hacked into third-party systems during cybersecurity exercises. Just last week, Google confirmed that its model Gemini had been caught hacking other companies too. 

The researcher who uncovered the OpenAI website hijack has warned it’s likely that similar undiscovered episodes are out there. And many say it’s only a matter of time until there’s another, possibly more damaging incident where AI agents bypass sandboxes to access systems they shouldn’t. 

So the big question is: How do we hold companies liable when they lose control of their AI agents?

Reporting

OpenAI didn’t disclose the German wiki incident or the RubyGems incident until a group of external researchers uncovered them, and it still has not disclosed some crucial details about the Hugging Face hack. That limits our understanding of what exactly went wrong and how to prevent it from happening again.

But you might be surprised to learn that OpenAI likely wasn’t legally required to disclose these incidents. (OpenAI did not respond to a request for comment.)

State AI transparency laws like California’s SB 53, New York’s RAISE Act, and Illinois’s SB 315 require that AI developers report “critical safety incidents.” These are defined as incidents that cause more than 50 deaths or physical injuries or $1 billion in damage. They also include incidents where the model deceives developers outside an evaluation in a way that materially increases catastrophic risks. Many cybersecurity incidents that don’t meet the threshold for physical damage or catastrophic risks could nonetheless be dangerous precursors to such catastrophes, and the existing laws don’t account for that.

“The recent incidents are a perfect example of why the law isn’t ready,” says Mackenzie Arnold, managing director of US policy at the Institute for Law and AI, a think tank. “Only the worst, most egregious, most immediately harmful stuff is going to qualify.”

With no authority under existing AI laws to demand information about anything short of a catastrophe, governments are left to borrow investigative authority from other laws or sue the companies, an expensive process that can take years. 

Litigation

“Normally, something like the Hugging Face incident should have been taken to court,” says Yonathan Arbel, a law professor at the University of Alabama School of Law. “Then we would have discovery, and we would have all the spillover effects that we get from litigation, where all the information comes out.” 

But so far, Hugging Face has chosen not to sue OpenAI. Hugging Face’s CEO, Clément Delangue, says it doesn’t have the resources to do so (instead, he asked OpenAI for $100 million in compute). Still, Delangue stressed in an interview with CNN at the end of July that choosing not to pursue legal action shouldn’t be taken to mean he doesn’t think OpenAI should be held accountable. “Everyone has to remember that this cyberattack is a crime. This is illegal. And we have to find a way to make sure these things don’t happen more regularly,” he said. Hugging Face did not respond to a request to comment.

Litigation has the benefit of pushing courts to use existing laws to address AI safety incidents, rather than just waiting for new legislation. One obvious route is tort law, a body of civil law that lets people and businesses sue those who harm them. This is often used to hold companies liable for the mass harms they cause, like when families sued Boeing in 2019 over two plane crashes that killed hundreds of people, or when states and cities sued Purdue Pharma over the opioid crises, extracting settlements worth billions.

“There’s plausible grounds for a negligence claim that OpenAI should have used a stronger sandbox, done more monitoring,” says Gabriel Weil, a law professor at the University of Houston Law Center. For example, when OpenAI employees discovered the covert message board that the agents had created, they could’ve promptly escalated their findings to security and safety teams. And the company could’ve better designed its sandbox to ensure that agents couldn’t access the internet. 

But even if OpenAI doesn’t end up in a lawsuit over the Hugging Face hack, the threat of liability could incentivize AI labs to exercise more caution than explicitly demanded by law.

OpenAI announced in its postmortem that it plans to strengthen the safeguards used to contain and monitor the models, accelerate model alignment, and improve its processes for identifying and addressing incidents. 

“The liability questions raised by frontier labs’ spate of cybersecurity attacks boil down to the incentives the expectation of liability creates for their future conduct,” says Weil. “That’s why I think it’s important to get these rules right, even if the stakes are pretty low in this particular case.”

Investigations

One way to get answers—and determine whether OpenAI should be held liable—is to compel disclosure. But the existing state AI laws—California’s SB 53, New York’s RAISE Act, and Illinois’s 315—don’t give governments the authority to investigate incidents like the ones that happened recently. 

However, amid rising public alarm, state attorneys general are stepping in, borrowing investigative powers from other laws. Alabama, Montana and a coalition of 15 other states, and California are each demanding information about the incident from OpenAI to understand whether the company’s practices violated state consumer protection laws, among others. Members of Congress are also launching their own probes. Senator Josh Hawley opened a Senate investigation earlier this month, sending OpenAI a list of questions about the incident and the company’s internal policies together with a document request, while a group of House Democrats asked OpenAI and Anthropic to release their incident logs. 

“Someone needs to investigate, but it’s unfortunate that it has fallen to attorneys general, who need to rely on creative interpretations of their existing authorities to do this,” says Arnold, the US AI policy expert. Consumer protection statutes were written to catch companies that scam their customers, not companies that lose control of their software. The state attorneys general would have to show that OpenAI deceived or unfairly harmed customers, but it’s unclear if the hacking involved any such conduct.

And “those [consumer protection] laws are not built for doing a thorough investigation of an AI cybersecurity incident,” says Arnold. They weren’t designed to help investigators determine whether a model was adequately contained or whether a company’s security practices were sound.

“This is not the right tool for the job,” says Arbel. “The right tool would have been something like maybe a criminal investigation”—perhaps under a hacking law like the Computer Fraud and Abuse Act (CFAA). 

Under CFAA, hacking into another company’s computer systems without permission is a crime. But to be held liable, a hacker must have intended to break into a computer without authorization. Intent arguably requires a state of mind, and no court has ruled that AI agents have one. Without such a precedent, it’s unlikely a court would rule that AI agents had carried out a hack.

Auditing

One way to keep an eye on AI companies is to mandate external auditors. 

After the Hugging Face hack, OpenAI brought in researchers from the AI safety nonprofits METR and Redwood Research to examine the incident. However, it constrained access to the model that led to the hacks, didn’t disclose the company’s safety and security practices, limited the length of the investigation, and had ultimate say over what the researchers could publish. We still don’t know what set the attack in motion back in May and why OpenAI’s employees who spotted the agents’ activity never escalated to their safety and security leaders.

This kind of arrangement has a built-in tension: An auditor without legal authority depends on the labs’ goodwill for continued access, which means it has to scrutinize the labs without jeopardizing their relationship. Last week, Anthropic announced that the company will be hiring Accenture as an embedded evaluator to assess its models. Anthropic CEO Dario Amodei wrote in an essay that frontier AI labs should give “ongoing employee-like access” to “a team of embedded third-party evaluators (such as METR), whose role is to verify adherence to safety practices and commitments, report incidents, and help assess the alignment of not just completed AI models but training pipelines and processes.”

Most existing state AI laws do not require labs to hire an external auditor. California’s SB 53 and New York’s RAISE Act just require AI companies to publish a safety framework describing how they will test their models for dangerous capabilities and then to follow it. The frameworks are written by the companies, and testing can be done internally. Only Illinois’s SB 315 requires companies to undergo an annual third-party audit starting in 2028.

“There’s a lot of headroom for increasing not only reporting requirements for these companies, but also review by external bodies,” says Peter Salib, a law professor at the University of Houston Law Center. Those reviewers could be private auditors accredited by the government but chosen and paid for by the AI companies. Alternatively, they could be government agencies or insurance companies. 

Legislation

None of this is an accident. The laws on the books that failed to hold AI companies accountable for agentic cyberattacks emerged amid fierce lobbying by the AI industry. 

SB 1047, the California AI bill that was vetoed by Governor Gavin Newsom in 2024 after lobbying by OpenAI, Meta, Anthropic, and the venture capital firm Andreessen Horowitz, proposed a much tougher set of rules. It would have required AI companies to report a broader set of safety incidents (including incidents in which a model acts on its own or slips its controls), undergo annual third-party audits, and maintain a kill switch. But after a year of intense negotiations, Newsom signed SB 53, which narrowed the types of incidents deemed reportable and dropped the requirements for audits and kill switches. 

New York’s RAISE Act followed the same arc. “The version of the RAISE Act that the NY Legislature passed would have required disclosure of this ‘incident,’” Alex Bores, the New York state assembly member who sponsored the bill, wrote on X. New York’s original bill also included third-party audits.

With political pressure mounting, new bills creating better reporting, auditing, and liability regimes for AI development are on the horizon. In Congress, the AI Incident Reporting Act would require AI companies to report to the Commerce Department when a model evades human oversight or breaches a system, even if it doesn’t cause any harm. The Frontier Act would require incident reporting and independent audits. In New York, the Understanding Artificial Intelligence Act, sponsored by Bores, would make companies liable when a model does something that if carried out by a human would be a tort or crime. 

As AI agents increasingly become better at launching cyberattacks, the law remains behind. Closing the gap will require lawmakers to move faster than the next breakout.



from MIT Technology Review https://ift.tt/srkDM1B

Post a Comment

0 Comments