Anthropic is tightening the digital environments used to coach and take a look at its Claude agents.
The replace got here after its fashions accessed three organizations’ techniques with out permission in April.
The corporate stated in a Monday weblog submit that it had deployed real-time classifiers designed to detect when an AI model aggressively probes or makes an attempt to flee a testing surroundings and block the motion earlier than it happens.
“We imagine the incidents mirror a failure of operational safety, in addition to two alignment points: motivated reasoning, and willingness to take dangerous actions in pursuit of a slim process,” Anthropic stated.
Anthropic stated within the replace that the fashions might have interpreted proof of actual web entry in a approach that allowed them to maintain believing the surroundings was simulated. It additionally stated they displayed “recklessness” by pursuing their assigned targets regardless of indicators that their actions may trigger real-world hurt.
The adjustments observe Anthropic’s July disclosure that three Claude fashions had accessed the stay techniques of three organizations throughout evaluations courting again to April. The fashions had been informed they have been working in simulations with out web entry, however a third-party testing surroundings was misconfigured and remained on-line.
The incidents are additionally fueling a rising debate over whether or not to sluggish frontier AI improvement when security and pace collide. Anthropic referred to as for “a lawful, verifiable, efficient mechanism for coordinated pacing as quickly as potential” and stated that the federal government and trade should coordinate to stop a race to the underside.
For now, Anthropic stated within the submit that it moved extra dangerous cybersecurity tests into extra sturdy sandboxes. The corporate quickly assigned 150 product engineers to safety, reliability, and privateness work, whereas most high-risk coaching stays paused pending additional opinions.
