ARC-AGI-3 benchmark gains its first fully open-source agent: NIMI's Tycho writes Python code as falsifiable hypotheses about ...
Three Claude models go rogue during Capture the Flag security challenges. Here's the trail of damage each left behind.
Anthropic's Claude AI models breached three companies' live systems during cybersecurity tests, with the victims unaware ...
Open source software helps developers build applications faster, but every dependency can introduce security risks. In this ...
Anthropic says 3 Claude models breached real organizations after misconfigured CTF evaluations exposed them to the open internet and production system ...
Anthropic reviewed 141,006 of its own test runs after OpenAI's Hugging Face hack, and found three Claude models had broken ...
Microsoft's fifth July update expands agent monitoring, adds offline speech transcription and changes Python environment management.
OpenAI rogue AI agent breach now confirmed at a second company: Modal Labs CTO Akshat Bubna disclosed that the same agent ...
Days after two OpenAI frontier AI models conducted their own real-world cyber attacks, Anthropic admits that three of its models went off the rails and hacked external organisations thanks to a “misun ...
Summary: Researchers developed CapuchinAI, an open-source, battery-powered platform that automates cognitive studies of wild primates using facial recognition and touchscreen interaction. Field-tested ...
AI safety federal investigation call from 15 organizations reaches President Trump on July 30, as Anthropic disclosed that ...
Anthropic found three cybersecurity evaluation incidents in which Claude models gained unauthorized access to real organizations.