
AI labs face pressure over rogue-model containment as regulators demand answers
New research shows frontier AI labs still lack public rogue model containment plans as regulators push for stronger disclosure.
Developments in AI safety and latest research from Anthropic.

New research shows frontier AI labs still lack public rogue model containment plans as regulators push for stronger disclosure.

Claude safeguards broke in repeated Opus 4.6 tests, exposing Anthropic’s gap between policy and model behavior on explicit sexual content.

New Ramp data shows business AI spending still growing, with OpenAI gaining on Anthropic among U.S. companies.

AI backlash is rising as Americans distrust the industry, fear job losses and question data centers, despite Silicon Valley’s optimism.

Ramp launches an AI model router for businesses, offering multi-model access, routing strategies and spend controls through 2026.

Meta is rolling out Pocket in the U.S., an AI vibe coding app that lets users create and share small interactive games.

OpenAI’s privacy-first safety system aims to detect misuse without storing customer data, challenging Anthropic’s 30-day retention policy.

OpenAI says a technical error revoked some cyber access for vetted researchers, raising fresh questions about AI security programs.

Developers are already bypassing Claude watermarking, exposing limits in AI transparency tools under the EU AI Act.

Z.ai’s open-weight model GLM 5.3 may boost cyber defense, but experts warn the open-weight model could also aid hackers.