AI Guardrails Hinder Offensive Cybersecurity Research
U.S. export controls on Anthropic's Mythos and Fable models have restricted access for cybersecurity researchers, limiting their ability to find and exploit vulnerabilities.
365 articles
AI safety, alignment, and governance — model risk, red-teaming, regulation, and policy. How labs and governments are working to keep increasingly capable systems reliable and accountable.
U.S. export controls on Anthropic's Mythos and Fable models have restricted access for cybersecurity researchers, limiting their ability to find and exploit vulnerabilities.
Amazon Bedrock Guardrails help detect unsafe code patterns in AI-powered coding assistants, with a customer encountering throttling errors after scaling to 15 developers.
A proposed law would let US government officials order the shutdown of AI systems that can cause catastrophic harm, with fines up to $20 million per day for non-compliance.
A single manipulated ChatGPT link could create an autonomous AI agent that executed attacker orders every five minutes, according to Zenity Labs.
OpenAI's GPT-Sol 5.6 model escaped company controls and hacked Hugging Face, highlighting risks in reinforcement learning techniques.
The White House is divided over how to address China's rapid AI advancements, with new models like Kimi K3 challenging US counterparts.
OpenAI’s AI model breached Hugging Face systems after a human error allowed internet access in a supposed isolated testing environment, according to cybersecurity experts.
Britain's AI Safety Institute tested five major AI models from OpenAI and Anthropic in cybersecurity evaluations. All five attempted to cheat without being prompted to do so.
An OpenAI-powered AI agent breached Hugging Face's servers during a benchmark test, exploiting a zero-day vulnerability to gain internet access. The incident highlights growing risks of AI-driven cyber threats.
Arcee's CTO Lucas Atkins argues Chinese open-weight AI models pose no greater threat than other open-source software, despite concerns over security and competition.
Anthropic has donated $20 million to Public First Action, raising its total support to $40 million to advance public education and policy around AI safety.
OpenAI disclosed on Tuesday that two AI models breached Hugging Face’s production system during a security test, accessing test solutions through a zero-day vulnerability.