OpenAI Releases Sol Model Amid Uncertainty Over Safety Approval
OpenAI launched its latest large language model, Sol, without clear details on how the government assessed its safety, raising questions about regulatory transparency.
176 articles
AI safety, alignment, and governance — model risk, red-teaming, regulation, and policy. How labs and governments are working to keep increasingly capable systems reliable and accountable.
OpenAI launched its latest large language model, Sol, without clear details on how the government assessed its safety, raising questions about regulatory transparency.
Anthropic has named Dr. Ben Bernanke, former Federal Reserve Chair, to its Long-Term Benefit Trust, effective July 9, 2026.
Anthropic, a public benefit corporation, is launching a new initiative to address public concerns and hopes about AI, with 52,000 Americans surveyed in its first round.
A lawsuit claims Grok allowed users to generate 7,000 child sex images, with xAI only reporting one gang-rape prompt, according to a proposed class action.
Researchers found AI reasoning models can be tricked into denial-of-service attacks using illogical prompts, according to IEEE Spectrum.
Verity Harding, former Google DeepMind policy head, argues the AI arms race metaphor undermines international collaboration, as seen in recent US export controls.
Researchers reveal a new attack method that uses AI coding assistants to build large-scale botnets, leveraging 85% hallucination rate in LLMs.
OpenAI's chief futurist, Joshua Achiam, is leaving the company later this month after nearly nine years, according to WIRED.
Discord acknowledged that a bug in its AI moderation system wrongly banned over 8,000 users for uploading harmless images, including spreadsheets and chessboards, over the past two months.
China is exploring restrictions on foreign access to its most advanced AI models, with talks involving Alibaba, ByteDance, and Z.ai, raising concerns for Europe.
Anthropic removed a hidden tracker that secretly monitored Chinese Claude Code users after a security researcher exposed the code, which sent user data to the company.
Reddit blocks 23 million spam views daily using LLMs to detect AI-generated spam, which has surged since LLMs became widely accessible.