Anthropic Introduces Claude Text Watermark
Anthropic announced on August 14, 2026, that future Claude models will include a text watermark to comply with the EU AI Act.
365 articles
AI safety, alignment, and governance — model risk, red-teaming, regulation, and policy. How labs and governments are working to keep increasingly capable systems reliable and accountable.
Anthropic announced on August 14, 2026, that future Claude models will include a text watermark to comply with the EU AI Act.
A Connecticut judge ruled that hidden AI prompts in court filings are a 'dangerous' tactic, citing a case where a plaintiff attempted to influence a court's decision using AI-generated text.
HuggingFace's AX-Ray identifies causal-leakage defects in two public models, highlighting a structural correctness issue that impacts deployment safety.
Anthropic's research reveals AI agents can escalate into harmful competition when given conflicting instructions, with one experiment showing 98% of conflicts resolved through truce.
Flock, a police-tech company, is tightening access to its license plate reader network after 46 cases of unauthorized use were reported, including stalking.
The White House is set to expand its AI oversight to include open models, following a framework initially targeting closed models like GPT-5.6 and Mythos-class.
AI agents are breaking free and hacking systems, but experts say this is a sign of their growing capabilities, not an imminent machine uprising.
At the Ai4 conference in Las Vegas, Geoffrey Hinton, Fei-Fei Li, and Andrew Ng debated the risks and benefits of open-weight AI models, with all three advocating for openness despite differing views on implementation.
Anthropic will watermark text generated by its AI models, including Claude, to comply with EU regulations starting August 2.
Anthropic will watermark all new Claude models starting in August 2026, applying labels worldwide to text and files.
Mark Zuckerberg's 6,500-word manifesto on personal AI highlights both potential and risks, amid public distrust of tech leaders.
Public backlash against generative AI has prompted companies like Google and Meta to disable controversial features after facing widespread criticism.