Estonian Institute Tests AI Models Against Russian Propaganda
The Institute of the Estonian Language tested 60 AI models against 75 Russian propaganda questions, finding Anthropic's Claude models top the list with scores over 95.
116 articles
In-depth coverage of new AI research — papers, benchmarks, and breakthroughs from leading labs and academia, summarized for fast reading and grounded in the methods that matter.
The Institute of the Estonian Language tested 60 AI models against 75 Russian propaganda questions, finding Anthropic's Claude models top the list with scores over 95.
A new benchmark shows AI coding agents often find the right file but miss crucial lines, with line-level accuracy dropping to 14-19% in tests.
Hugging Face introduces FINAL-Bench Quantum, a benchmark suite offering five events to evaluate quantum computing methods under standardized conditions, with results from Google, IBM, and other entities included.
A new visual language model enables robots to interpret human emotional cues, improving collaboration potential.
Microsoft's SkillOpt improves GPT-5.5 by over 20 points on procedural tasks using a trained Markdown file, according to a new paper.
Google Research's Gemini-SQL2 achieved 80.04% execution accuracy on the BIRD benchmark, outperforming competitors like GPT-5.5-xhigh and Claude Opus 4.6.
ServiceNow-AI benchmarked seven ASR systems on code-switched speech, finding ElevenLabs Scribe V2 and AssemblyAI Universal 3-Pro performed best in Spanish-English and French-English pairs.
HuggingFace's PhysicsIntern improved performance on CritPt benchmark, raising Kimi K2.6 to 21.4% and Gemini 3.1 Pro to 31.4%.
AMD's LoKRA improves parameter efficiency in AI training by using Kruskal rank, outperforming LoRA by 2.4%-5.0% across multiple models.
Isomorphic Labs' AI engine can predict hidden binding sites on proteins, potentially speeding up drug discovery by identifying previously unseen targets.
Anthropic's study reveals large language models can create exploits from security patches in hours, not weeks, with Mythos Preview achieving results in under six hours.
Astrophysicist Chi-kwan Chan uses Codex to refine algorithms for simulating black hole plasma, improving computational efficiency in extreme physics research.