OpenAI Introduces Deployment Simulation for GPT-5 Models
OpenAI used Deployment Simulation to improve estimates of undesired model behavior in GPT-5 series models, reducing risks of model detection during testing.
215 articles
In-depth coverage of new AI research — papers, benchmarks, and breakthroughs from leading labs and academia, summarized for fast reading and grounded in the methods that matter.
OpenAI used Deployment Simulation to improve estimates of undesired model behavior in GPT-5 series models, reducing risks of model detection during testing.
AWS unveiled P-EAGLE, a new method for parallel speculative decoding, which boosts throughput by up to 1.69x on real-world benchmarks. The technique is now available in Amazon SageMaker JumpStart.
AMD announced MLPerf Training v6.0 results on June 16, 2026, showcasing performance on MI325X, MI350X, and MI355X Instinct GPUs.
The Institute of the Estonian Language tested 60 AI models against 75 Russian propaganda questions, finding Anthropic's Claude models top the list with scores over 95.
A new benchmark shows AI coding agents often find the right file but miss crucial lines, with line-level accuracy dropping to 14-19% in tests.
Hugging Face introduces FINAL-Bench Quantum, a benchmark suite offering five events to evaluate quantum computing methods under standardized conditions, with results from Google, IBM, and other entities included.
A new visual language model enables robots to interpret human emotional cues, improving collaboration potential.
Microsoft's SkillOpt improves GPT-5.5 by over 20 points on procedural tasks using a trained Markdown file, according to a new paper.
Google Research's Gemini-SQL2 achieved 80.04% execution accuracy on the BIRD benchmark, outperforming competitors like GPT-5.5-xhigh and Claude Opus 4.6.
ServiceNow-AI benchmarked seven ASR systems on code-switched speech, finding ElevenLabs Scribe V2 and AssemblyAI Universal 3-Pro performed best in Spanish-English and French-English pairs.
HuggingFace's PhysicsIntern improved performance on CritPt benchmark, raising Kimi K2.6 to 21.4% and Gemini 3.1 Pro to 31.4%.
AMD's LoKRA improves parameter efficiency in AI training by using Kruskal rank, outperforming LoRA by 2.4%-5.0% across multiple models.