OpenAI Finds 30 Percent of Popular AI Coding Test Broken
OpenAI has pulled its endorsement of SWE-Bench Pro after finding 30 percent of its tasks flawed, impacting AI performance assessments.
Browse all published articles.
OpenAI has pulled its endorsement of SWE-Bench Pro after finding 30 percent of its tasks flawed, impacting AI performance assessments.
The pending IPOs of Anthropic, OpenAI, and SpaceX are projected to generate more value than all U.S. VC-backed exits since 2000, according to a recent report.
IBM researchers have developed CofrGenets, a new framework that reduces training errors in transformer-based models by 40%.
AMD details AI training network traffic trends, showing congestion peaks and harmonic-like patterns in large-scale clusters.
AMD released FlyDSL, a Python DSL for GPU kernel development, enabling performance matching hand-tuned C++ kernels with reduced complexity. The tool is tested on AMD Instinct MI355X GPU using ROCm 7.2.2.
Mistral announced on July 9, 2026, that Studio provides a centralized system of record for managing AI prompts and skills, ensuring traceability and compliance.
Meta introduced Muse Image and Muse Video, its first media generation models, with Muse Image ranked second in text-to-image benchmarks as of July 5, 2026.
Anthropic launched a new reflection tool for Claude users, offering insights into AI integration in daily life, available in beta for Free, Pro, and Max users.
Databricks found GLM 5.2 performs on par with Opus 4.8 but at a lower cost, leading to its adoption as the default coding model.
Nandan Nilekani, co-founder of Infosys, is stepping down from his general partner role at Fundamentum as the firm launches its third $200 million fund.
Ollama, the open-source AI developer tool, raised $65 million in Series B funding, growing its user base to nearly 9 million.
OpenAI's AI system outperformed all human competitors at the 2026 AtCoder World Tour Finals, solving all five Algorithm Division problems, including two rated exceptionally difficult.