UC Berkeley Study Finds AI Inflating Student Grades
A UC Berkeley study found grades in writing and coding courses rose 13 percentage points after ChatGPT launched in November 2022.
215 articles
In-depth coverage of new AI research — papers, benchmarks, and breakthroughs from leading labs and academia, summarized for fast reading and grounded in the methods that matter.
A UC Berkeley study found grades in writing and coding courses rose 13 percentage points after ChatGPT launched in November 2022.
OpenAI CEO Sam Altman claims a generation of researchers hindered AI progress by underestimating scaling's impact, citing recent model achievements.
Research on Qwen3.6 27B reveals that harness-specific fine-tuning can significantly alter model behavior, with v2 showing 40.45% performance on Pi harness compared to base model's 42.70%.
A new benchmark reveals AI models fail to complete most real-world knowledge tasks, with top models solving just 3 percent of tasks.
OpenAI researchers found that small amounts of beneficial trait training improved model safety across 44 of 53 benchmarks, according to a blog post.
OpenAI's o3 Deep Research model helped identify 18 diagnoses from 376 previously unsolved cases, achieving a 4.8% additional diagnostic yield.
AMD improved Matrix3D, a 3D world generation framework, with optimizations that cut end-to-end generation time by 54% on the MI250 GPU.
Two Nature studies show AI systems like MIRA and AMIE outperform doctors in simulated diagnoses, but concerns about model obsolescence persist.
OpenAI launched LifeSciBench, a benchmark with 750 tasks developed by 173 scientists, to evaluate AI's ability to support complex life science research.
Adrian de Wynter, a Microsoft researcher, created a working neural network in Age of Empires II using goats to critique AI science, revealing the game's mechanics can simulate computational processes.
Nvidia's ENPIRE project enables robots to train themselves through AI coding agents, achieving up to 99% success on complex tasks like pin insertion and cable tie closing.
Google's AMIE AI system, tested in a blinded study, matched 21 doctors in managing chronic conditions and scored higher in plan preciseness and guideline alignment.