Epoch AI Study Shows AI Detectors Struggle With Style Imitation
AI text detectors like Pangram and GPTZero fail to catch 13% of style-mimicking AI texts, according to Epoch AI research.
215 articles
In-depth coverage of new AI research — papers, benchmarks, and breakthroughs from leading labs and academia, summarized for fast reading and grounded in the methods that matter.
AI text detectors like Pangram and GPTZero fail to catch 13% of style-mimicking AI texts, according to Epoch AI research.
Moonshot's Kimi K3 scored 1,679 in frontend code benchmark, surpassing Fable 5's 1,631, according to the Code Arena: Frontend test.
Researchers developed EgoBabyVLM, a test that challenges AI models to learn like infants, using footage from babies' cameras. The test reveals current models struggle to interpret messy, real-world data.
Researchers uncovered the original ELIZA source code from MIT archives, showing the chatbot could assume various personas beyond its therapist role.
Hugging Face evaluated 13 open models on three Swiss legal benchmarks with 55,861 scored samples each, showing the top five models differ by only 2.1 points.
Anthropic's new research reveals a hidden 'J-space' within AI models that influences their reasoning, offering deeper insight into how they process information.
AMD announces ROCm-optimized video sparse attention with 3.13× speedup over prior methods, reducing training-time latency to 73.06 ms.
Anthropic's new research suggests Claude may have two distinct language processing routes, with one resembling access consciousness. The study shows the model can continue Spanish text while naming French authors, indicating potential for explicit reasoning.
A Danish team used quantum computing to enhance AI models, generating more effective peptides for vaccine development.
Cohere's new technique improves large language model inference speed by adapting to hardware constraints, achieving up to 23% faster performance at high batch sizes.
Anthropic has identified a hidden area within its Claude Opus 4.6 model called J-space, offering new insights into how the AI processes information.
OpenAI has pulled its endorsement of SWE-Bench Pro after finding 30 percent of its tasks flawed, impacting AI performance assessments.