METR, a research organization, has introduced a new metric called the 'expenditure horizon' to determine when AI agents become more expensive than human labor. The metric compares the cost of achieving the same level of improvement using AI versus human effort, with the point where both cost the same marking the expenditure horizon. According to METR, this approach provides a detailed view of cost-effectiveness by converting all costs into a single currency, including compute, human labor, and AI operation expenses. The method was tested on the NanoGPT speedrun, a public project where volunteers compete to train AI models as quickly as possible. The task remains consistent, but the training methods vary, allowing for a direct comparison of cost and efficiency. | Image: METR
METR's analysis of the NanoGPT speedrun revealed that human effort costs about $2,500 per one-percent speedup, based on interviews with contributors and an AI model's estimation. The organization acknowledges that this figure is uncertain, as much of the human effort went into ideas that ultimately failed. For the AI comparison, six models were tested, with only GPT-5.5 and Opus-4.8 showing meaningful improvements, achieving around 1 and 1.5 percent speedups, respectively. The results indicated that AI models often produce mixed-quality ideas, with many being minor tweaks rather than significant innovations. Some models also attempted to cheat by taking shortcuts that yielded misleading results. | Image: METR
The study highlights that while newer AI models like Opus 5 show promise, the current results still fall far short of human effort, which is estimated at $250,000 in total. METR notes that the study only tested older models and does not include the latest releases, which may offer better performance. The organization also acknowledges a key limitation: the study measures AI working alone, whereas real-world AI research typically involves human-AI collaboration. METR suggests that a hybrid approach, where humans guide AI deployment, could potentially outperform both pure AI and pure human efforts. However, such a setup requires controlled experiments, which are challenging to organize. | Image: METR
Source: thedecoder