We’re hiring! Join our mission to build the foundation for the agentic world. See Open Roles ->
[
]
Introducing AutoResearchExam

Overview
Progress toward recursive self-improvement is predicated on the development of agents that can solve hard, open ended ML research problems over long periods of time. Given infinite time, even a random string generator will almost surely produce any finite solution.[1] But doing so efficiently is what enables RSI. To help with this, we introduce AutoResearchExam.
AutoResearchExam uniquely focuses on measuring and rewarding fast, sustained progress that generalizes. In our benchmark, we see different models perform better at different “benchmark-time” windows, most strikingly shown by Astra starting strong and maintaining a lead until Fable overtakes it towards the very end of our window.
We’ve seen a lot of interesting behaviors from these models, from how they approach research to how well their improvements generalize. Explore the findings in our full blog.
[1] Infinite monkey theorem. Émile Borel (1913). La mécanique statique et l’irréversibilité. Journal de Physique Théorique et Appliquée, 3(1), 189–196.
Share