OpenAI's Deep Research smashes records for the world's hardest AI exam, with ChatGPT o3-mini and DeepSeek left in its wake
image via TechRadar
February 4, 2025, 12:24 PM
- •OpenAI's Deep Research has topped the leaderboard of Humanity's Last Exam, scoring 26.6% accuracy
- •Deep Research has search capabilities which make comparisons slightly unfair, as the other AI models don't
- •Humanity's Last Exam is an excellent benchmark, and one that will prove invaluable as AI models develop
The world's hardest AI exam, Humanity's Last Exam, was launched less than two weeks ago, and we've already seen a huge jump in accuracy, with ChatGPT o3-mini and now OpenAI's Deep Reasoning topping the leaderboard. Deep Research has search capabilities which make comparisons slightly unfair, as the other AI models don't. While a score of 26.6% on Humanity's Last Exam is seriously impressive, especially considering how far the benchmark's leaderboard has come in just a couple of weeks, it's still a low score in absolute terms – no one would claim to have passed a test with anything less than 50% in the real world.
Entities Mentioned
John-Anthony Disotto
Topics Covered
Artificial IntelligenceComputingSoftware
Comments (0)
No comments yet.