UNCOS

OpenAI's Deep Research smashes records for the world's hardest AI exam, with ChatGPT o3-mini and DeepSeek left in its wake

OpenAI's Deep Research smashes records for the world's hardest AI exam, with ChatGPT o3-mini and DeepSeek left in its wake

image via TechRadar

February 4, 2025, 12:24 PM

  • OpenAI's Deep Research has topped the leaderboard of Humanity's Last Exam, scoring 26.6% accuracy
  • Deep Research has search capabilities which make comparisons slightly unfair, as the other AI models don't
  • Humanity's Last Exam is an excellent benchmark, and one that will prove invaluable as AI models develop

The world's hardest AI exam, Humanity's Last Exam, was launched less than two weeks ago, and we've already seen a huge jump in accuracy, with ChatGPT o3-mini and now OpenAI's Deep Reasoning topping the leaderboard. Deep Research has search capabilities which make comparisons slightly unfair, as the other AI models don't. While a score of 26.6% on Humanity's Last Exam is seriously impressive, especially considering how far the benchmark's leaderboard has come in just a couple of weeks, it's still a low score in absolute terms – no one would claim to have passed a test with anything less than 50% in the real world.

Read original article

Entities Mentioned

John-Anthony Disotto

Topics Covered

Artificial IntelligenceComputingSoftware

Comments (0)

No comments yet.