Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard(artificialanalysis.ai)
319 points by aarondong 16 hours ago | 180 comments
tl;dr: Anthropic's Claude Opus 5 has taken the top spot on Artificial Analysis's Intelligence Index v4.1, which aggregates nine benchmarks including GDPval-AA v2, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, and GPQA Diamond. The leaderboard also tracks cost per task, output speed, latency, token usage, and context window size across proprietary and open-weight models.
HN Discussion:
  • Confirms Opus 5's superiority by highlighting it outperforms competitors even at lower effort settings
  • Cost-effectiveness undermines the ranking since cheaper models match Opus 5's performance
  • Claude's censorship and safeguards make its top ranking practically meaningless for real use
  • Personal user experience praising Opus 5's improved behavior over previous models
  • ~Benchmarks are incomplete without testing long-context performance with irrelevant filler