OpenAI's GPT-6 Astra on ARC-AGI-3(arcprize.org)
232 points by vignesh_warar 6 days ago | 146 comments
tl;dr: GPT-6 Astra scored 62.7% ($26K) on ARC-AGI-3 Semi-Private using ARC's Standard harness and 99.9% ($19K) using the Provider Adapter harness, both state-of-the-art. Notably, Astra beat the median human's action efficiency on 96% of levels, using 51.7% fewer actions on average, and developed its own compact algebraic notation to model game mechanics. ARC's authors emphasize this is meaningful progress toward generalization but not proof of AGI, given the benchmark's bounded scope.
HN Discussion:
  • Benchmark performance doesn't equate to true intelligence or AGI
  • Suspicion of benchmark gaming via custom harnesses or data leakage
  • Cost-per-puzzle trajectory suggests AI will soon undercut human labor
  • Article fails to clarify what capabilities remain out of reach
  • ~AGI goalposts will keep moving regardless of achievements