Are AI labs pelicanmaxxing?(dylancastillo.co)
570 points by dcastm 18 hours ago | 221 comments
tl;dr: To test whether AI labs are gaming Simon Willison's famous "pelican on a bicycle" SVG benchmark, the author generated 1,008 SVGs across 7 frontier models using a grid of 8 animals × 6 vehicles, then scored them with an LLM judge. The results show no evidence of pelicanmaxxing: pelicans rank 6th of 8 animals, bicycles rank 5th of 6 vehicles, and no lab performs disproportionately well on the specific combination. The more likely explanation is broader "SVGmaxxing" (optimizing SVG generation generally), which this methodology can't detect.
HN Discussion:
  • Praise for the robust methodology and appreciation that someone quantitatively tested the pelicanmaxxing hypothesis
  • The right-facing bicycle observation is explained by real-world photography conventions showing the drivetrain
  • ~Alternative pattern spotted: models appear to be Ottermaxxing on the 'otter on a plane' benchmark instead
  • SVGmaxxing isn't really a problem since improving SVG generation is a legitimately useful skill
  • Sharing related experiments or anecdotes about model behavior on similar prompts