| Are AI labs pelicanmaxxing?(dylancastillo.co) | |
| 570 points by dcastm 18 hours ago | 221 comments | |
tl;dr: To test whether AI labs are gaming Simon Willison's famous "pelican on a bicycle" SVG benchmark, the author generated 1,008 SVGs across 7 frontier models using a grid of 8 animals × 6 vehicles, then scored them with an LLM judge. The results show no evidence of pelicanmaxxing: pelicans rank 6th of 8 animals, bicycles rank 5th of 6 vehicles, and no lab performs disproportionately well on the specific combination. The more likely explanation is broader "SVGmaxxing" (optimizing SVG generation generally), which this methodology can't detect. | |
HN Discussion:
| |