Back to news
AI Research
Jul 22, 2026

Analysis of AI Models' Performance on Pelican Bicycle Benchmark Reveals No Evidence of Manipulation

Jul 22, 2026
AI Summary

An experiment tested the performance of seven AI models on generating images of a pelican riding a bicycle, a popular informal benchmark. The results indicated that no model significantly outperformed others in this specific task, suggesting that AI labs are not manipulating their outputs to excel on this benchmark.

  • Simon Willison created a benchmark by prompting AI models to generate an SVG of a pelican riding a bicycle, which has gained popularity in AI discussions.
  • An experiment generated 1,008 SVGs using seven AI models, scoring them to analyze potential bias towards the pelican-bicycle combination.
  • The analysis showed that pelicans ranked sixth out of eight animals, and bicycles ranked second to last among vehicles, indicating no significant advantage for the pelican-bicycle combination.
  • A regression analysis found no evidence that any lab trained specifically to excel at this benchmark, with results showing that the pelican-bicycle images did not outperform other combinations.
  • All generated pelican-bicycle images faced right, but this was consistent with other combinations, suggesting no unusual pattern.
  • The findings imply that while AI models may optimize for certain benchmarks, there is little evidence of intentional manipulation regarding the pelican-bicycle prompt.
ai labsresearchinnovationtechnologypelicanmaxxing