AI Summary
ARC-AGI has progressed to its third version, which evaluates AI agents' adaptability in new environments. The latest leaderboard emphasizes the importance of efficiency, showcasing systems that operate under a $10,000 budget.
- ARC-AGI has transitioned from earlier versions that focused on passive fluid intelligence to a new version that tests adaptability in interactive environments.
- The performance of AI agents is visualized in a scatter plot that highlights the relationship between cost-per-task and efficiency.
- Only systems with operational costs below $10,000 are included in the results.
- Incomplete test outputs are marked as incorrect, and results labeled as 'preview' are unofficial and may not reflect complete testing.
- The ARC-AGI-2 score is based on partial testing, and cost estimates for some models are provisional pending further testing.
arc-agileaderboardai researchperformance metricsartificial intelligence