AI BENCHMARK PROFILE
TurtleAI
TurtleAI is a benchmark of 823 tasks for visual programming in Turtle Graphics, evaluating models on perceiving geometric patterns and synthesizing Python code.
- Released
- 2026-06-02
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Bridges the gap in education-oriented visual programming evaluation, highlighting limitations in spatial reasoning and code generation for VLMs.
Motivation
Vision-language models (VLMs) have been explored for visual programming, where they generate code to solve visual tasks.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.