Benchmark Radar
AI BENCHMARK PROFILE

TurtleAI

General AIMultimodal PerceptionCoding & Software Engineering

TurtleAI is a benchmark of 823 tasks for visual programming in Turtle Graphics, evaluating models on perceiving geometric patterns and synthesizing Python code.

Released
2026-06-02
Readiness
Paper only
Primary field
General AI

Why it matters

Bridges the gap in education-oriented visual programming evaluation, highlighting limitations in spatial reasoning and code generation for VLMs.

Motivation

Vision-language models (VLMs) have been explored for visual programming, where they generate code to solve visual tasks.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.