Benchmark Radar
AI BENCHMARK PROFILE

PPT-Eval

General AIKnowledge & ReasoningMicrosoft

PPT-Eval is a benchmark of 120 PowerPoint tasks across 12 files for computer-use agents. It covers content creation and editing, with rubric-based evaluation that awards partial credit and provides natural language feedback.

Released
2026-06-30
Readiness
Inspectable
Primary field
General AI

Why it matters

Provides a realistic, multimodal testbed for computer-use agents. The rubric-based scoring captures partial progress and correlates with human judgment, enabling nuanced comparison.

Motivation

Creating and editing slides is a rich, multimodal activity that is ubiquitous in professional and educational settings, making it an ideal testbed for real-world computer-use agents.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.