StepJack
StepJack evaluates computer-use agents against multi-step indirect prompt injection attacks, where adversarial goals are decomposed into innocuous sub-steps across a chain of pages. It comprises 480 test examples across platforms and instruction types, with attack success rate as the primary metric.
- Released
- 2026-08-06
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
The benchmark addresses the evaluation gap in agent safety against sophisticated, staged prompt injection attacks that previous single-step benchmarks fail to capture, offering a standardized way to assess and compare defenses for computer-use agents.
Motivation
Computer-use agents (CUAs) face a growing threat from indirect prompt injection, where adversarial instructions are planted in the environment such as web pages.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.