AI BENCHMARK PROFILE
II-Bench
II-Bench evaluates computer-use agents against low-harm adversarial tasks across three platforms, with 444 examples and a testing framework.
- Released
- 2026-08-03
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Exposes security blind spots in human-in-the-loop defenses for computer-use agents.
Motivation
Computer-use agents (CUAs), which empower large language models to autonomously operate operating systems and the web, are increasingly vulnerable to indirect prompt injection attacks.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.