Benchmark Radar
AI BENCHMARK PROFILE

II-Bench

General AIMultimodal PerceptionCoding & Software Engineering

II-Bench evaluates computer-use agents against low-harm adversarial tasks across three platforms, with 444 examples and a testing framework.

Released
2026-08-03
Readiness
Paper only
Primary field
General AI

Why it matters

Exposes security blind spots in human-in-the-loop defenses for computer-use agents.

Motivation

Computer-use agents (CUAs), which empower large language models to autonomously operate operating systems and the web, are increasingly vulnerable to indirect prompt injection attacks.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.