Benchmark Radar
AI BENCHMARK PROFILE

StakeBench

CybersecurityFinance & EconomicsKnowledge & Reasoning

StakeBench is a stakeholder-centric benchmark for prompt-injection attacks in web agents for online shopping. It decomposes risk into 12 attack objectives across three stakeholder classes, with 264 adversarial cases across 12 product categories.

Released
2026-06-11
Readiness
Paper only
Primary field
Cybersecurity

Why it matters

Prompt-injection risk is victim-dependent. StakeBench captures asymmetric consequences for different stakeholders, which is overlooked by attack-centric evaluations.

Motivation

LLM-based web agents are increasingly deployed in real-world settings such as e-commerce, where they interact extensively with untrusted web content while executing actions that carry direct financial consequences.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.