NetInjectBench
NetInjectBench evaluates LLM agents for network operations under indirect prompt injection. It comprises 130 scenarios: 40 benign, 40 weak-attack, 40 strong-attack, and 10 approved high-impact changes, with separated untrusted artifact text, trusted policy metadata, and evaluation labels. Scoring measures unsafe tool-action rate, usefulness, and overblocking across models and defenses.
- Released
- 2026-07-11
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Network operations increasingly rely on tool-using LLM agents, which are vulnerable to indirect prompt injections via untrusted artifacts. This benchmark quantifies safety and usefulness trade-offs in a realistic network-domain environment, providing a concrete means to assess defense effectiveness and authorization boundaries.
Motivation
Tool-using large language model (LLM) agents are attractive for network operations, but tickets, alerts, logs, runbooks, and ChatOps messages can carry indirect prompt injections.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.