Benchmark Radar
AI BENCHMARK PROFILE

SovereignPA-Bench

Robotics & Autonomous SystemsKnowledge & Reasoning

SovereignPA-Bench evaluates user-owned personal agents on 120 sovereignty stress scenarios, measuring task success, alignment, privacy, consent, evidence, manipulation, burden, and auditability in evolving intent and platform mediation contexts.

Released
2026-07-06
Readiness
Paper only
Primary field
Robotics & Autonomous Systems

Why it matters

Personal agents must balance task completion with user sovereignty. This benchmark quantifies trade-offs across privacy, consent, and manipulation, enabling principled agent design.

Motivation

Personal agents are becoming persistent user-owned intermediaries: they remember preferences, filter platform-mediated information, use tools, and negotiate with services.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.