Janus
JANUS is a benchmark with 160 scenarios across 8 domains, each providing a fixed pool of favorable and adverse facts and paired neutral and goal-directed prompts, to evaluate fact-grounded, goal-conditioned pragmatic distortion in LLM outputs.
- Released
- 2026-06-09
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Existing benchmarks primarily detect direct deception, missing subtler misleading communication that stays factually accurate. JANUS measures whether LLMs distort net impressions when incentivized, offering a more practical assessment of risks in real-world applications where selective presentation of facts can mislead stakeholders.
Motivation
LLM deception is often evaluated through direct markers such as fabricated claims, explicit lies, or strategic concealment.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.