Benchmark Radar
AI BENCHMARK PROFILE

StabilityBench

General AIMathematics & Formal Sciences

StabilityBench is a benchmark operator that converts single-turn benchmark queries into multi-turn interaction histories with injected user simulations, evaluating LLM performance stability across demographic proxies and sycophantic baits on existing benchmarks.

Released
2026-07-17
Readiness
Paper only
Primary field
General AI

Why it matters

Static benchmarks may not capture real-world conversational variability; StabilityBench highlights performance instability under realistic multi-turn conditions, motivating more realistic evaluation settings.

Motivation

AI Assistants are increasingly deployed in high-stakes settings, such as healthcare or government services.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.