Benchmark Radar
AI BENCHMARK PROFILE

DelegateCI-Bench

Health & Life SciencesKnowledge & Reasoning

DelegateCI-Bench evaluates privacy-conscious query rewriting for LLM delegation, with 3,167 samples combining synthetic data across 20 task types, real user queries from WildChat, and a medical challenge set. Systems rewrite queries to suppress non-essential sensitive information.

Released
2026-06-02
Readiness
Paper only
Primary field
Health & Life Sciences

Why it matters

Addresses a gap in privacy benchmarks by focusing on task-based necessity rather than type-based PII redaction. Measures privacy-utility tradeoff in delegation, supporting safer LLM use.

Motivation

As LLMs become increasingly woven into everyday workflows, user queries sent to cloud hosted LLMs routinely mix task-essential content with task non-essential sensitive disclosures, yet type based PII redaction is context agnostic and may raise two issues: over disclosing untyped sensitive context and over removing answer bearing spans.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.