WeSCE
WeSCE is a benchmark for quantifying security drift in LLM-driven code editing. It consists of 400 executable programs derived from real-world code, covering feature addition, removal, bug fixing, and refactoring. The benchmark proposes a continuous risk representation and drift measures that capture changes in overall risk, worst-case severity, and vulnerability distribution.
- Released
- 2026-08-15
- Readiness
- Paper only
- Primary field
- Cybersecurity
Why it matters
Code editing with weak-security constraints can introduce vulnerabilities, but existing evaluations lack a systematic measure. WeSCE offers a standardized way to quantify security drift, potentially aiding in selecting models and prompts that minimize security risks during code modifications.
Motivation
In this work, we introduce WeSCE, a benchmark for quantifying security drift in code editing under weak-security constraints, where tasks specify only functional objectives without explicit security requirements.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.