Benchmark Radar
AI BENCHMARK PROFILE

Vul4Py

CybersecurityCoding & Software Engineering

Vul4Py evaluates automated vulnerability repair in Python across 100 real vulnerabilities from 60 open-source projects, with paired exploit and functional oracles.

Released
2026-08-01
Readiness
Paper only
Primary field
Cybersecurity

Why it matters

It addresses the gap of missing functional regression checks in existing Python AVR benchmarks, providing a more reliable comparison of repair methods.

Motivation

Automated Vulnerability Repair (AVR) has advanced rapidly across program analysis, machine learning, and Large Language Models (LLMs), but a verifiable, head-to-head comparison of AVR approaches on Python is still missing.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.