Benchmark Radar
AI BENCHMARK PROFILE

AssertLLM2

Industrial & EngineeringKnowledge & Reasoning

Evaluates LLM generation of SystemVerilog assertions from structured design specifications and RTL, across two tasks: bug-prevention and bug-hunting. The benchmark includes 83 real-world designs with golden and mutated RTL, and assesses syntactic validity, formal provability, coverage, and mutation-based bug detection.

Released
2026-05-26
Readiness
Paper only
Primary field
Industrial & Engineering

Why it matters

Fills a gap in realistic evaluation for assertion generation by using full specifications and buggy RTL, supporting comparison of LLM capabilities for hardware verification tasks.

Motivation

Assertion-based verification (ABV) is a cornerstone of modern hardware design, yet manually translating design intent into formal SystemVerilog Assertions (SVAs) remains labor-intensive and error-prone.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.