Benchmark Radar
AI BENCHMARK PROFILE

NetlistBench

Robotics & Autonomous SystemsRobotics & Embodied Intelligence

NetlistBench evaluates LLM reliability in recognizing and manipulating SPICE netlists, covering parameter and connectivity recognition, edits, hierarchical operations, equivalence judgment, and compound editing, with deterministic structure-aware scoring.

Released
2026-08-12
Readiness
Paper only
Primary field
Robotics & Autonomous Systems

Why it matters

It addresses the gap in assessing LLM reliability for simulator-facing netlist tasks, distinct from high-level design reasoning, and provides a structured evaluation to guide trustworthy circuit design automation.

Motivation

Large Language Models (LLMs) are increasingly used in circuit design workflows, yet their reliability on simulator-facing SPICE netlist recognition and manipulation remains poorly understood and is rarely separated from high-level design reasoning.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.