NetlistBench: Benchmarking LLM Reliability in SPICE Netlist Recognition and Manipulation
A new benchmark named NetlistBench has been launched to assess the reliability of large language models (LLMs) in interpreting and altering SPICE netlists, which serve as textual representations of electronic circuits. This benchmark, outlined in a paper on arXiv (2608.12197), seeks to fill a gap in the existing knowledge: even though LLMs are increasingly utilized in circuit design processes, their effectiveness in simulator-related netlist tasks is seldom distinguished from high-level design reasoning. NetlistBench comprises 2,342 instances across 24 task categories, including parameter recognition, connectivity edits, hierarchical functions, equivalence assessments, and long-horizon compound edits. A deterministic structure-aware oracle evaluates model outputs. Testing six non-thinking LLMs revealed significant performance variations based on operation-level structural complexity, with simple local edits reaching 96%–100% accuracy, while device addition accuracy fell to 41%–83%, and equivalence judgment tasks experienced further decline. This benchmark aims to establish a structured, verifiable method for evaluating LLM capabilities in this area, underscoring the necessity for enhanced reliability in complex netlist manipulation tasks.
Key facts
- NetlistBench is a structure-verified benchmark for SPICE netlist recognition and manipulation.
- It contains 2,342 cases across 24 task families.
- Tasks include parameter and connectivity recognition and edits, hierarchical operations, equivalence judgment, and long-horizon compound editing.
- Model outputs are evaluated by a deterministic structure-aware oracle.
- Across six non-thinking LLMs, simple local edits reached 96%–100% accuracy.
- Device addition accuracy dropped to 41%–83%.
- The benchmark is presented in a paper on arXiv with identifier 2608.12197.
- The paper is announced as a cross-type announcement.
Entities
Institutions
- arXiv