Imperfect LLM-Generated RTL Enhances Netlist Representation Learning
A recent study published on arXiv (2603.09161) presents a novel framework that utilizes imperfect Register-Transfer-Level (RTL) code generated by Large Language Models (LLMs) to enhance netlist representation learning. The limited availability of labeled datasets, primarily due to intellectual property concerns and the high costs of annotation, has restricted previous research to small circuits with pristine labels, which hampers scalability. The authors note that even if the RTL produced by LLMs is functionally flawed, the resulting netlists still exhibit structural patterns that reflect the intended functionality. They propose an economical data augmentation and training framework that effectively utilizes these imperfect LLM-generated RTLs as training data, creating a comprehensive end-to-end pipeline. The full paper can be accessed at https://arxiv.org/abs/2603.09161.
Key facts
- Paper arXiv:2603.09161 proposes using imperfect LLM-generated RTL for netlist representation learning.
- Labeled datasets for netlist representation are scarce due to IP protection and annotation costs.
- Existing work focuses on small-scale circuits with clean labels.
- LLM-generated RTL can be produced at scale but is often functionally incorrect.
- Key insight: even imperfect RTL preserves structural patterns indicative of intended functionality.
- Framework is cost-effective and uses data augmentation and training.
- Pipeline is end-to-end from LLM-generated RTL to netlist representation learning.
- Paper is available on arXiv.
Entities
Institutions
- arXiv