NANQ: Noise-Aware Mixed-Precision Quantization for Analog CIM
The newly introduced framework, NANQ, tackles the issue of hardware noise in analog compute-in-memory (CIM) systems utilized for neural network inference. While analog CIM is known for its energy efficiency, it faces challenges from device variation and read noise, which adversely affect low-bit quantized models. Current quantization techniques focus on minimizing ideal quantization errors but overlook the hardware noise floor, resulting in suboptimal precision distribution. NANQ effectively models weight noise that depends on magnitude, derived from actual responses of an eFlash CIM array, and transforms this noise profile into an adaptive quantization density. This allows for finer resolution in low-noise areas and prevents ineffective precision in high-noise zones. It also determines layer-wise bit-widths by pinpointing each layer's precision saturation threshold under hardware noise. Experimental results on an eFlash CIM SoC reveal that with 2-bit weight-magnitude quantization, NANQ significantly enhances accuracy compared to existing approaches. The research paper can be found on arXiv under identifier 2608.02700.
Key facts
- NANQ is a noise-aware mixed-precision non-uniform quantization framework for analog CIM.
- It models magnitude-dependent weight noise from measured responses of an eFlash CIM array.
- It converts noise profile into adaptive quantization density.
- It assigns layer-wise bit-widths by identifying precision saturation points under hardware noise.
- On-chip experiments on an eFlash CIM SoC show improvements under 2-bit weight-magnitude quantization.
- The paper is on arXiv with identifier 2608.02700.
- Analog CIM enables energy-efficient neural network inference but suffers from device variation and read noise.
- Existing CIM-oriented quantization methods ignore hardware noise floor.
Entities
Institutions
- arXiv