ReQuant: Fixed-Grid Discrete Refinement for Post-Training Quantization
A new arXiv paper introduces ReQuant, a backpropagation-free fixed-grid refinement procedure for post-training quantization (PTQ) of large language models. The method addresses a gap in existing PTQ techniques, which typically treat quantized weights as final after initial optimization. ReQuant operates as a complementary stage, iteratively revisiting discrete weight assignments on a fixed quantization grid to reduce mean squared reconstruction error while preserving the quantized format. It is agnostic to the PTQ initializer and accepts any existing quantized model as a starting point. The paper is available on arXiv under the identifier 2608.07019.
Key facts
- ReQuant is a backpropagation-free fixed-grid refinement procedure for post-training quantization.
- It targets large language models to reduce memory and computational cost.
- Existing PTQ methods typically treat quantized weights as final after initial optimization.
- ReQuant iteratively revisits discrete weight assignments on a fixed quantization grid.
- Accepted updates strictly reduce mean squared reconstruction error.
- It preserves the quantized format and remains on the original grid.
- ReQuant is agnostic to the PTQ initializer.
- The paper is available on arXiv with identifier 2608.07019.
Entities
Institutions
- arXiv