QuaSAR: New Method for Stable Activation-Aware Quantization Compensation
A recent publication on arXiv (2608.14149) presents QuaSAR, a technique designed for quantization compensation through stable activation-aware rank truncation. This study tackles a significant issue in current training-free post-training quantization (PTQ) approaches that rely on closed-form residual compensation to enhance model accuracy. In low-bit configurations (W4A4), the goodness-of-fit gating mechanism, which identifies layers for compensation based on positive residual fit scores, struggles to differentiate between unpredictable quantization errors and failures in numerical solvers. The authors reveal that rank-deficient input activations can result in ill-conditioned or numerically singular Gram matrices, leading to instability in the closed-form solver and erroneously negative fit scores. Consequently, layers that could benefit from compensation are incorrectly deemed uncompensable and eliminated. QuaSAR counters this by utilizing stable activation-aware rank truncation for dependable compensation. This method is crucial for enhancing the accuracy of quantized models without further training, particularly for large models on devices with limited resources. The findings underscore the necessity of numerical stability in PTQ and offer a viable solution. The paper can be accessed at https://arxiv.org/abs/2608.14149.
Key facts
- Paper arXiv:2608.14149 introduces QuaSAR.
- QuaSAR addresses instability in goodness-of-fit gating for PTQ.
- Under W4A4, rank-deficient activations cause ill-conditioned Gram matrices.
- Existing gates misclassify layers as uncompensable due to solver failure.
- QuaSAR uses stable activation-aware rank truncation.
- Method is training-free and aims to restore accuracy.
- Paper is on arXiv preprint server.
- Source URL: https://arxiv.org/abs/2608.14149
Entities
Institutions
- arXiv