LoRD: A Calibration Fix for Overconfident Log Anomaly Detectors
Log anomaly detectors, which utilize language models and are essential for monitoring extensive computing systems, exhibit a troubling level of overconfidence. An arXiv preprint (2608.17965) reveals that these detectors often assign too much confidence to incorrect predictions, particularly in the presence of severe class imbalance in anomalous logs. The research highlights a significant reliability issue: high confidence in wrong predictions persists despite traditional calibration metrics suggesting otherwise. To address this, the authors introduce LoRD (Log Reconstruction and Distance), a streamlined post-hoc calibration framework. LoRD develops reliability models based on the latent representations of accurately classified validation samples, providing a practical calibration solution for operational monitoring without the need for extensive retraining. The preprint addresses the disparity between detector efficacy and reliable confidence.
Key facts
- arXiv preprint 2608.17965 addresses miscalibration in log anomaly detection.
- Language model-based detectors assign excessive confidence to incorrect predictions.
- Confidence on errors remains high even when calibration metrics look good.
- Researchers propose LoRD (Log Reconstruction and Distance) as a post-hoc calibration framework.
- LoRD uses latent representations of correctly classified validation samples.
- The work targets operational monitoring in large-scale computing systems.
- The paper is classified as cross-type on arXiv.
- Severe class imbalance aggravates the overconfidence problem.
Entities
Institutions
- arXiv