Too Sure to Be Safe: Model Calibration for Reliable Log Anomaly Detection

August 18, 2026 ยท Grace Period ยท ๐Ÿ› the 2026 IEEE International Conference on Data Mining

โณ Grace Period
This paper is less than 90 days old. We give authors time to release their code before passing judgment.
Authors Bin Li, Dongdong Wang, Siyang Lu arXiv ID 2608.17965 Category cs.LG: Machine Learning Cross-listed cs.AI, cs.SE Citations 0 Venue the 2026 IEEE International Conference on Data Mining
Abstract
Online log anomaly detection is critical for maintaining the reliability of large-scale computing systems. Although recent language model-based log anomaly detectors achieve strong detection performance, their confidence estimates remain poorly calibrated. We show that these detectors frequently assign excessive confidence to incorrect predictions, particularly for anomalous logs under severe class imbalance. Moreover, confidence on erroneous predictions remains persistently high even when conventional calibration metrics indicate good calibration, creating a critical reliability gap for operational monitoring systems. To address this issue, we propose Log Reconstruction and Distance (LoRD), a lightweight post-hoc calibration framework for reliable log anomaly detection. LoRD learns prediction-route-specific reliability models from latent representations of correctly classified validation samples and estimates prediction reliability through route-wise reconstruction distances. Based on the estimated reliability, LoRD selectively recalibrates high-risk predictions to suppress overconfident errors while preserving reliable predictions. Extensive experiments on four large-scale log benchmark datasets and multiple language model-based detectors demonstrate that LoRD consistently improves confidence reliability and substantially reduces overconfident anomaly-related errors without sacrificing anomaly detection performance.
Community shame:
Not yet rated
Community Contributions

Found the code? Know the venue? Think something is wrong? Let us know!

๐Ÿ“œ Similar Papers

In the same crypt โ€” Machine Learning