ARTFEED — Contemporary Art Intelligence

Explainable Online Failure Prediction for Linux Operating Systems

ai-technology · 2026-08-04

A paper on arXiv (2608.00651) presents a practical approach to Online Failure Prediction (OFP) in Linux operating systems, emphasizing the need for diagnostic insight beyond mere prediction accuracy. The authors argue that without explainability, operators cannot trust alerts or decide on responses, and high accuracy may be due to workload-specific noise. They built an explainable OFP pipeline combining consensus-based feature selection, temporal onset analysis, subsystem-level causal analysis, and complementary diagnostic mechanisms. Under strict cross-workload conditions with frozen training artifacts, the pipeline achieved 91-94% detection on unseen workloads without retraining. The study highlights the importance of interpretability in AI-driven system monitoring.

Key facts

  • Paper on arXiv: 2608.00651
  • Focus on Linux operating systems
  • Combines consensus-based feature selection, temporal onset analysis, subsystem-level causal analysis, and complementary diagnostic mechanisms
  • Achieved 91-94% detection on unseen workloads without retraining
  • Strict cross-workload conditions with frozen training artifacts
  • Emphasizes diagnostic insight for practical adoption
  • Addresses risk of models exploiting workload-specific noise
  • Published as a cross-type announcement

Entities

Institutions

  • arXiv

Sources