A preprint accepted to NeurIPS 2026 argues that common benchmarks for industrial root cause analysis (RCA) conflate retrieval and reranking failures, according to research posted to arXiv on September 30.

Root cause analysis is critical for preventing safety incidents and costly downtime in complex monitored systems, the authors note, but the standard top-k accuracy metric can obscure whether a method retrieves the correct cause or simply ranks it poorly.

The paper audits four widely used benchmark suites. On tests involving complex faults, the authors report that statistical baselines mis-rank the true cause 79–100% of the time. Graph-based methods, which rely on inferred causal structures, never clearly outperform the best statistical baseline regardless of whether their causal graphs are learned from short fault windows, candidate pools, or multi-day normal-operation data. On simpler benchmarks where the fault manifests strongly at its origin, retrieval is nearly solved at 98–100%.

Guided by the decomposition, the authors build a two-stage pipeline: a multi-signal retriever narrows the candidate pool and a large language model (LLM) reranker scores them. In a fixed configuration, the pipeline matches or beats the best baseline’s top-1 accuracy on all six benchmark suites, by as much as 12 percentage points, with no causal graph or labelled data needed. When all methods rank the same retrieved candidates with the true cause guaranteed present, adding a short system-description document lets the reranker lead the best baseline by 7 to 18 points on every benchmark.

The reported accuracy figures are the authors’ own results and have not been independently replicated or fully peer-reviewed. The paper has been accepted for the NeurIPS 2026 Evaluations & Datasets Track, though the version posted to arXiv has not yet undergone full peer review. The authors have released code online.