Geometric Error Detection for Pre-trained Perception Models Without Domain Knowledge
A new arXiv preprint (2608.04190) proposes a method to make pre-trained perception models more robust to distributional shift in novel environments. The authors demonstrate that combining models via majority voting or similar ensemble methods does not recover accuracy under such shifts, and can be brittle to coordinated failures. Prior metacognitive approaches relied on hand-authored domain-knowledge cues, such as object-size priors or segmentation masks, which do not transfer to genuinely novel scenes. The new method, termed Label Vector Pools (LVP), exploits vector-space geometry: per-model LVPs are built from each model's own training embeddings, and error-detection rules are derived from the geometry of detections relative to training-determined prototypes. This approach achieves parity with domain-knowledge rules to within 0.002 F1 on the test set, without requiring any domain knowledge. The method remains neurosymbolic, and the geometric rules share a single log... (truncated for brevity). The paper is available on arXiv under the identifier 2608.04190.
Key facts
- The paper is an arXiv preprint with identifier 2608.04190.
- It addresses accuracy degradation of pre-trained perception models under distributional shift.
- Majority voting and similar combiners trade recall for precision and are brittle to coordinated failures.
- Prior metacognitive methods rely on hand-authored domain-knowledge cues that do not transfer to novel scenes.
- The proposed method, Label Vector Pools (LVP), uses per-model training embeddings to build error-detection rules.
- Geometric rules derived from detections relative to training-determined prototypes reach parity with domain-knowledge rules.
- The performance difference is within 0.002 F1 on the test set.
- The approach remains neurosymbolic, sharing a single log... (incomplete).
Entities
Institutions
- arXiv