ARTFEED — Contemporary Art Intelligence

Probes Detect Corrupted Context but Fail to Predict Errors in Language Models

ai-technology · 2026-08-11

A new study on arXiv (2608.07528) reveals a gap between how well linear probes can detect corruption in language models and their ability to predict errors. While these probes can accurately identify corrupt contexts, they fail to assess the accuracy of answers in multi-step arithmetic problems. The research, which included pre-registered hypotheses, found that models using structured confidence formats often settle on two similar error rates. Furthermore, the persistence of probes over multiple steps doesn't effectively differentiate between right and wrong answers, challenging the 'persistence beats peak' idea. This pattern is seen across different model types. The effectiveness of probe-based monitoring varies based on the model and error type, with branch-and-pick showing promise, especially with Llama-3.1-8B.

Key facts

  • Linear probes detect corrupted context with near-perfect accuracy.
  • Probe detection does not translate into reliable failure prediction.
  • The dissociation is observed across multi-hop arithmetic chains.
  • Structured confidence formats collapse to two values with indistinguishable error rates.
  • Probe persistence across hops fails to separate correct from incorrect outcomes.
  • The pre-registered 'persistence beats peak' hypothesis was refuted.
  • The pattern generalises across model families including reasoning models.
  • Branch-and-pick is net-positive across models and uniquely non-breaking on Llama-3.1-8B (4 rescued, 0 broken).

Entities

Institutions

  • arXiv

Sources