LLMs Predict Their Own Failures to Cut Inference Costs
A recent paper on arXiv (2602.09924v4) investigates the ability of large language models (LLMs) to anticipate their performance on math and coding challenges by examining internal activations prior to task execution. Researchers employed linear probes on these pre-generation activations to predict success specific to various policies, revealing that these probes surpassed traditional metrics such as question length and TF-IDF. Utilizing the E2H-AMC dataset, which compares both human and model results on the same tasks, they found that models possess a unique interpretation of difficulty that differs from human assessments, with this gap widening during prolonged reasoning. Their findings suggest that strategically routing queries among multiple models can enhance performance beyond the best individual model while significantly lowering inference costs. The paper, classified as a replace-cross type, is accessible on arXiv, indicating potential advancements for cost-effective inference in LLMs in practical applications.
Key facts
- Paper: arXiv:2602.09924v4
- Announce Type: replace-cross
- Research investigates predicting LLM success from pre-generation activations
- Linear probes trained on pre-generation activations
- Probes outperform surface features like question length and TF-IDF
- Uses E2H-AMC dataset with human and model performance
- Models encode a model-specific difficulty distinct from human difficulty
- Routing queries across models can exceed best model while reducing inference cost
Entities
Institutions
- arXiv