Logit-Based Energy Scoring Improves Scientific Hypothesis Ranking in LLMs
A recent arXiv preprint, titled 'Do LLMs Know a Good Hypothesis When They See One?' (2608.17270), critiques traditional methods for assessing scientific hypotheses produced by large language models. The authors contend that current evaluation techniques, which either utilize LLMs as evaluators or depend on semantic similarity, often favor well-known concepts over innovative ones. To counter this, they propose a logit-based energy scoring system that measures a model's inherent confidence in a hypothesis without comparative evaluation. Testing seven language models against 1,323 papers from 12 fields, they found that intrinsic scoring achieved a 33.0% Hit@1, surpassing the 16.6% from prompted listwise ranking. The best-performing configuration, a 1-billion-parameter model with logit-based energy scoring, reached 53.1%. The results indicate that this intrinsic scoring method could enhance the reliability and scalability of hypothesis ranking in AI-driven scientific processes. The preprint can be accessed at https://arxiv.org/abs/2608.17270.
Key facts
- Paper proposes logit-based energy scoring for scientific hypothesis ranking
- Benchmarked seven language models on 1,323 papers across 12 disciplines
- Each paper paired with its correct hypothesis and fifteen incorrect alternatives
- Intrinsic scoring reached 33.0% Hit@1 pooled, versus 16.6% for prompted listwise ranking
- Strongest configuration: 1-billion-parameter model with logit-based scoring, 53.1% Hit@1
- Maximum performance selected post hoc across 14 model-by-scorer combinations
- Existing methods use LLM-as-judge or semantic similarity, favoring familiar ideas
- Study addresses trustworthy AI-enabled scientific workflows
Entities
Institutions
- arXiv