Structured Compositional Reasoning Improves LLM Performance on Compound Logic Questions
A new framework for structured compositional reasoning enhances large language model (LLM) performance on complex logic questions. By breaking down compound answer options into individual atomic answers and employing contrastive hypothesis scoring, this method utilizes an operator-constrained integer linear program to compile scores effectively. Evaluations on the LOGICAL-COMMONSENSEQA and the newly introduced LOGICAL-SATA benchmark show significant improvements, with Macro-F1 scores rising from 48.3 to 77.0 and from 47.0 to 75.6, respectively. Notably, the largest advancements were seen in NEITHER/NOR responses. The research paper is accessible on arXiv under the identifier 2608.12836.
Key facts
- Framework decomposes compound answer options into atomic answers
- Uses contrastive hypothesis scoring for each atomic answer
- Operator-constrained integer linear program composes scores
- Evaluated on LOGICAL-COMMONSENSEQA and new LOGICAL-SATA benchmark
- Macro-F1 improved from 48.3 to 77.0 on LOGICAL-COMMONSENSEQA
- Macro-F1 improved from 47.0 to 75.6 on LOGICAL-SATA
- Largest gains on NEITHER/NOR options
- Paper available on arXiv (2608.12836)
Entities
Institutions
- arXiv