LLM-Generated Commonsense Axioms for NLI: A Reference-Free Evaluation
A recent study published on arXiv (2507.15100) explores the capability of Large Language Models (LLMs) to produce factual commonsense axioms for Natural Language Inference (NLI) and assesses their effectiveness using the SNLI and ANLI benchmarks. Utilizing Llama-3.1-70B and gpt-oss-120b, the research presents a novel reference-free approach through an LLM-as-Judge framework, as traditional factuality metrics are inadequate for commonsense axioms that do not have clear textual references. Results indicate a significant disparity between the models: gpt-oss-120b predominantly generates accurate axioms, while Llama tends to produce more inaccuracies. Additionally, the study examines three prompting strategies, emphasizing the potential and challenges of leveraging LLMs for commonsense knowledge in NLI tasks.
Key facts
- The study is from arXiv:2507.15100.
- It evaluates LLMs on generating commonsense axioms for NLI.
- Benchmarks used: SNLI and ANLI.
- Models tested: Llama-3.1-70B and gpt-oss-120b.
- A reference-free LLM-as-Judge method is introduced.
- gpt-oss-120b outperforms Llama in axiom accuracy.
- Three prompting pipelines are evaluated.
- The abstract is truncated, missing the full description of the third pipeline.
Entities
Institutions
- arXiv