Legal Deception Detection: Benchmarking NLP Models Against General-Domain State-of-the-Art
A recent preprint on arXiv (2607.29066) offers an extensive survey and comparative study of NLP-driven Automatic Deception Detection (ADD), particularly within the legal sector. The research traces the transition from feature-based machine learning to Large Language Model (LLM) methodologies and performs a consolidated empirical assessment across seven datasets—comprising two from the legal field and five from general domains. The analysis compares six fine-tuned transformer models alongside seven LLMs using four prompting techniques. Findings reveal significant domain sensitivity: fine-tuned models perform better in data-abundant general domains, whereas few-shot LLMs are effective in resource-scarce legal contexts. Importantly, Chain-of-Thought prompting frequently falls short compared to direct classification. These results highlight the necessity for domain-specific strategies in deception detection, which is vital for legal contexts, law enforcement, and online security. The study is authored by a team of researchers and can be accessed on arXiv.
Key facts
- The paper is titled 'Semantics of Subterfuge: Benchmarking Legal Deception Detection Against General-domain State-of-the-Art'.
- It is published on arXiv with identifier 2607.29066.
- The study surveys and compares NLP-based Automatic Deception Detection (ADD) methods.
- It evaluates six fine-tuned transformer models and seven LLMs.
- Seven datasets are used: two legal and five general-domain.
- Four prompting strategies are tested.
- Fine-tuned models perform best in general domains, while few-shot LLMs are competitive in legal settings.
- Chain-of-Thought prompting often underperforms direct classification.
Entities
Institutions
- arXiv