AI's Internal Representation of Falsehood vs Impossibility: A Study of Gemma 3 4B IT
A recent study on arXiv (2608.12852) explores whether AI can tell apart false statements from those that are impossible. The research used the multimodal open-weight model Gemma 3 4B IT, testing it with 85 prompts across 17 philosophical categories. Each prompt was framed as a truth or different types of falsehoods. Interestingly, the model often mistook contingent falsehoods for contradictions, labeling 12 out of 15 false statements as contradictions. However, a linear truth probe could effectively differentiate true from impossible statements (AUC 0.93), but struggled with impossible versus false statements (AUC 0.20). On the other hand, using an impossibility probe, the model distinguished necessary from contingent falsehoods perfectly at AUC 1.00. This suggests that while the model's outputs may mix these concepts, its internal workings might actually grasp them, offering insights into AI's understanding of language and logic. The study, by an unnamed author, is also significant for future AI design and evaluation.
Key facts
- Study on arXiv:2608.12852
- Model: Gemma 3 4B IT
- 85 prompts from 17 philosophical families
- 15 topics with five modalities each
- Model labels 12 of 15 false statements as 'contradiction'
- Truth probe AUC 0.93 for impossible vs true
- Impossibility probe AUC 1.00 for necessary vs contingent falsehood
- Probe peaks at layer 15 with balanced accuracy 0.97
Entities
Institutions
- arXiv