LLM Reality Monitoring Fails When Conversation Memory Is Stressed
A recent investigation published on arXiv examines whether large language models (LLMs) can differentiate between their own outputs and user inputs, a skill known as reality monitoring. Researchers conducted two experiments involving six LLMs and discovered that the accuracy of source attribution is influenced by the structure of conversational memory. When memory demands are low, models demonstrate nearly flawless accuracy for their self-generated content. However, this shifts to a fragile advantage for external items when episodic delays occur. Feedback indicates two types of errors: in certain models, internal and external evaluations interchange, while in others, accuracy increases but confidence becomes disconnected from correctness—issues not captured by current benchmarks. This pattern highlights the importance of active parameter count over total count. The results imply that as AI systems engage in autonomous, multi-turn interactions, their failure to track sources may result in misinterpreting their own errors as user-provided information, potentially leading to hallucinations and confabulation.
Key facts
- arXiv paper 2607.23927
- Two experiments with six LLMs
- Ceiling accuracy for self-generated content under minimal memory demands
- Reverses to external-item advantage with episodic delay
- Two failure modes: judgment swap and confidence decoupling
- Active parameter count implicated, not aggregate
- Links to hallucinations and confabulation
- Implications for autonomous multi-turn AI roles
Entities
Institutions
- arXiv