FACTWASH: Open-Source Tool Detects AI Rewrites That Erase Evidence
An open-source tool named factwash has been introduced by researchers to address a specific failure mode in AI systems, known as 'factwashing.' This issue arises when AI rewrites maintain assertions but eliminate the supporting evidence, such as sources, certainty, and temporal validity. Factwashing can convert conversations into memories or documents into answers, potentially misrepresenting hearsay as established truth. The tool functions deterministically, utilizing named flags and evidence instead of an LLM judge, and was created to determine when a simple check is adequate versus when a more advanced model is necessary. The findings are published in a paper on arXiv (ID: 2608.03372), marking progress toward enhancing the transparency of AI-generated content.
Key facts
- Factwash is an open-source write-time gate that catches factwashing deterministically.
- Factwashing is the failure mode where AI rewrites keep a claim but remove what made it checkable.
- The tool uses named flags and evidence rather than an LLM judge.
- Explicit negation cues are nearly enumerable, and a word list reaches 0.91 F1 on untuned text.
- Hedging and attribution have open-ended realizations, causing vocabulary approaches to plateau near half recall.
- A one-question LLM witness recovers +17 and +15 points of cue-detection recall at equal precision.
- The research is published on arXiv with ID 2608.03372.
- The tool's deployment may only lower a verdict, not fully eliminate factwashing.
Entities
Institutions
- arXiv