Wnuan: A Three-Stage Pipeline for Enterprise QA with Proprietary Knowledge
A recent study on arXiv introduced Wnuan, an innovative three-phase approach aimed at improving question answering within proprietary enterprise knowledge while maintaining general capabilities. The framework generates specific guidance from existing documents, employs supervised fine-tuning with general data replay, and integrates reinforcement learning to address remaining errors. In a benchmark test involving 707 questions, the primary 32B model enhanced the acceptable-answer rate from 52.76% to 80.06% after fine-tuning, and to 91.51% with reinforcement learning. Furthermore, a 100-update protocol indicated that residual-error sampling surpassed both full-pool and random sampling by 3.11 and 2.97 points, respectively.
Key facts
- Wnuan is a three-stage pipeline for enterprise QA.
- Stages: task-oriented supervision, SFT with general-data replay, RL on residual errors.
- WnuanBench has 707 questions.
- Primary 32B route: AAR improves from 52.76% to 80.06% after SFT, to 91.51% after RL.
- Residual-error sampling outperforms full-pool and size-matched random sampling by 3.11 and 2.97 points.
- General-benchmark average decreases by 5.17 points, concentrated in instruction following.
- Automatic evaluation ensemble agrees with authoritative assessment.
- Paper ID: arXiv:2608.01862.
Entities
Institutions
- arXiv