HyWA: Hypernetwork-Based Personalized VAD for Voice Assistants
A new method called HyWA enables personalized voice activity detection (PVAD) for full-duplex voice assistants without altering the underlying VAD architecture. Conventional VADs trigger on any speech, causing unwanted activations from nearby conversations or assistant playback. HyWA uses a hypernetwork to generate speaker-conditioned weights at enrollment, preserving the original model's acoustic interface and inference topology. This approach reduces engineering and requalification costs associated with architectural changes. The method is detailed in arXiv paper 2510.12947, announced as a replace-cross update. HyWA requires no per-user optimization, making it efficient for deployment. The work addresses a key limitation in voice-assistant pipelines for smart devices, improving user experience and computational efficiency.
Key facts
- HyWA is a hypernetwork-based weight-adaptation method for personalized voice activity detection (PVAD).
- It converts an established VAD into a PVAD while preserving its acoustic interface and inference topology.
- HyWA generates speaker-conditioned weights once at enrollment and requires no per-user optimization.
- Conventional VADs respond to speech from any speaker, leading to unwanted triggers and wasted resources.
- Existing PVAD methods require architectural changes that increase engineering and requalification costs.
- The paper is available on arXiv with ID 2510.12947.
- The announcement type is replace-cross.
- HyWA targets full-duplex voice assistants in smart devices.
Entities
Institutions
- arXiv