ARTFEED — Contemporary Art Intelligence

HIVE Study: Voice Input Perturbations Reduce LLM Accuracy More Than Keyboard

ai-technology · 2026-08-06

A new study from arXiv (ID 2608.03970) introduces HIVE (Human Input-Variation Engine), a suite of perturbations simulating voice transcription and QWERTY keyboard input errors, to evaluate the robustness of instruction-tuned large language models (LLMs). The study finds that voice transcription perturbations, including disfluency from conventional transcription and restructuring from AI-backed dictation, lower accuracy across all tested models, with the structure of the transcription (not fillers) being the primary cause. Keyboard perturbations are less costly, and models can absorb many before accuracy drops. Both error types trace back to the survival of question tokens: destroying tokens is what hurts performance. The paper presents seven findings in total, though only three are detailed in the abstract. The research highlights the distinct signatures of typing and speaking as input channels and their differential impact on LLM performance, with implications for human-AI interaction design. The study is authored by researchers affiliated with the arXiv preprint server, though specific author names are not provided in the available content. The paper was announced as new on arXiv, with the identifier 2608.03970, suggesting a publication date in August 2026 (based on the arXiv naming convention). The research contributes to understanding how input modalities affect AI model reliability, particularly in voice-based interfaces.

Key facts

  • HIVE (Human Input-Variation Engine) is a suite of voice transcription and QWERTY keyboard perturbations.
  • Voice transcription perturbations lower accuracy across every instruction-tuned model tested.
  • The structure of the transcription, not fillers, carries the cost in voice perturbations.
  • QWERTY keyboard perturbations cost less, and models absorb many before accuracy falls.
  • Both perturbation types trace back to how many question tokens survive the perturbation.
  • The study presents seven findings, with three detailed in the abstract.
  • The paper is available on arXiv with ID 2608.03970.
  • The research focuses on human input channels: typing and speaking.

Entities

Institutions

  • arXiv

Sources