ARTFEED — Contemporary Art Intelligence

Perturbation-Based DoS Attack Targets End-to-End Speech Language Models

ai-technology · 2026-08-13

A recent paper on arXiv (arXiv:2608.10405v1) introduces a novel denial-of-service (DoS) attack targeting end-to-end (E2E) speech language models. This attack takes advantage of the models' susceptibility to subtle acoustic disturbances, which can compel them to produce excessively lengthy outputs, resulting in increased computational demands and resource usage. Unlike traditional text-based DoS attacks that depend on prompt engineering techniques, this approach fine-tunes acoustic disturbances to elicit prolonged outputs without modifying the original prompt. The research indicates that earlier studies have mainly concentrated on the security of automatic speech recognition (ASR) and text-to-speech (TTS) systems, leaving E2E speech LLMs' DoS vulnerabilities largely unaddressed. This work emphasizes the necessity for strong defenses in the evolving domain of speech-focused AI systems.

Key facts

  • A perturbation-based DoS attack targets end-to-end speech language models.
  • The attack optimizes imperceptible acoustic perturbations to induce long outputs.
  • Existing text-based DoS attacks cannot be transferred to continuous speech inputs.
  • Prior speech model security research focused on ASR and TTS, not DoS.
  • The paper is available on arXiv with identifier 2608.10405.
  • The attack causes significant computational overhead and resource consumption.
  • The method does not rely on prompt manipulation.
  • The vulnerability of E2E speech LLMs to DoS attacks was previously unexplored.

Entities

Institutions

  • arXiv

Sources