ARTFEED — Contemporary Art Intelligence

GRM: Utility-Aware Jailbreak Attacks on Audio LLMs via Gradient-Ratio Masking

ai-technology · 2026-08-10

A recent paper on arXiv (ID: 2604.09222v2) presents GRM, a method for jailbreak attacks that takes utility into account, specifically aimed at Audio Large Language Models (ALLMs). This research fills a void in current perturbation-based jailbreak techniques, which often fail to manage the frequency bands involved in the perturbation. While these perturbations can trigger unsafe outputs, applying a universal perturbation across various inputs may reduce effectiveness in standard tasks, potentially revealing the attack to users or automated detection systems. The study found that varying the perturbation coverage from partial to full-band affects the Jailbreak Success Rate (JSR) in a non-linear way, with increased utility degradation as coverage expands. The authors suggest that targeted frequency bands can produce more effective attacks with less utility loss. Consequently, they introduce GRM, a frequency-masked, utility-focused approach. This work falls under AI and machine learning, highlighting concerns for the security of audio-based AI systems. The announcement is labeled 'replace-cross', indicating a revised version, and the research team is affiliated with unspecified institutions. The paper is accessible on arXiv.

Key facts

  • Paper ID: arXiv:2604.09222v2
  • Announcement type: replace-cross
  • Proposes GRM (Gradient-Ratio Masking) for jailbreak attacks on Audio LLMs
  • Existing perturbation-based jailbreaks do not control frequency bands
  • Full-band perturbations degrade utility and leave behavioral footprints
  • JSR varies non-monotonically with coverage; utility degradation grows with coverage
  • Selected bands can yield stronger attacks with less utility degradation
  • Research focuses on Audio Large Language Models (ALLMs)

Entities

Institutions

  • arXiv

Sources