ARTFEED — Contemporary Art Intelligence

OmniDelta: Skill-Driven Token Compression for OmniLLMs

other · 2026-07-29

A new paper on arXiv (2607.25669) introduces OmniDelta, a training-free framework for token compression in Omni-modal Large Language Models (OmniLLMs). OmniLLMs process text, audio, and video, but long token sequences cause high memory and inference costs. Existing methods select important tokens under fixed budgets, neglecting budget allocation. OmniDelta addresses this with skill-driven, intent-aware inter-modal allocation and content-aware intra-modal allocation. It constructs audio and video skill pools to shift retained-token budgets based on query demand, then reallocates budgets over audio segments and video frames using local complexity. The paper shows direct query-to-audio/video similarity is unreliable for budget allocation, and uniform intra-modal budgets can miss key evidence. OmniDelta is training-free and aims to improve efficiency.

Key facts

  • Paper arXiv:2607.25669 proposes OmniDelta for token compression in OmniLLMs.
  • OmniDelta is a training-free, skill-driven framework.
  • It couples intent-aware inter-modal allocation with content-aware intra-modal allocation.
  • It constructs audio and video skill pools to shift retained-token budgets.
  • Budgets are reallocated over audio segments and video frames using local complexity.
  • Direct query-to-audio/video similarity is unreliable for inter-modal budget allocation.
  • Uniform intra-modal budgets can miss key evidence while retaining redundant content.
  • The paper addresses the underexplored budget-allocation problem in token compression.

Entities

Institutions

  • arXiv

Sources