ARTFEED — Contemporary Art Intelligence

DeepSeek's DSpark Speeds AI Inference by 85%, Cuts Chip Strain

ai-technology · 2026-06-28

Chinese AI start-up DeepSeek has unveiled DSpark, a speculative decoding framework that accelerates AI inference by up to 85%, addressing bottlenecks in serving large models. The framework uses a lightweight draft model to propose candidate responses, which are then verified in batches by a larger model, reducing token-by-token delays that lead to low GPU utilization and high user wait times. DSpark also employs a semi-autoregressive generation method to produce small token chunks instead of single tokens, and a confidence-based scheduling system that dynamically adjusts verification based on computing demand. The announcement comes amid China's push to overcome US AI curbs on advanced chips. DeepSeek published the research on Saturday.

Key facts

  • DeepSeek unveiled DSpark, a speculative decoding framework.
  • DSpark speeds up AI inference by up to 85%.
  • It uses a lightweight draft model to propose candidate responses.
  • A larger model verifies the candidates in batches.
  • Semi-autoregressive generation produces small token chunks.
  • Confidence-based scheduling adjusts verification dynamically.
  • The framework reduces GPU utilization and user-perceived waiting time.
  • Research was published on Saturday amid China's push to overcome US AI curbs.

Entities

Institutions

  • DeepSeek
  • Shutterstock

Locations

  • China
  • United States

Sources