Chinese AI Firms Struggle with Chip Shortage as Inference Demand Surges
Chinese artificial intelligence firms are encountering significant computing limitations as they shift from developing models to large-scale implementation. Domestic chips are inadequate for complex tasks such as coding. Although efforts have been made to optimize inference workloads for local hardware, the generation of high-quality tokens still depends on Nvidia processors, which are in short supply due to export limitations. Guan Jiawei, vice-president of the inference optimization startup Approaching.AI, pointed out a 'bipolarisation' in demand, with a vast disparity between the need for high-quality tokens and their availability. Domestic processors can only manage lower-quality outputs, hindering commercial viability. The situation is worsened by a dramatic increase in token usage, with China's daily token calls surpassing 140 trillion in March, a rise of over 1,000-fold since early 2024, as reported by the National Data Administration.
Key facts
- Chinese AI companies are optimizing software to cope with surging demand for inference.
- Inference can be adapted to domestic hardware, but complex tasks like coding still require Nvidia chips.
- Guan Jiawei, vice-president of Approaching.AI, notes demand for high-quality tokens far outstrips supply.
- Domestic processors cannot reliably deliver stringent performance metrics for high-tier tasks.
- High-quality tokens still depend on Nvidia due to weak monetisation of low-quality tiers.
- China's average daily token calls exceeded 140 trillion in March, up more than 1,000-fold from early 2024.
- The compute squeeze is exacerbated by AI becoming more agentic.
- Restricted access to Nvidia processors limits the pool of high-end chips.
Entities
Institutions
- Approaching.AI
- National Data Administration
- Nvidia
Locations
- China