SparSEEty Attack Exploits LLM Serving Side Channels to Extract Tokens
A novel security vulnerability, referred to as SparSEEty, has been introduced to extract tokens from large language model (LLM) serving systems that utilize activation sparsity. This attack takes advantage of side channels resulting from neuron weight accesses that depend on input. Researchers successfully demonstrated this attack on an LLM serving system secured within an Intel TDX confidential virtual machine (CVM). By utilizing side channels during inference, the attack creates a neuron-activation oracle and subsequently reverses the activation traces to recover input tokens, culminating in a comprehensive token extraction assault. This study reveals a critical security flaw in LLM serving systems that optimize performance by disregarding weight accesses for inactive neurons. The findings are published in a paper on arXiv under identifier 2608.02995.
Key facts
- SparSEEty is a token extraction attack exploiting input-dependent neuron weight accesses.
- The attack targets LLM serving systems that use activation sparsity for optimization.
- It constructs a neuron-activation oracle via side channels during LLM inference.
- The attack inverts activation traces to reconstruct input tokens.
- Demonstrated against an LLM serving system in an Intel TDX confidential virtual machine.
- The research is available on arXiv with ID 2608.02995.
- The attack addresses three key challenges (details not provided in source).
- The vulnerability arises from optimizations that omit weight accesses for inactive neurons.
Entities
Institutions
- Intel
- arXiv