LoRAScan: Detecting Backdoor Prompts in Low-Rank Adapters via Activation Spikes
A recent study published on arXiv (2608.06795) presents LoRAScan, a technique aimed at identifying backdoor prompts in low-rank adapters (LoRA) used in large language models. This cross-type submission highlights a supply-chain vulnerability where untrusted adapters may lead models to produce harmful outputs, such as malicious code, political propaganda, or hidden advertisements when activated by concealed inputs. Current defenses either weaken backdoor signals by integrating adapters with base models or inadequately manage potentially compromised adapters. LoRAScan detects unique latent-space signatures, particularly down-projection activation spikes, associated with trigger inputs to identify harmful adapters. This method is adapter-aware and eliminates the need for separate mitigation. Authored by unnamed researchers, the paper enhances AI security in large language model supply chains.
Key facts
- Paper ID: arXiv:2608.06795
- Announcement type: cross
- Title: LoRAScan: Detecting Backdoor Prompts in Low-Rank Adapters for Large Language Models via Down-Projection Activation Spikes
- Focus: detecting backdoor prompts in LoRA adapters
- Threat: untrusted adapters can cause harmful outputs
- Existing defenses: adapter-agnostic and adapter-aware methods have limitations
- LoRAScan uses down-projection activation spikes as a signature
- Published on arXiv (source URL: https://arxiv.org/abs/2608.06795)
Entities
Institutions
- arXiv