Pretrained Model Exposure Amplifies Jailbreak Risks in Finetuned LLMs
A recent study published on arXiv (ID 2512.14751) examines whether fine-tuned large language models (LLMs) carry over jailbreak vulnerabilities from their pretrained counterparts. Released as a replace-cross announcement, the research introduces a realistic threat model where an attacker has complete access to a publicly available pretrained LLM but cannot access its proprietary fine-tuned versions. The empirical findings indicate that adversarial prompts optimized for the pretrained model are highly transferable to its fine-tuned derivatives, revealing inherited vulnerabilities. Additionally, representation-level probing indicates that these transferable prompts are linearly separable within the hidden states of the pretrained model, suggesting that relevant structural information is encoded from the outset. This highlights significant security concerns regarding the release of pretrained models, emphasizing the necessity for strong protections throughout the pretraining-to-finetuning process.
Key facts
- Study on arXiv ID 2512.14751
- Announcement type: replace-cross
- Focuses on finetuned LLMs inheriting jailbreak vulnerabilities
- Threat model: attacker has full access to pretrained model, no access to finetuned derivatives
- Adversarial prompts transfer from pretrained to finetuned models
- Representation-level probing shows linear separability in hidden states
- Findings highlight security risks of releasing pretrained models
Entities
Institutions
- arXiv