DRAA: Dynamic Routing Adaptive Alignment Against White-Box Attacks on Large Foundation Models
A recent paper available on arXiv (ID 2608.02674) introduces a defense mechanism known as Dynamic Routing Adaptive Alignment (DRAA) aimed at safeguarding large foundation models (LFMs) against white-box attacks. This cross-type submission highlights the evolving safety challenges, transitioning from black-box jailbreaks to white-box attacks that specifically target and disrupt internal safety neurons or pathways. Current defenses depend on static safety units, which are susceptible to precise route-level assaults. DRAA offers dynamic compensatory routes to ensure strong refusal behavior when safety pathways are compromised. The approach identifies the model's safety route by comparing internal activations from safe and unsafe calibration samples, subsequently masking this route to create causal failure scenarios, which are then analyzed to form failure-aware preference pairs. Extensive testing showcases DRAA's effectiveness, although the abstract does not provide conclusive results. Authored by a team of researchers, this paper is hosted on arXiv, indicating it has not yet been peer-reviewed, and is significant for the AI safety community, especially for those focused on enhancing the alignment of large language models.
Key facts
- Paper arXiv:2608.02674 proposes DRAA framework
- DRAA addresses white-box attacks on large foundation models
- Existing defenses use static safety units or fixed refusal pathways
- DRAA introduces dynamic compensatory routes
- Method identifies safety route via activation contrast
- Masks safety route to induce causal failure cases
- Constructs failure-aware preference pairs
- Experiments demonstrate effectiveness
Entities
Institutions
- arXiv