Study Quantifies Trustworthiness Costs of Domain Adaptation in Small Language Models
An empirical study identified as arXiv:2608.00042 presents the inaugural systematic analysis of trustworthiness costs linked to domain adaptation in small language models (SLMs) across various domains and architectures. This research involved three SLM architectures—TinyLlama 1B, Gemma-2 2B, and Llama 3.2 1B—examining three fields (healthcare, legal, finance), two types of training data (benign and adversarially altered), and four fine-tuning approaches (baseline LoRA, Safety-DPO, Dark Experience Replay, and Task Arithmetic LoRA). The study assesses how these factors influence factual calibration and adversarial resilience. Results indicate that while fine-tuning improves performance, it incurs significant trustworthiness costs, varying by architecture, domain, and method. The implications stress the balance between performance and trustworthiness in critical sectors. The paper can be accessed on arXiv with the identifier 2608.00042.
Key facts
- First systematic cross-domain, cross-architecture study on trustworthiness costs of domain adaptation in SLMs
- Evaluates three SLM architectures: TinyLlama 1B, Gemma-2 2B, Llama 3.2 1B
- Covers three domains: healthcare, legal, finance
- Tests two training-data conditions: benign and adversarially perturbed
- Compares four fine-tuning strategies: baseline LoRA, Safety-DPO, Dark Experience Replay, Task Arithmetic LoRA
- Trustworthiness measured via TruthfulQA MC2 for factual calibration and adversarial robustness
- Paper available on arXiv with identifier 2608.00042
- Highlights trade-offs between performance and trustworthiness in high-stakes SLM deployment
Entities
Institutions
- arXiv