Black-Box Language Model Adaptation via Logit Bias
Researchers propose a minimal method for adapting language models to domain-specific tasks without fine-tuning or prompt engineering. The approach uses a single context-independent logit-bias vector added at every decoding step, learned via a black-box method from a KL-regularized reinforcement learning objective. A closed-form inverse-propensity estimator is derived from rollouts, rewards, and token probabilities. The method addresses privacy concerns by avoiding model weight modification. The paper is available on arXiv.
Key facts
- Method uses a single logit-bias vector added at every decoding step
- Black-box learning without model weights modification or gradients
- Derived from KL-regularized reinforcement learning objective
- Closed-form inverse-propensity estimator from rollouts, rewards, and token probabilities
- Aims to improve performance on domain-specific tasks and address privacy concerns
- Paper available on arXiv with ID 2607.22837
Entities
Institutions
- arXiv