Activation Oracles Develop Concept-Specific Blind Spots
A recent study published on arXiv (2607.23379) indicates that Activation Oracles (AOs)—language models designed to interpret another model's internal activations—can exhibit blind spots specific to certain concepts. In a regulated Taboo Word Guessing scenario, the subject models were adjusted to utilize a concealed concept internally while refraining from direct revelation. Surprisingly, the fine-tuned AOs acted as anti-readers, consistently struggling to retrieve the concept that had been a constant during their training. This finding suggests that AOs do not function as impartial interpreters; rather, their performance is influenced by the training data and objectives they are exposed to.
Key facts
- Activation Oracles are language models trained to answer questions about another model's internal activations.
- The study uses a controlled Taboo Word Guessing setting.
- Subject models were fine-tuned to internally use a hidden concept while avoiding direct disclosure.
- Fine-tuned AOs can become concept-specific anti-readers.
- AOs selectively fail to recover the concept present during their own training.
- AOs are shaped by training data and objectives, not neutral readouts.
- The paper is on arXiv with ID 2607.23379.
- The research highlights limitations in using AOs for interpretability.
Entities
Institutions
- arXiv