ARTFEED — Contemporary Art Intelligence

Suppressing Evaluation-Awareness Latents in LLMs via Input-Only Optimization

ai-technology · 2026-07-29

A recent preprint on arXiv (2607.25907) presents a technique aimed at reducing evaluation-awareness latents in large language models through input-only optimization, eliminating the need for model access during inference. This method modifies Fluent Dreaming and EPO by incorporating a negated feature term, merging GCG-style token optimization with a self-cross-entropy fluency regularizer. The target latent, a linearly interpretable and controllable evaluation-awareness feature highlighted in previous research, poses risks to safety evaluation validity if models react differently during testing. Tests conducted on Llama-3.2-3B and Llama-3.1-8B effectively suppressed the latent across five target constructs. The suppression is notably strong (z≈-7), and a causally-validated Llama Scope SAE feature can be completely and selectively disabled.

Key facts

  • Method suppresses evaluation-awareness latents via input-only optimization
  • No inference-time model access required
  • Adapts Fluent Dreaming / EPO with negated feature term
  • Uses GCG-style token optimization plus self-cross-entropy fluency regularizer
  • Target latent is linearly readable and steerable
  • Experiments on Llama-3.2-3B and Llama-3.1-8B
  • Five target constructions tested: CAA direction, subspace norm, SAE feature, single MLP neuron, behavioral logit
  • Latent robustly suppressible (z≈-7)

Entities

Institutions

  • arXiv

Sources