ARTFEED — Contemporary Art Intelligence

Scale-Aware Memorization Diagnostics for Code LLMs

ai-technology · 2026-08-15

A new arXiv preprint (2608.12771) investigates the reliability of memorization detection techniques in large language models for code, revealing that traditional probing methods fail at scale. The study finds that encoder-style probes using perturbations like synonym fuzzing or dead-code insertion, and decoder-style probes relying on log probabilities, degrade severely when applied to larger models, even on benchmarks known to be contaminated. The authors argue that memorization and representation load should be treated as separate phenomena, and they propose using invertible mathematical transforms on numeric problems to isolate these factors. The research highlights the need for scale-aware diagnostics in evaluating code LLMs, as current methods may underestimate memorization in scaled models. The preprint was announced on arXiv with a cross-type classification, and the full text is available at the provided URL.

Key facts

  • The preprint is arXiv:2608.12771v1.
  • The study focuses on large language models for code.
  • Traditional encoder-style probes (synonym fuzzing, dead-code insertion) fail at scale.
  • Decoder-style probes using log probabilities also degrade in performance.
  • The authors propose separating representation load from memorization.
  • Invertible mathematical transforms are applied to numeric problems to isolate factors.
  • The research suggests current memorization diagnostics are not scale-aware.
  • The paper is available on arXiv.

Entities

Institutions

  • arXiv

Sources