ARTFEED — Contemporary Art Intelligence

Final pretraining data shapes model behavior after SFT

ai-technology · 2026-07-29

There’s a new study on arXiv (2607.25063) that looks into how the last bit of pretraining data might affect language model checkpoints even after they've gone through supervised fine-tuning (SFT). The researchers took one partially pretrained checkpoint and made six different versions, each using a unique dataset of 500 million tokens. These datasets included things like generic web text and safety text. After they all went through the same SFT, they found that even though the checkpoints scored similarly on benchmarks, they responded differently to further alignment tasks, like preference optimization. This suggests that the last data the models see before tuning can influence their behavior in ways that typical assessments might miss, raising questions about how interchangeable these checkpoints really are.

Key facts

  • arXiv paper 2607.25063
  • Six branches from one partially pretrained checkpoint
  • Each branch trained on 500 million tokens from a single data source
  • Data sources: generic web text, filtered web text, normative discourse, safety text, mathematical text, synthetic educational text
  • SFT and post-training were identical across branches
  • Checkpoints performed similarly on benchmarks after SFT
  • Differences emerged in response to further alignment (preference optimization)
  • Final pretraining window influences post-training beyond SFT

Entities

Institutions

  • arXiv

Sources