ARTFEED — Contemporary Art Intelligence

Nanbeige4.2-3B: A Compact 3B Model for Agentic Tasks

ai-technology · 2026-07-27

A team of researchers has unveiled Nanbeige4.2-3B, a compact general agentic model featuring 3 billion non-embedding parameters. This model was pretrained from scratch using a Looped Transformer architecture on 28 trillion tokens, allowing for increased capacity without additional parameters. It excels in various tasks, including code-agent, office-agent, and complex tool usage, while also showing strong reasoning capabilities in mathematics, coding, and science. To enhance supervised fine-tuning (SFT) data and trajectory construction, the researchers broadened the diversity of executable environments and task assets through real-world applications and large-scale synthesis. Their reinforcement learning (RL) pipeline incorporates mixed-mode RLHF, length-controlled reasoning RL, and agentic RL, leading to improved performance and reduced failure rates. Comprehensive evaluations confirm the model's effectiveness.

Key facts

  • Nanbeige4.2-3B has 3B non-embedding parameters.
  • Pretrained from scratch on 28T tokens.
  • Uses Looped Transformer to reuse layer stack.
  • Strong performance on code-agent, office-agent, and tool-use tasks.
  • Competitive reasoning in math, coding, and science.
  • SFT data expanded via real-world deployment and large-scale synthesis.
  • RL pipeline includes mixed-mode RLHF, length-controlled reasoning RL, and agentic RL.
  • Agentic RL uses outcome and process rewards for long-horizon training.

Entities

Sources