Nanbeige4.2-3B: A Compact 3B Model for Agentic Tasks
A team of researchers has unveiled Nanbeige4.2-3B, a compact general agentic model featuring 3 billion non-embedding parameters. This model was pretrained from scratch using a Looped Transformer architecture on 28 trillion tokens, allowing for increased capacity without additional parameters. It excels in various tasks, including code-agent, office-agent, and complex tool usage, while also showing strong reasoning capabilities in mathematics, coding, and science. To enhance supervised fine-tuning (SFT) data and trajectory construction, the researchers broadened the diversity of executable environments and task assets through real-world applications and large-scale synthesis. Their reinforcement learning (RL) pipeline incorporates mixed-mode RLHF, length-controlled reasoning RL, and agentic RL, leading to improved performance and reduced failure rates. Comprehensive evaluations confirm the model's effectiveness.
Key facts
- Nanbeige4.2-3B has 3B non-embedding parameters.
- Pretrained from scratch on 28T tokens.
- Uses Looped Transformer to reuse layer stack.
- Strong performance on code-agent, office-agent, and tool-use tasks.
- Competitive reasoning in math, coding, and science.
- SFT data expanded via real-world deployment and large-scale synthesis.
- RL pipeline includes mixed-mode RLHF, length-controlled reasoning RL, and agentic RL.
- Agentic RL uses outcome and process rewards for long-horizon training.
Entities
—