Palmyra x6 Technical Report Released on arXiv: Agentic Tool-Use Model Achieves Top Benchmark Scores
A new technical report on arXiv introduces Palmyra x6, a comprehensive language model aimed at enterprise use. This model was created by enhancing a Mixture-of-Experts framework with Anchored Supervised Fine-Tuning, trained on a small dataset of 626 verified synthetic tool-use trajectories. The training involved just one epoch, a low learning rate, and a KL divergence anchor, using a combination of Muon and Adam optimizers. Palmyra x6 outperformed the previous default model for Writer Agent tasks, achieving a score of 0.785 on BFCL Core and the highest average across six benchmarks. It also excelled in bias and safety evaluations. The report recommends a careful training approach, and the code and data are available, categorized under Computer Science, Computation and Language.
Key facts
- Palmyra x6 is optimized for enterprise-oriented agentic tasks.
- The model was built by post-training a Mixture-of-Experts base model.
- Anchored Supervised Fine-Tuning was performed on 626 verified synthetic tool-use trajectories.
- Training used a Muon + Adam hybrid optimizer with a single epoch and low learning rate.
- A KL anchor to the frozen base model was maintained throughout training.
- Palmyra x6 shows substantial gains over the previous default model for Writer Agent.
- It scored the highest on BFCL Core at 0.785.
- It posted the highest six-benchmark mean among comparators.
- The model was competitive or leading in bias and safety evaluations.
- The report is available on arXiv and indexed by Semantic Scholar.
Entities
Institutions
- arXiv
- Semantic Scholar
- arXivLabs