ARTFEED — Contemporary Art Intelligence

DMAPO: Data-Centric Multi-Evaluator Agreement for Preference Optimization

ai-technology · 2026-07-29

There's this new method called DMAPO, which stands for Data-centric Multi-evaluator Agreement for Preference Optimization. It's designed to improve how language models align by using a small, reliable set of on-policy answers. Instead of changing the training goals while keeping the data the same, DMAPO generates responses from the target policy and evaluates them based on usefulness, factual correctness, and clarity with the help of specific evaluators. They then refine the responses by keeping only the most agreed-upon ones. From a pool of 54,236 candidates from Mistral-7B, only 1,871 (3.45%) made the cut. The model KTO, trained on these, scored 7.50 on MT-Bench and outperformed text-davinci-003 with a 95.5% win rate.

Key facts

  • DMAPO stands for Data-centric Multi-evaluator Agreement for Preference Optimization.
  • Method generates candidate responses from the target policy.
  • Evaluates helpfulness, factuality, and conciseness with rubric-specialized evaluators.
  • Applies a process-critic correction.
  • Retains only high-consensus desirable or undesirable examples.
  • Accepts 1,871 of 54,236 Mistral-7B candidates (3.45%).
  • KTO trained on this set reaches 7.50 on MT-Bench.
  • Achieves 95.5% length-controlled win rate against text-davinci-003.
  • Achieves 57.3% IFEval prompt accuracy.
  • GPT-4o yields net win rate of 23.3 points on 129 held-out prompts and 24.0 points on 200 out-of-distribution LMSY.

Entities

Sources