RubricForge: Evolving Reward-Free Judging Rubrics to Reduce Over-Crediting in Agent Evaluation
A new arXiv paper (2608.13564) introduces RubricForge, a method for automatically generating judging rubrics for evaluating language-model agents. The approach addresses the problem of over-crediting in agent evaluation, where existing judges (such as G-Eval or fine-tuned models) tend to credit fluent but unsuccessful trajectories as successes. RubricForge evolves a judge rubric using reflective evolution against a small set of ground-truth-labeled trajectories, maximizing agreement with the environment reward. The optimized rubric is frozen and applied to held-out trajectories in a single model call without environment access. The resulting artifact is human-readable text, ensuring transparency in verdicts. This work is relevant to the field of AI evaluation, offering a more reliable and cost-effective alternative to hand-written rubrics or fine-tuned judges.
Key facts
- Paper arXiv:2608.13564 introduces RubricForge.
- RubricForge evolves a judge rubric using reflective evolution against labeled trajectories.
- The method maximizes agreement with the environment reward.
- The optimized rubric is frozen and applied to held-out trajectories in one model call.
- No environment access is required during application.
- The artifact is human-readable text.
- Existing judges like G-Eval hand-write scoring rubrics or fine-tune judge weights.
- The approach aims to reduce over-crediting of fluent but unsuccessful trajectories.
Entities
Institutions
- arXiv