SkillTV-Bench: New Benchmark for Skill-Aware Agentic Execution Verification
A new benchmark, SkillTV-Bench, has been introduced to evaluate the performance of judges in verifying skill-augmented agentic executions. The benchmark comprises 681 cases of real agent trajectories from 50 tasks across eleven domains. It is designed to assess both LLM-as-a-Judge and Agent-as-a-Judge methods in skill-aware trajectory verification. The accompanying SkillTV-Evolve framework externalizes verification knowledge as a reusable JudgeSkill, guiding agent judges to plan effective verification strategies. This development addresses the shift in evaluation from final-response scoring to verification of complete executions, incorporating procedural knowledge encoded in task-time skills. The benchmark is detailed in a paper on arXiv (arXiv:2608.05573), announced as a new submission.
Key facts
- SkillTV-Bench includes 681 cases of real agent trajectories.
- The benchmark covers 50 tasks across eleven domains.
- It evaluates LLM-as-a-Judge and Agent-as-a-Judge methods.
- SkillTV-Evolve externalizes verification knowledge as a reusable JudgeSkill.
- The paper is available on arXiv with ID 2608.05573.
- The benchmark focuses on skill-augmented agentic execution verification.
- It shifts evaluation from final-response scoring to verification of complete executions.
- The benchmark incorporates procedural knowledge from task-time skills.
Entities
Institutions
- arXiv