ARTFEED — Contemporary Art Intelligence

SkillTV-Bench: New Benchmark for Skill-Aware Agentic Execution Verification

ai-technology · 2026-08-07

A new benchmark, SkillTV-Bench, has been introduced to evaluate the performance of judges in verifying skill-augmented agentic executions. The benchmark comprises 681 cases of real agent trajectories from 50 tasks across eleven domains. It is designed to assess both LLM-as-a-Judge and Agent-as-a-Judge methods in skill-aware trajectory verification. The accompanying SkillTV-Evolve framework externalizes verification knowledge as a reusable JudgeSkill, guiding agent judges to plan effective verification strategies. This development addresses the shift in evaluation from final-response scoring to verification of complete executions, incorporating procedural knowledge encoded in task-time skills. The benchmark is detailed in a paper on arXiv (arXiv:2608.05573), announced as a new submission.

Key facts

  • SkillTV-Bench includes 681 cases of real agent trajectories.
  • The benchmark covers 50 tasks across eleven domains.
  • It evaluates LLM-as-a-Judge and Agent-as-a-Judge methods.
  • SkillTV-Evolve externalizes verification knowledge as a reusable JudgeSkill.
  • The paper is available on arXiv with ID 2608.05573.
  • The benchmark focuses on skill-augmented agentic execution verification.
  • It shifts evaluation from final-response scoring to verification of complete executions.
  • The benchmark incorporates procedural knowledge from task-time skills.

Entities

Institutions

  • arXiv

Sources