Study Finds 91.8% of Agent Skills Contain Defects, Hindering Reusability
A recent investigation into 138,133 public SKILL.md files sourced from 20,556 repositories has found that 91.8% of Agent Skills harbor at least one defect, which hampers their reusability. This study, published on arXiv (2608.08453), employs a two-tier defect classification based on established specifications and best practices. The primary issues identified are typical packaging flaws, including inadequate routing metadata, excessive or unhelpful content, and disorganized resources, rather than security vulnerabilities. A deterministic routing stress test involving 20,000 skills shows that those with valid routing metadata are considerably more reusable. The results expose a disparity between the ideal standard for Agent Skills—SKILL.md files designed to facilitate LLM agents in reusing procedures—and the reality that many public skills stem from isolated tasks or conversations. The study offers consistent estimates across both lenient and strict criteria (88.8-94.6%), underscoring the widespread nature of defects. This research is vital for the AI community, as it highlights frequent challenges in skill development and suggests improvements for enhancing reusability, potentially shaping future standards for Agent Skills.
Key facts
- 138,133 public SKILL.md files analyzed
- 20,556 repositories examined
- 91.8% of skills contain at least one defect
- Defect rates stable across thresholds: 88.8-94.6%
- Dominant failures: weak routing metadata, bloated bodies, poor resource organization
- Routing stress test conducted on 20,000 skills
- Study available on arXiv with ID 2608.08453
- Two-tier defect taxonomy used
Entities
Institutions
- arXiv