License Drift in Open-Source AI: 35.5% of Models Re-licensed Permissively
A recent study published on arXiv highlights widespread license violations within the open-source AI community. The research team performed a comprehensive audit encompassing 364,000 datasets and 1.6 million models hosted on Hugging Face, along with 140,000 projects from GitHub. Their findings revealed that 35.5% of transitions from models to applications bypass restrictive licensing by adopting more permissive terms, which could lead to legal risks for organizations. Additionally, the researchers developed a flexible rule engine that incorporates nearly 200 SPDX and model-specific clauses, successfully identifying 86.4% of licensing conflicts.
Key facts
- First end-to-end audit of licenses for datasets and models on Hugging Face and their downstream integration into GitHub projects
- Covered 364,000 datasets, 1.6 million models, and 140,000 GitHub projects
- 35.5% of model-to-application transitions eliminate restrictive license clauses by relicensing under permissive terms
- Prototyped an extensible rule engine encoding almost 200 SPDX and model-specific clauses
- Rule engine can solve 86.4% of license conflicts
- Hidden license conflicts pose serious legal and ethical risks
- Study published on arXiv with ID 2509.09873
- Research highlights systemic non-compliance in open-source AI
Entities
Institutions
- Hugging Face
- GitHub
- arXiv