Don't Trust the Label: License Laundering in AI Supply Chains
This paper investigates the survival of licenses in the supply chain of AI artifacts, finding that a large percentage of artifacts lack declared licenses and that obligation-bearing licenses have low survival rates.
The study provides new insights into the survival of licenses in the AI supply chain and identifies the prevalence of artifacts without declared licenses.
Before reading this…
Applications
- →Model publishers
- →Rights holders
- →Platform owners
To understand this paper, make sure you know these concepts first:
- Understanding of AI artifacts and their supply chainfind papers →
- Familiarity with software licensingfind papers →
Abstract
More Like ThisAI artifacts move through a multi-platform supply chain, spanning datasets and models on Hugging Face and applications on GitHub. While each artifact carries a license whose obligations should propagate through redistribution, no study has yet measured whether those obligations survive the chain or are stripped and replaced as artifacts move downstream. We trace 232,270 dataset$\rightarrow$model$\rightarrow$application chains and quantify two forms of license laundering: when artifacts with no declared license acquire definitive labels downstream, and when one declared license category replaces another during redistribution. We find that 62.3% of chains pass through at least one artifact with no declared license (concentrated in a small set of foundational datasets), and that every obligation-bearing license category falls below 7% end-to-end survival while the Permissive category reaches 95.1%. Based on these findings, we provide actionable recommendations for practitioners, model publishers, rights holders, and platform owners.