osaid-october-2024-training-data-disclosure
IN premise — entries/2026/06/21/wiki-LLaMA-chunk-2.md
Created 2026-06-21T09:50:09+00:00
The Open Source AI Definition (OSAID), published by OSI in October 2024, requires open-source AI to disclose training data details, which Meta does not do for Llama
Summary
The Open Source Initiative's 2024 definition of open-source AI sets a bar that includes sharing details about what data the model was trained on, and Meta's Llama models fall short of that standard because those details are not published. This matters because it means Llama occupies a gray zone: widely called "open source" in practice, but not actually meeting the formal community definition of what that term should mean.
Dependents
These beliefs depend on this one:
- IN open-weight-models-face-unresolved-definitional-tensions — The "open" AI ecosystem faces unresolved tensions: Llama's license restricts large platforms and prohibits competitive training use, the FSF classified it as nonfree software, and the OSAID requires training data disclosure that most "open" models do not provide.
- OUT osaid-could-resolve-open-weight-governance-gap — The Open Source AI Definition (OSAID, October 2024) — requiring training data disclosure as a condition of the "open-source AI" label — could resolve the open-weight governance gap by establishing a clear, enforceable standard that retires the definitional tensions currently preventing coherent policy and enabling principled governance of model distribution.