Open source
Software is open source when anyone can inspect, modify and redistribute the code. The Open Source Initiative extended that to AI in October 2024 with the Open Source AI Definition, which asks for three things under terms that allow anyone to use, study, modify and share the system for any purpose: the parameters, the training and inference code, and detailed information about the training data. Note the third one. The definition asks for a full description of the data and its provenance, not the dataset itself, and plenty of people think that bar is set too low. The point of open source is the process, not just the artifact. When the build is visible, an independent group can test a developer's claims about bias, failure modes and the populations a model was never meant for, instead of taking them on trust. Few released models clear the bar. OLMo from the Allen Institute for AI does. Llama, Gemma and Mistral's open releases are open weights under vendor licences, and calling those open source sets the wrong expectation about what can be verified.
Terms like this come up in real clinical scenarios across the HelloAI courses: bite-sized modules with verifiable certificates. An account takes one minute, no password needed.
Sign in →