HelloAI glossary

Open source

Software is open source when anyone can inspect, modify and redistribute the code. The Open Source Initiative extended that to AI in October 2024 with the Open Source AI Definition, which asks for three things under terms that allow anyone to use, study, modify and share the system for any purpose: the parameters, the training and inference code, and detailed information about the training data. Note the third one. The definition asks for a full description of the data and its provenance, not the dataset itself, and plenty of people think that bar is set too low. The point of open source is the process, not just the artifact. When the build is visible, an independent group can test a developer's claims about bias, failure modes and the populations a model was never meant for, instead of taking them on trust. Few released models clear the bar. OLMo from the Allen Institute for AI does. Llama, Gemma and Mistral's open releases are open weights under vendor licences, and calling those open source sets the wrong expectation about what can be verified.

In the clinic

A vendor pitches a documentation assistant as built on an open source model, and your AI governance committee takes that to mean outside researchers have already checked it for bias. Open the release itself. If it includes the weights, training code and a description of where the training data came from, independent groups can test the developer's claims and you can look for their findings. If it ships only weights under the developer's own license, every claim about the training data rests on the developer's word. Either way, the vendor adds its own tuning and any outside findings come from other patients, so budget for bias testing of the assistant on yours.

Go beyond the definition

Terms like this come up in real clinical scenarios across the HelloAI courses: bite-sized modules with verifiable certificates. An account takes one minute, no password needed.

Sign in →
See all terms →Still unclear? Ask the team →