HelloAILearn
GlossarySign in →
HelloAI glossary

Vision-language model

A model trained on images and text together, so it can describe an image, answer a question about it, or point to the part of it a sentence refers to. This is the class of model behind generating a report from a chest X-ray, or answering a question about a scan in words rather than returning a label. Training usually pairs images with the reports already written about them, which is why radiology moved first: the paired data was sitting there. The known weakness is that fluency and grounding come apart, so the model can produce a well-written report describing a finding nobody can locate in the image. Any clinical use needs a way to check the words against the pixels.

In the clinic

A draft report arrives already written, in the house style of the archive it learned from, including a sentence about a finding that is not on the scan. It is fluent, it is formatted correctly, and it is wrong in a way that is harder to catch than a blank field would be. Any deployment needs the words tied back to the pixels, either through a tool that points at what it is describing or a reading step that assumes nothing on the page is confirmed. Fluency is the part that got optimized. Accuracy was not.

Go beyond the definition

Terms like this come up in real clinical scenarios across the HelloAI courses: bite-sized modules with verifiable certificates. An account takes one minute, no password needed.

Sign in →
See all terms →
Learning Again