Vision encoders learn and store conceptual semantic information like canonical colors independently of visual input, and this knowledge is tightly linked to object recognition—providing a measurable way to understand what abstract concepts models actually learn.
This paper investigates whether vision encoders in vision-language models (VLMs) learn conceptual information beyond what's visible in images. Using canonical color (like 'bananas are yellow') as a test case, researchers found that models can decode object colors from grayscale images, suggesting they learn abstract semantic knowledge.