How effectively a model can jointly learn and model both image and text tokens together during training.