A batch of image and caption pairs scored against each other, with the diagonal pushed up and everything else pushed down.

Learning from image–text pairs

Collected image–text pairs supply a training signal. These invented scores illustrate the objective; no encoders run or train here.

Collected pairs are positive examples. Other combinations are treated as negatives; they are not guaranteed semantic mismatches. A caption paired with one image may also describe another.