A batch of image and caption pairs scored against each other, with the diagonal pushed up and everything else pushed down.
Learning from image–text pairs
Collected image–text pairs supply a training signal. These invented scores illustrate the objective; no encoders run or train here.
Collected pairs are positive examples. Other combinations are treated as negatives; they are not guaranteed semantic mismatches. A caption paired with one image may also describe another.