A log-log plot of model loss versus model size. The relationship is a straight, downward-sloping line — a power law. A slider moves a marker from a small model to a huge one; the loss readout drops steadily along the line. Each time the model grows 10x, the loss falls by a fixed, predictable amount.
Why bigger is (predictably) better
On log–log axes, loss vs. size is a straight line. Scale up → loss slides down.
Model size
100 M
Loss
2.40
10× the model → a steady, predictable drop in loss.