This scripted illustration moves probability toward an observed target. It computes the loss; it does not train weights or calculate gradients.
“The cat sat ___”
Each click transfers 32% of the remaining probability mass to “on”. This chosen rule lowers this example’s loss; a real training update need not improve every example.