Two scripted responses to one question: a wrong direct answer and a correct worked answer. No model is running here. The example illustrates how an intermediate result can support a later step.
pretraining
predict the next word
post-training
follow + prefer
answer-time
spend compute thinking
prompt
23 apples. They used 20 to make lunch and bought 6 more. How many apples do they have?
output
The worked example carries 23 − 20 = 3 forward into 3 + 6 = 9. A real model can answer directly and correctly, or write incorrect reasoning. These fixed examples illustrate the format; they are not an accuracy comparison.