principles.fyi · Topic

Post-training

Turning a model that just predicts the next word into one that's actually helpful, honest, and harmless — instruction tuning, reward models, RLHF, DPO, and thinking longer at answer-time.