principles.fyi · the brain · concept

compute-optimal scaling (Chinchilla)

Match how big your brain is to how much you read.

grow model size and training data together for a fixed compute budget, roughly 20 training tokens per parameter (e.g. Chinchilla: 70B params, 1.4T tokens)

When you can only spend so much effort training, you have to choose: build a bigger model, or feed it more text? Chinchilla showed the best results come from growing both together, in step. The surprise was that famous earlier models were too big and hadn't read enough — like a giant brain that only skimmed a few books. A right-sized model that reads a lot can beat a huge one that reads too little.

Appears in

Nearby in the brain