An interactive treemap of a large language model's pretraining corpus, using The Pile's rough composition: web text (Pile-CC) is the largest slice, followed by PubMed, books, web crawl backbone, code from GitHub, ArXiv papers, Wikipedia, StackExchange, legal text, patents and more. Hover or click any rectangle to read a one-line description of that data source in a panel below. A Quality and safety filter toggle greys out and removes the low-quality and unsafe slices — boilerplate, near-duplicate pages, and toxic text — visibly shrinking and cleaning the pile to show that what survives filtering is what the model actually learns from. The take-away: a model is mostly whatever it read.

What an LLM eats

The pretraining diet, sized by share. Poke a slice to see what it is.

Hover a slice

Each rectangle is one source the model read. Bigger box = more of the diet.

100%
of the pile kept
13
sources shown

A model is mostly what it read. Clean the diet — drop boilerplate, near-duplicates and toxic text — and the pile gets smaller, but every byte that survives is text worth learning from.