An interactive treemap of a large language model's pretraining corpus, using The Pile's rough composition: web text (Pile-CC) is the largest slice, followed by PubMed, books, web crawl backbone, code from GitHub, ArXiv papers, Wikipedia, StackExchange, legal text, patents and more. Hover or click any rectangle to read a one-line description of that data source in a panel below. A Quality and safety filter toggle greys out and removes the low-quality and unsafe slices — boilerplate, near-duplicate pages, and toxic text — visibly shrinking and cleaning the pile to show that what survives filtering is what the model actually learns from. The take-away: a model is mostly whatever it read.
What an LLM eats
The pretraining diet, sized by share. Poke a slice to see what it is.
Hover a slice
Each rectangle is one source the model read. Bigger box = more of the diet.
100%
of the pile kept
13
sources shown
A model is mostly what it read. Clean the diet — drop boilerplate,
near-duplicates and toxic text — and the pile gets smaller, but every byte that
survives is text worth learning from.