Skip to content
Back to chapters

How LLMs are Trained

Discover how neural networks learn and improve through human feedback

Training Data & Scale

LLMs are trained on billions of words from the internet. Explore the data sources.

GPT-3 was trained on approximately 333 billion words from these sources. Click a source to learn more.

Common Crawl410B words
Quality: Low · Weight in training: 60%
WebText219B words
Quality: High · Weight in training: 22%
Books112B words
Quality: High · Weight in training: 8%
Books255B words
Quality: High · Weight in training: 8%
Wikipedia3B words
Quality: Very High · Weight in training: 3%
Total: ~333 Billion words · ~500 Billion tokens · ~680 GB of text

RLHF — Learning from Human Feedback

Rate model outputs and see how preferences shape behavior.

RLHF (Reinforcement Learning from Human Feedback) works by showing humans two model outputs and asking which is better. Your preferences help the model learn what "good" looks like. Try it yourself:

Question:

How do I make my code run faster?

Total preferences collected: 0