Models & Algorithms••KR

Karpathy's “Let's Reproduce GPT-2” in One Read: The Four Parts, Their Numbers, and What Changed Since

A condensed walk through the 4-hour video: building GPT-2 (124M), taking a step from 1,000 ms to 90 ms on one A100, the GPT-3 training settings, and a 10B-token run that beat OpenAI's checkpoint on HellaSwag. Plus one check we ran: nanoGPT's exact GELU moves logits by up to 5 against GPT-2's.

Karpathy's “Let's Reproduce GPT-2” in One Read: The Four Parts, Their Numbers, and What Changed Since

Karpathy's "Let's Reproduce GPT-2" in One Read: The Four Parts, Their Numbers, and What Changed Since

Andrej Karpathy's Let's reproduce GPT-2 (124M) went up on 9 June 2024. It runs 4 hours and 1 minute and had about 1.2 million views when I checked on 6 October 2026. It starts from an empty Python file and ends with a model that beats the smallest GPT-2 on the benchmark it uses. The code is in karpathy/build-nanogpt, one commit per step of the video (44 commits in all, counting later fixes).

This post condenses it. It follows the video's four sections (the data, evaluation and launch steps at the end of Section 3 are folded into the results section here), keeps the numbers Karpathy shows on screen, and marks where each one comes from: the video (with a timestamp), the repository, or a run of mine. Watch the video if you want to build it with him. Read this if you want the shape of it, or a map before you start.

The short version

  • Section 1 builds GPT-2 and checks it against OpenAI's weights before training anything. Load the released checkpoint into your own module, sample from it, and only then switch to random weights.
  • Section 2 makes one training step about 11 times faster on one A100: 1,000 ms to 90 ms for a batch of 16 × 1,024 tokens, through TF32, bfloat16, torch.compile, Flash Attention and a padded vocabulary.
  • Section 3 copies the training settings from the GPT-3 paper: AdamW, gradient clipping, warmup and cosine decay, weight decay on matrices only, a batch of about 0.5M tokens, then gradient accumulation and eight GPUs.
  • Section 4 trains on 10B tokens of FineWeb-Edu in about two hours on 8 A100s. The result beats OpenAI's GPT-2 (124M) on HellaSwag, and a 40B-token run gets to 33.24%, close to GPT-3 (124M).

Section 1: build GPT-2, then prove it is GPT-2 (00:13:47)

Karpathy writes the model as a plain PyTorch module whose parameter names match Hugging Face's GPT-2, so that OpenAI's released weights load straight into it (28:08). He then writes the forward pass, a sampling loop, and generates text from the real checkpoint (37:02). Only when that works does he switch to random weights and start training, first by overfitting one batch of Tiny Shakespeare (56:42), then with a small data loader.

Two details from this section are worth noting, because the speedrun described at the end of this post later replaced both (with an untied output layer and zero-initialized projections):

  • The token embedding and the output layer share one matrix (1:06:14). For GPT-2 small that matrix is 50,257 × 768, about 38.6M parameters, roughly 31% of the 124M total.
  • Initialization follows GPT-2's code: normal with standard deviation 0.02, and the projections that write into the residual stream scaled down by 1/√(2 × number of layers) so the stream does not grow with depth (1:13:47).

One check we ran: the GELU in nanoGPT is not GPT-2's

I started this series from karpathy/nanoGPT (commit 3adf61e, MIT licence), the older repository that, by the video's own description, the build ends up "about 90% similar" to. Before training anything I did what Section 1 does: loaded OpenAI's 124M weights into it and compared logits with Hugging Face transformers 4.56.2 on the same 8 × 1,024 tokens.

They did not match. In float32 on the CPU, the largest logit difference was 5.08. The cause is one line: nanoGPT uses the exact GELU, while OpenAI's GPT-2, Hugging Face's gelu_new and build-nanogpt's own nn.GELU(approximate='tanh') use the tanh approximation. With the tanh form the largest difference dropped to 0.0013 and the loss on those tokens was 3.160816 in both.

How much does it matter? With the same GPT-2 weights, switching the GELU moved the loss on those tokens by only 0.00007 nats (3.160886 against 3.160816). That is all I measured: I did not train a model with the exact GELU, so whether training with it ends up anywhere different is not tested here. But a check that is meant to say "this is GPT-2" fails, and it fails loudly at the level of single logits. If you load GPT-2 weights into nanoGPT, switch the GELU first.

With matching weights, OpenAI's GPT-2 (124M) scores a loss of 3.297 on the held-out FineWeb-Edu tokens I use for the rest of this series (20.9M tokens). That is a reference line, not a target: OpenAI trained on different data (WebText).

Section 2: from 1,000 ms to 90 ms per step (01:22:18)

All timings in this section are one A100 SXM 80GB and a batch of 16 sequences of 1,024 tokens. The milestones are in the video's chapter titles.

ChangeStep timeWhat it does
float32 startabout 1,000 msevery number in 32 bits
TF32 matmuls333 mstensor cores run fp32 matmuls with a shorter mantissa; one line, set_float32_matmul_precision('high')
bfloat16 autocast300 msactivations in 16 bits; same exponent range as fp32, so no gradient scaler
torch.compile130 msremoves Python overhead and fuses elementwise kernels so tensors make fewer trips to memory
Flash Attention96 msnever writes the T × T attention matrix to memory; more FLOPs, much less memory traffic
vocabulary 50,257 → 50,30493 ms50,304 is 128 × 393; kernels handle "nice" sizes in full tiles
fused AdamW and the Section 3 settings90 msone kernel for the optimizer update

The lesson Karpathy draws at each row is the same: the GPU is mostly waiting on memory, not on arithmetic. TF32 promises 8 times the fp32 FLOPs in the datasheet but gave about 3 times here (1:38:24), because the step was not limited by FLOPs. Every later row is a way of moving fewer bytes.

90 ms for 16,384 tokens is about 182,000 tokens per second on one GPU. As a rough cross-check, my harness for this series (the same model with a micro-batch of 64, on two A100 PCIe cards) measured about 361,000 tokens per second in its first training run (the mean of the step times logged from step 10 to 1,580), or about 180,000 per card. The conditions are not the same (different A100 variant, batch and code), so treat it as the same ballpark and no more.

Section 3: the training settings, from the GPT-3 paper (02:14:55)

GPT-2's paper says little about how it was trained, so Karpathy takes the settings from the GPT-3 paper, whose smallest model is the same size:

  • AdamW with β = (0.9, 0.95) and ε = 1e-8; gradients clipped to a norm of 1.0
  • learning rate warmed up linearly over 715 steps (375M tokens, as in the GPT-3 paper), then cosine decay to 10% of the peak; peak 6e-4
  • weight decay 0.1, applied only to 2-D parameters (matmul weights and embeddings), not to biases and LayerNorm gains
  • a batch of 524,288 tokens (2^19, about 0.5M as in the paper). That does not fit on one GPU, so it comes from gradient accumulation (2:34:09): run several small batches, add up their gradients, step once. Each micro-batch loss has to be divided by the number of micro-batches; he shows with a small example that without it the gradients come out too large (2:42:43)
  • DistributedDataParallel across 8 GPUs (2:46:52): each GPU takes different micro-batches, and gradients are averaged before the step. With 64 × 1,024 tokens per GPU and 8 GPUs, no accumulation is needed

He also says something the rest of this series picks up. Near the end (3:51:35) he notes that people had found the peak learning rate can be about three times higher, and that the GPT-3 settings may be "extremely conservative".

Section 4: the run and the results (03:43:05)

The setup for the run comes at the end of Section 3, from 3:10:21; I put it here with the results. The data is FineWeb-Edu, a filtered educational subset of Common Crawl, in its 10B-token sample: 100 shards of 100M tokens. On 8 A100s the run processed about 1.5 million tokens per second at about 330 ms per step, so 19,073 steps (one pass over 10B tokens) were planned at about 1.7 hours (3:21:39 to 3:22:13). Later he calls it a run of "roughly two hours" (3:48:48); the run you see has torch.compile turned off, because at the time it broke his evaluation and sampling code (3:38:55).

The results the next morning:

  • Validation loss on FineWeb-Edu went below the OpenAI GPT-2 (124M) checkpoint evaluated on the same data. Karpathy calls this "not an exactly fair comparison", because GPT-2 was trained on a different distribution.
  • HellaSwag, the comparable benchmark, climbed from 25% (chance) past OpenAI's GPT-2 (124M), which his evaluation script scores at 29.5% (3:37:16). He points out that GPT-2 was trained on about 100B tokens against his 10B.
  • A 40B-token run (four passes over the same data, about 8 hours) reached 33.24% on HellaSwag, just short of GPT-3 (124M), which was trained on 300B tokens (3:52:42).

He lists the caveats himself (3:45:30 to 3:47:10): FineWeb-Edu is English-only and light on code and math, so capacity is not spent elsewhere; HellaSwag is old enough that parts of it may sit in web data; and the training loss shows a periodic wobble, which he suspects comes from reading the shards in a fixed order without shuffling.

What changed after the video

nanoGPT is retired. Its README has said since November 2025 that it is "now very old and deprecated" and points to nanochat, Karpathy's newer project that goes from tokenizer to a chat model. The video still teaches the fundamentals well, but for new work the code to start from is nanochat. This series starts from nanoGPT anyway, on purpose: the point is a faithful GPT-2 baseline to measure later changes against, and nanoGPT, with the one GELU fix above, is exactly that.

The same target became a speedrun. modded-nanogpt took the loss Karpathy's llm.c GPT-2 run reached (3.28 on FineWeb validation, 45 minutes on 8 H100s) and turned it into a record table. The latest record in its README, from 30 August 2026, is 0.665 minutes. The README also says the target now takes under 330M training tokens, against 10B for the llm.c run.

The early records in that table read like a list of what changed in language models after GPT-2: rotary position embeddings with a tuned learning rate, the Muon optimizer, ReLU², zero-initialized projections and QK-norm, an untied output layer, logit softcapping. Each record is measured in wall-clock time on H100s, and many of them change several things at once. So the table tells you the order things were added. It does not tell you how much each change is worth on its own.

That is the question this series takes up, one change at a time, on two A100s, with the same GPT-2 starting point as above. Before comparing anything, the next post measures how much the final loss moves when nothing changes but the random seed. That number decides which differences in the later posts count as real.

If you want to know when that measurement is published, the newsletter form below is the way to hear about it.

Stay Updated

Follow us for the latest posts and tutorials

Subscribe to Newsletter

Related Posts