GPT-2 with ReLU² Instead of GELU: 0.051 Lower Loss After 1B Tokens, From the Same Starting Weights
Three GPT-2 (124M) runs with ReLU² in the MLP ended at 3.623 mean validation loss after 1B FineWeb-Edu tokens, against 3.674 for GPT-2's GELU, at GPT-2's learning rate. Each pair started from identical weights, so the activation is the only difference. The gap clears the series' pre-registered bar (0.010) five times over and was below it at no checkpoint.

GPT-2 with ReLU² Instead of GELU: 0.051 Lower Loss After 1B Tokens, From the Same Starting Weights
We replaced the activation inside GPT-2's MLP blocks, the tanh-approximated GELU, with ReLU², which zeroes negative inputs and squares the positive ones. Nothing else changed: same width, same parameter count, same billion tokens in the same order, same learning rate. Three runs of each. The ReLU² runs ended at a mean validation loss of 3.623; the GELU runs at 3.674.
The gap is 0.051 nats. This series sets, before each set of runs, the smallest gap it will call real. For this comparison that bar was 0.010, so ReLU² clears it about five times over.
This comparison is also the cleanest in the series so far. An activation function has no parameters, so with the same seed the two versions start from exactly the same weights; the logs confirm it. In each pair, the only thing that differs from the first step to the last is what happens to the numbers between the MLP's two matrices.
The short answer
- ReLU² ended 0.051 lower than GELU (3.623 against 3.674, three runs each). The bar was 0.010.
- Every ReLU² run was below every GELU run at all eight checkpoints, and each seed's ReLU² run ended 0.048 to 0.057 below the GELU run that started from the same weights.
- The gap was largest early and was still closing at the end: 0.144 at 0.26B tokens, 0.074 at 0.52B, 0.051 at 1.0B. These runs do not say where it would settle.
- Speed: no measured difference. Two of the three ReLU² runs trained 2.2% faster than the GELU runs; the third was slowed for part of its run by other jobs sharing the GPUs. Runs of the same setup have differed by up to 3% in this series, so we do not count 2.2% as a speedup.
What changed
The model is the same GPT-2 small as in the earlier posts: 12 layers, width 768, 12 heads, a 1,024-token context, learned position embeddings, LayerNorm, tied input and output embeddings, nanoGPT code. Each MLP block widens to 3,072, applies the activation, and projects back to 768.
| GELU (GPT-2) | ReLU² | |
|---|---|---|
| Activation | x · Φ(x), tanh approximation | max(0, x)² |
| Negative inputs | small negative outputs, down to about −0.17 | exactly 0 |
| Parameters | 124,475,904 | 124,475,904 |
| Initial weights for seeds 0, 1, 2 | identical to GELU's (same hashes in the logs) |
Everything else matches the earlier posts: FineWeb-Edu shards in the same order, 1,907 steps of 524,288 tokens (1.0B tokens), AdamW with peak learning rate 6e-4, 100 warmup steps and cosine decay to 10%, bfloat16, two A100 80GB cards. The GELU side is the three seed runs from the first measurement post, reused as planned.
Before the runs we checked on the CPU that the two start from the same loss. On the same 4,096 tokens as in the earlier posts, the initial losses were 10.940 (GELU) and 10.950 (ReLU²) for seed 0, and 11.031 and 11.035 for seed 1.
The rule, fixed before the runs
We wrote down this comparison's question, runs and rule on 7 October at 13:16, the day before the first ReLU² run started (8 October, 08:56). The rule is the same as in the previous posts: call the two different only if their mean final losses differ by more than 2 × σ × √(1/3 + 1/3), where σ pools the seed-to-seed variance of every configuration this series has compared so far (GELU, RoPE, RMSNorm, RMSNorm with QK norm, ReLU²). The script that applies the rule also refuses to run if a ReLU² run did not start from the same weights as its GELU partner.
| Final validation loss, seeds 0 / 1 / 2 | Mean | Standard deviation | |
|---|---|---|---|
| GELU | 3.6681, 3.6858, 3.6690 | 3.6743 | 0.0100 |
| ReLU² | 3.6192, 3.6289, 3.6208 | 3.6230 | 0.0052 |
Pooled σ = 0.00629 (10 degrees of freedom), so the bar is 2 × 0.00629 × √(2/3) = 0.0103. The difference is −0.0513.
The three pairs, seed by seed: 3.6681 → 3.6192 (−0.049), 3.6858 → 3.6289 (−0.057), 3.6690 → 3.6208 (−0.048). The seed that was worst with GELU was also worst with ReLU², by a similar margin.
Over the whole run

| Tokens seen | GELU | ReLU² | Gap |
|---|---|---|---|
| 0.13B | 5.5714 | 5.5164 | −0.0550 |
| 0.26B | 4.7146 | 4.5706 | −0.1440 |
| 0.39B | 4.1726 | 4.0744 | −0.0982 |
| 0.52B | 3.9468 | 3.8726 | −0.0742 |
| 0.66B | 3.8211 | 3.7582 | −0.0629 |
| 0.79B | 3.7389 | 3.6829 | −0.0560 |
| 0.92B | 3.6918 | 3.6396 | −0.0522 |
| 1.00B | 3.6743 | 3.6230 | −0.0513 |
(Means of three runs. At every checkpoint the highest ReLU² run was below the lowest GELU run.)
The gap opens fast, peaks at 0.26B tokens and then shrinks, quickly at first and slowly at the end: by 0.070 between 0.26B and 0.52B tokens, by 0.018 between 0.52B and 0.79B, by 0.005 between 0.79B and 1.0B, and by 0.0009 over the last 82M tokens. The ReLU² runs reached the GELU runs' final loss (3.674) somewhere between 0.79B and 0.92B tokens.
That shape resembles what RoPE did in post 2: a large early advantage that narrows as training goes on. We did not test why either one narrows, and with one budget we cannot say whether the ReLU² gap levels off near 0.05 or keeps closing.
Where ReLU² comes from
So et al. (2021) found ReLU² by searching over the building blocks of a Transformer for cheaper training. Their abstract attributes most of the improvement of the resulting architecture, Primer, to two changes: squaring the ReLU activations and adding a depthwise convolution after the query, key and value projections. In the modded-nanogpt speedrun, ReLU² arrived in record 5 together with QK norm, zero-initialized projections and padded embeddings, so its share of that record's gain was not measured on its own. This post measures it alone, at GPT-2's learning rate.
The series so far, each change measured against the same three GPT-2 runs except where noted:
| Change | Gap in final loss | Verdict |
|---|---|---|
| RoPE instead of learned positions (post 2) | −0.122 | lower |
| Parameter-free RMSNorm instead of LayerNorm (post 3) | +0.038 | higher |
| QK norm on top of RMSNorm (against RMSNorm alone) | −0.012 | lower, by 0.001 over the bar |
| ReLU² instead of GELU (this post) | −0.051 | lower |
The gaps are not additive. Each was measured on its own, starting from GPT-2.
Speed
| GELU | ReLU² | |
|---|---|---|
| Training throughput per run | 361,173 / 361,211 / 361,185 tokens/s | 369,119 / 369,225 / 282,952 tokens/s |
| Wall time per run, including evaluation | 48.7 min each | 47.8 / 47.7 / 62.2 min |
The third ReLU² run's average is low because for part of the run its throughput fell as far as 121,714 tokens per second; its median logged step ran at 368,867, like the other two. The slow stretch ran from about 10:35 to 11:00 on 8 October. From 10:49 to 11:01 the GPUs were shared with another of our training jobs, a follow-up run for a different post, and the throughput recovered as soon as that job ended. What was slowing the run between 10:35 and 10:49 we did not record. Taking the two undisturbed runs, ReLU² trained 2.2% faster than GELU. But the GELU runs were measured on a different day, and in post 3 three runs of one configuration spread across 3%, so we do not call this a measured speedup. Throughput was never part of the verdict.
What this does not show
- One learning rate. Both sides used GPT-2's 6e-4. Post 5 is planned to change the optimizer and sweep the learning rate.
- One budget, one size. GPT-2 small at 1B tokens. The gap was still shrinking at the end.
- Why it helps. We measured loss and throughput only. How many MLP activations are exactly zero with ReLU², and whether that matters, was left out of the plan.
- One MLP shape. Width 3,072, no gated variants such as SwiGLU.
- Throughput belongs to this implementation and this machine.
Next
Next in the plan is the optimizer: Muon against AdamW, with the learning rate swept for each. Post 5's comparison runs have not started yet.
The plan, the scripts and the raw logs of all six runs are in the reproduction package below, no login needed. Its README has the commands that recompute every number in this post from those logs, without a GPU.
Files for this post
Reproduction package
The harness, raw logs and result CSVs behind this post. No account needed; the README has a command that recomputes the tables without a GPU.
gpt2-relu2-repro.zip · 356 KB