Models & Algorithms•SOTAAZ Lab••KR

GPT-2 with ReLU² Instead of GELU: 0.051 Lower Loss After 1B Tokens, From the Same Starting Weights

Three GPT-2 (124M) runs with ReLU² in the MLP ended at 3.623 mean validation loss after 1B FineWeb-Edu tokens, against 3.674 for GPT-2's GELU, at GPT-2's learning rate. Each pair started from identical weights, so the activation is the only difference. The gap clears the series' pre-registered bar (0.010) five times over and was below it at no checkpoint.

GPT-2 with ReLU² Instead of GELU: 0.051 Lower Loss After 1B Tokens, From the Same Starting Weights

GPT-2 with ReLU² Instead of GELU: 0.051 Lower Loss After 1B Tokens, From the Same Starting Weights

We replaced the activation inside GPT-2's MLP blocks, the tanh-approximated GELU, with ReLU², which zeroes negative inputs and squares the positive ones. Nothing else changed: same width, same parameter count, same billion tokens in the same order, same learning rate. Three runs of each. The ReLU² runs ended at a mean validation loss of 3.623; the GELU runs at 3.674.

The gap is 0.051 nats. This series sets, before each set of runs, the smallest gap it will call real. For this comparison that bar was 0.010, so ReLU² clears it about five times over.

This comparison is also the cleanest in the series so far. An activation function has no parameters, so with the same seed the two versions start from exactly the same weights; the logs confirm it. In each pair, the only thing that differs from the first step to the last is what happens to the numbers between the MLP's two matrices.

The short answer

  • ReLU² ended 0.051 lower than GELU (3.623 against 3.674, three runs each). The bar was 0.010.
  • Every ReLU² run was below every GELU run at all eight checkpoints, and each seed's ReLU² run ended 0.048 to 0.057 below the GELU run that started from the same weights.
  • The gap was largest early and was still closing at the end: 0.144 at 0.26B tokens, 0.074 at 0.52B, 0.051 at 1.0B. These runs do not say where it would settle.
  • Speed: no measured difference. Two of the three ReLU² runs trained 2.2% faster than the GELU runs; the third was slowed for part of its run by other jobs sharing the GPUs. Runs of the same setup have differed by up to 3% in this series, so we do not count 2.2% as a speedup.

What changed

The model is the same GPT-2 small as in the earlier posts: 12 layers, width 768, 12 heads, a 1,024-token context, learned position embeddings, LayerNorm, tied input and output embeddings, nanoGPT code. Each MLP block widens to 3,072, applies the activation, and projects back to 768.

GELU (GPT-2)ReLU²
Activationx · Φ(x), tanh approximationmax(0, x)²
Negative inputssmall negative outputs, down to about −0.17exactly 0
Parameters124,475,904124,475,904
Initial weights for seeds 0, 1, 2identical to GELU's (same hashes in the logs)

Everything else matches the earlier posts: FineWeb-Edu shards in the same order, 1,907 steps of 524,288 tokens (1.0B tokens), AdamW with peak learning rate 6e-4, 100 warmup steps and cosine decay to 10%, bfloat16, two A100 80GB cards. The GELU side is the three seed runs from the first measurement post, reused as planned.

Before the runs we checked on the CPU that the two start from the same loss. On the same 4,096 tokens as in the earlier posts, the initial losses were 10.940 (GELU) and 10.950 (ReLU²) for seed 0, and 11.031 and 11.035 for seed 1.

The rule, fixed before the runs

We wrote down this comparison's question, runs and rule on 7 October at 13:16, the day before the first ReLU² run started (8 October, 08:56). The rule is the same as in the previous posts: call the two different only if their mean final losses differ by more than 2 × σ × √(1/3 + 1/3), where σ pools the seed-to-seed variance of every configuration this series has compared so far (GELU, RoPE, RMSNorm, RMSNorm with QK norm, ReLU²). The script that applies the rule also refuses to run if a ReLU² run did not start from the same weights as its GELU partner.

Final validation loss, seeds 0 / 1 / 2MeanStandard deviation
GELU3.6681, 3.6858, 3.66903.67430.0100
ReLU²3.6192, 3.6289, 3.62083.62300.0052

Pooled σ = 0.00629 (10 degrees of freedom), so the bar is 2 × 0.00629 × √(2/3) = 0.0103. The difference is −0.0513.

The three pairs, seed by seed: 3.6681 → 3.6192 (−0.049), 3.6858 → 3.6289 (−0.057), 3.6690 → 3.6208 (−0.048). The seed that was worst with GELU was also worst with ReLU², by a similar margin.

Over the whole run

Two panels. Left: final validation loss for each seed, with a line from the GELU run to the ReLU² run that started from the same weights; all three lines slope down, from a GELU mean of 3.674 to a ReLU² mean of 3.623. Right: the gap between the two means at eight checkpoints from 0.13B to 1.0B training tokens: −0.055 at 0.13B, deepest at −0.144 at 0.26B, then shrinking to −0.051 at 1.0B, always far below the grey band of plus or minus 0.010.
Tokens seenGELUReLU²Gap
0.13B5.57145.5164−0.0550
0.26B4.71464.5706−0.1440
0.39B4.17264.0744−0.0982
0.52B3.94683.8726−0.0742
0.66B3.82113.7582−0.0629
0.79B3.73893.6829−0.0560
0.92B3.69183.6396−0.0522
1.00B3.67433.6230−0.0513

(Means of three runs. At every checkpoint the highest ReLU² run was below the lowest GELU run.)

The gap opens fast, peaks at 0.26B tokens and then shrinks, quickly at first and slowly at the end: by 0.070 between 0.26B and 0.52B tokens, by 0.018 between 0.52B and 0.79B, by 0.005 between 0.79B and 1.0B, and by 0.0009 over the last 82M tokens. The ReLU² runs reached the GELU runs' final loss (3.674) somewhere between 0.79B and 0.92B tokens.

That shape resembles what RoPE did in post 2: a large early advantage that narrows as training goes on. We did not test why either one narrows, and with one budget we cannot say whether the ReLU² gap levels off near 0.05 or keeps closing.

Where ReLU² comes from

So et al. (2021) found ReLU² by searching over the building blocks of a Transformer for cheaper training. Their abstract attributes most of the improvement of the resulting architecture, Primer, to two changes: squaring the ReLU activations and adding a depthwise convolution after the query, key and value projections. In the modded-nanogpt speedrun, ReLU² arrived in record 5 together with QK norm, zero-initialized projections and padded embeddings, so its share of that record's gain was not measured on its own. This post measures it alone, at GPT-2's learning rate.

The series so far, each change measured against the same three GPT-2 runs except where noted:

ChangeGap in final lossVerdict
RoPE instead of learned positions (post 2)−0.122lower
Parameter-free RMSNorm instead of LayerNorm (post 3)+0.038higher
QK norm on top of RMSNorm (against RMSNorm alone)−0.012lower, by 0.001 over the bar
ReLU² instead of GELU (this post)−0.051lower

The gaps are not additive. Each was measured on its own, starting from GPT-2.

Speed

GELUReLU²
Training throughput per run361,173 / 361,211 / 361,185 tokens/s369,119 / 369,225 / 282,952 tokens/s
Wall time per run, including evaluation48.7 min each47.8 / 47.7 / 62.2 min

The third ReLU² run's average is low because for part of the run its throughput fell as far as 121,714 tokens per second; its median logged step ran at 368,867, like the other two. The slow stretch ran from about 10:35 to 11:00 on 8 October. From 10:49 to 11:01 the GPUs were shared with another of our training jobs, a follow-up run for a different post, and the throughput recovered as soon as that job ended. What was slowing the run between 10:35 and 10:49 we did not record. Taking the two undisturbed runs, ReLU² trained 2.2% faster than GELU. But the GELU runs were measured on a different day, and in post 3 three runs of one configuration spread across 3%, so we do not call this a measured speedup. Throughput was never part of the verdict.

What this does not show

  • One learning rate. Both sides used GPT-2's 6e-4. Post 5 is planned to change the optimizer and sweep the learning rate.
  • One budget, one size. GPT-2 small at 1B tokens. The gap was still shrinking at the end.
  • Why it helps. We measured loss and throughput only. How many MLP activations are exactly zero with ReLU², and whether that matters, was left out of the plan.
  • One MLP shape. Width 3,072, no gated variants such as SwiGLU.
  • Throughput belongs to this implementation and this machine.

Next

Next in the plan is the optimizer: Muon against AdamW, with the learning rate swept for each. Post 5's comparison runs have not started yet.

The plan, the scripts and the raw logs of all six runs are in the reproduction package below, no login needed. Its README has the commands that recompute every number in this post from those logs, without a GPU.

Files for this post

Reproduction package

The harness, raw logs and result CSVs behind this post. No account needed; the README has a command that recomputes the tables without a GPU.

gpt2-relu2-repro.zip · 356 KB

Download

Stay Updated

Follow us for the latest posts and tutorials

Subscribe to Newsletter