Models & Algorithms•SOTAAZ Lab••KR

GPT-2 with Parameter-Free RMSNorm: 0.038 Higher Loss After 1B Tokens; QK Norm Won Back 0.012, Barely Past the Bar

Swapping all of GPT-2's LayerNorms for parameter-free RMSNorm raised the mean validation loss after 1B FineWeb-Edu tokens from 3.674 to 3.712, three runs each at GPT-2's learning rate. Adding QK norm on top brought it to 3.701, a 0.012 gain that clears this series' pre-registered bar (0.011) by 0.001. Throughput moved by +1.4% and −0.4%, and no run had an instability to fix.

GPT-2 with Parameter-Free RMSNorm: 0.038 Higher Loss After 1B Tokens; QK Norm Won Back 0.012, Barely Past the Bar

GPT-2 with Parameter-Free RMSNorm: 0.038 Higher Loss After 1B Tokens; QK Norm Won Back 0.012, Barely Past the Bar

I replaced every LayerNorm in GPT-2 (124M) with RMSNorm without learned parameters, the form the modded-nanogpt speedrun and nanochat use. Three runs, same billion tokens in the same order, same learning rate. The mean validation loss went up, from 3.674 to 3.712.

Then I added QK norm on top, which normalizes each attention head's queries and keys before they are multiplied. That brought the loss down to 3.701, a gain of 0.012. The bar this series set in advance for this comparison was 0.011, so it counts, by a margin of 0.001.

This is the third post in a series that starts from GPT-2 and measures, one at a time, what later training recipes changed. The previous post swapped learned position embeddings for RoPE and found a 0.122 gain. This one is a different kind of result: one change made things worse on its own, and the other was close to the line.

The short answer

  • Parameter-free RMSNorm ended 0.038 higher than GPT-2's LayerNorm (3.712 against 3.674, three runs each). The bar was 0.011, so this is a real difference, in the wrong direction. From 0.26B tokens on, every RMSNorm run was above every LayerNorm run at every checkpoint.
  • QK norm on top ended 0.012 lower than RMSNorm alone (3.701 against 3.712). That clears the bar by 0.001. Earlier in training the gap changed sign: at 0.39B tokens QK norm was behind. I count it because the rule says to, and I would not lean on it.
  • The two together still ended 0.026 above GPT-2's LayerNorm.
  • Speed and stability barely moved. Throughput changed by +1.4% (RMSNorm) and −0.4% (with QK norm), and no run in any of the three conditions showed a gradient spike to fix.

What changed

The model is the same GPT-2 small as before: 12 layers, width 768, 12 heads, a 1,024-token context, tied input and output embeddings, learned position embeddings, nanoGPT code with GPT-2's tanh GELU.

LayerNorm (GPT-2)Parameter-free RMSNorm+ QK norm
What each norm doessubtract the mean, divide by the standard deviation, then a learned gain and bias per channeldivide by the root mean square; no mean, no gain, no biasthe same
Where25 norms: two in each of the 12 blocks, one before the output layerthe same 25the same 25, plus queries and keys normalized per head in every attention layer
Parameters124,475,904124,437,504 (the norms' 38,400 gains and biases removed)124,437,504

Everything else is identical to the previous posts: FineWeb-Edu shards read in the same order, 1,907 steps of 524,288 tokens (1.0B tokens), AdamW with peak learning rate 6e-4, 100 warmup steps, cosine decay to 10%, bfloat16, two A100 80GB cards. The attention scale stays at 1/√64 with QK norm, and QK norm has no learned scale. The logs confirm that each compared pair differs only in the switch named, and that all nine runs saw the same first batch. The LayerNorm side is the three seed runs from the first measurement post, reused as planned.

Before the runs I checked on the CPU that the three versions start from the same place. On the same 4,096 tokens, the initial losses for seed 0 were 10.940 (LayerNorm), 10.938 (RMSNorm) and 10.944 (with QK norm). So none of the differences below comes from a different starting point.

RMSNorm here changes two things at once: it drops the mean subtraction and it drops the learned gain and bias. These runs cannot tell which of the two costs the loss. A version with a learned gain would separate them; I did not run it.

The rule, fixed before the runs

Two comparisons, both written down and committed at 13:07, before the first run of this post started at 15:26. The script that applies the rule was committed at 13:55, also before the first run.

  • (a) RMSNorm against LayerNorm
  • (b) RMSNorm plus QK norm against RMSNorm alone

Each counts as a difference only if the gap between the means exceeds 2 × σ × √(1/3 + 1/3), where σ pools the seed-to-seed variance of every configuration run so far: LayerNorm, RoPE, RMSNorm and RMSNorm with QK norm. I also wrote down that making two comparisons raises the chance that one of them clears the bar by luck. That warning applies to (b).

Final validation loss, seeds 0 / 1 / 2MeanStandard deviation
LayerNorm3.6681, 3.6858, 3.66903.67430.0100
RMSNorm3.7128, 3.7194, 3.70453.71220.0075
RMSNorm + QK norm3.6996, 3.6983, 3.70403.70060.0030

Pooled σ = 0.00653 (8 degrees of freedom), so the bar is 2 × 0.00653 × √(2/3) = 0.0107.

ComparisonGapBarVerdict
(a) RMSNorm − LayerNorm+0.03790.0107different: RMSNorm higher
(b) + QK norm − RMSNorm−0.01160.0107different: QK norm lower
for reference: + QK norm − LayerNorm+0.02640.0107different: higher

Over the whole run

Two panels. Left: final validation loss of each run as a dot, three per condition, with a line at the mean: LayerNorm 3.674, parameter-free RMSNorm 3.712, RMSNorm with QK norm 3.701. Right: the two pre-registered gaps at eight checkpoints from 0.13B to 1.0B training tokens. RMSNorm minus LayerNorm stays between +0.029 and +0.043, ending at +0.038, always above the grey band of plus or minus 0.011. QK norm minus RMSNorm swings from −0.024 at 0.26B to +0.006 at 0.39B and then drifts down to −0.012 at the end, just below the band.
Tokens seenLayerNormRMSNorm+ QK norm(a) RMSNorm − LayerNorm(b) QK norm − RMSNorm
0.13B5.57145.60065.5922+0.0292−0.0083
0.26B4.71464.74524.7209+0.0306−0.0243
0.39B4.17264.21524.2208+0.0425+0.0057
0.52B3.94683.98483.9841+0.0380−0.0008
0.66B3.82113.85883.8518+0.0377−0.0070
0.79B3.73893.77543.7666+0.0365−0.0089
0.92B3.69183.72943.7192+0.0376−0.0102
1.00B3.67433.71223.7006+0.0379−0.0116

(Means of three runs. The verdicts above use the final row only; the bar is defined for final losses.)

The two comparisons look nothing alike. RMSNorm's cost appears at the first checkpoint and settles at about 0.037 to 0.038 from 0.52B tokens on; the runs never overlap after 0.13B. QK norm's effect does not hold still: it helps at 0.26B, hurts at 0.39B, is nothing at 0.52B, and then grows slowly to the end. At the final checkpoint all three QK norm runs are below all three RMSNorm runs, but in between, from 0.39B to 0.92B, they overlap.

A gap that was still growing at the last checkpoint might keep growing with more training, or it might not. With this shape and a 0.001 margin over the bar, I read (b) as "probably a small gain, not established". More seeds per side would narrow it down; they are not in the plan.

Why use these at all?

Neither change was introduced to lower the loss of a 124M model at a gentle learning rate, and that is worth keeping in mind before reading this post as "RMSNorm is bad".

  • RMSNorm is mostly about cost. Skipping the mean and the parameters makes the norm cheaper. On this code that showed up as a 1.4% throughput gain (366,346 against 361,190 tokens per second), too small to change the comparison.
  • QK norm is mostly about stability. Wortsman et al. (2023) report that the growth of attention logits, one of the instabilities QK normalization was used against in large models (Dehghani et al., 2023), also appears in small models at high learning rates, and that the mitigations used at large scale work there too. At 6e-4 there was nothing to stabilize: after warmup, the largest gradient norm in any of the nine runs was 1.23, and the mean was 0.39 to 0.45 in every run. QK norm cost 0.4% in throughput.

In the modded-nanogpt speedrun, parameter-free RMSNorm is already in the earliest record whose code is in the repository (record 4, October 2024), alongside the Muon optimizer; no record names it as its change. QK norm arrived in record 5 together with ReLU², zero-initialized projections and padded embeddings. Neither was measured there on its own at GPT-2's learning rate, which is what this post did. Post 5 changes the optimizer and learning rate, and the last post combines everything; whether RMSNorm's cost survives in that company is a separate question.

What this does not show

  • One learning rate. All runs used 6e-4. QK norm is mainly used to keep training stable at higher learning rates; this post did not test that.
  • Which half of RMSNorm costs the loss. Dropping the mean and dropping the gains were changed together.
  • One budget, one size. GPT-2 small at 1B tokens.
  • Comparison (b) is marginal. It clears the bar by 0.001, and it is one of two comparisons made on the same runs.
  • Throughput belongs to this implementation (PyTorch's rms_norm under torch.compile).

Next

Next in the plan is GELU against ReLU² in the MLP, again against the same three GPT-2 runs. Its runs have not started yet.

The plan, the scripts and the raw logs of all nine runs are in the reproduction package below, no login needed. Its README has the commands that recompute every number in this post from those logs, without a GPU.

Files for this post

Reproduction package

The harness, raw logs and result CSVs behind this post. No account needed; the README has a command that recomputes the tables without a GPU.

gpt2-norm-repro.zip · 289 KB

Download

Stay Updated

Follow us for the latest posts and tutorials

Subscribe to Newsletter