Models & Algorithms•SOTAAZ Lab••KR

GPT-2 with RoPE Instead of Learned Position Embeddings: 0.122 Lower Loss After 1B Tokens, 5% Slower

Three GPT-2 (124M) runs with rotary position embeddings ended at 3.552 mean validation loss after 1B FineWeb-Edu tokens, against 3.674 for three runs with GPT-2's learned position embeddings. The gap clears this series' pre-registered bar (0.012) by ten times, at GPT-2's learning rate. RoPE ran 5.3% slower on this code.

GPT-2 with RoPE Instead of Learned Position Embeddings: 0.122 Lower Loss After 1B Tokens, 5% Slower

GPT-2 with RoPE Instead of Learned Position Embeddings: 0.122 Lower Loss After 1B Tokens, 5% Slower

I took GPT-2 (124M), removed its learned position embeddings, put rotary position embeddings (RoPE) in their place and changed nothing else. Three runs of each, same billion tokens in the same order, same learning rate. The RoPE runs ended at a mean validation loss of 3.552. The GPT-2 runs ended at 3.674.

That is a gap of 0.122 nats. The previous post measured how far apart runs land when only the random seed changes, and set the bar a difference has to clear in this series. With the runs here, that bar is 0.012. RoPE clears it about ten times over.

This is the second change in a series that starts from GPT-2 and adds, one at a time, what later training recipes changed. In the modded-nanogpt speedrun, rotary embeddings arrived in record 2 together with a retuned learning rate, and the two were credited jointly. Here RoPE is the only change, at GPT-2's own learning rate of 6e-4.

The short answer

  • RoPE ended 0.122 nats lower (3.552 against 3.674, mean of three runs each). The pre-registered bar for this comparison was 0.012, so this counts as a real difference.
  • Every RoPE run beat every learned-position run at every checkpoint. The worst RoPE run finished at 3.554; the best learned-position run at 3.668.
  • The gap was largest early and was still closing at the end: 0.443 at 0.13B tokens, 0.161 at 0.52B, 0.122 at 1.0B. These runs do not say where it would settle with more training.
  • RoPE cost 5.3% in training throughput on this code (342,000 against 361,000 tokens per second on two A100s). Its runs dropped below the learned-position runs' final loss somewhere between 0.66B and 0.79B tokens, well before that cost adds up.

What changed and what did not

The model is the same GPT-2 small as in the previous post: 12 layers, width 768, 12 heads of 64 dimensions, a 1,024-token context, tied input and output embeddings, nanoGPT code with GPT-2's tanh GELU. One switch differs.

Learned positions (GPT-2)RoPE
Position informationa 1,024 × 768 table added to the token embeddings at the inputqueries and keys rotated by an angle that depends on position, in every attention layer
Parameters124,475,904123,689,472 (the table's 786,432 removed, 0.6%)
RoPE details—base 10,000, all 64 dimensions of each head, queries and keys only

Everything else is identical: FineWeb-Edu shards read in the same order, 1,907 steps of 524,288 tokens (1.0B tokens), AdamW with peak learning rate 6e-4, 100 warmup steps and cosine decay to 10%, bfloat16, two A100 80GB cards. The training logs confirm that each pair of runs differs in that one setting and that every run saw the same first batch.

The learned-position side is not new: it is the three seed runs from the previous post, reused as planned. The RoPE runs use the same seeds, 0, 1 and 2. That does not give them the same initial weights, because without the position table the random number generator is consumed differently during initialization.

Before any RoPE run I checked the implementation on the CPU for the property that defines it: the attention score between a query and a key should depend only on how far apart they are. Three query-key pairs, each seven positions apart, gave the same dot product (−14.552), and rotating a vector did not change its length.

The rule, fixed before the runs

The previous post fixed the rule: call two configurations different only if their mean final losses differ by more than 2 × σ × √(1/n₁ + 1/n₂), where σ pools the seed-to-seed standard deviation of both configurations. I wrote this comparison's question, runs and rule down and committed them at 10:11, before the first RoPE run started at 12:50. The script that applies the rule was committed later, at 13:55, after the first RoPE run had finished, so I do not claim it was written blind. It implements the rule as committed and adds nothing to it.

RunsFinal validation lossMeanStandard deviation
Learned positionsseeds 0, 1, 23.6681, 3.6858, 3.66903.67430.0100
RoPEseeds 0, 1, 23.5543, 3.5519, 3.54943.55190.0025

Pooled σ = 0.0073, so the bar is 2 × 0.0073 × √(2/3) = 0.0119. The difference is −0.1224.

The bar came out lower than the 0.0163 the previous post quoted for three seeds, because the three RoPE runs happened to sit closer together (standard deviation 0.0025) and pooling pulls σ down. Three runs pin a standard deviation down loosely, so I would not read the tighter RoPE spread as a property of RoPE. It does not affect the verdict here: the gap is more than seven times either bar.

Over the whole run

Two panels. Left: validation loss at eight checkpoints from 0.13B to 1.0B training tokens for three learned-position runs (orange) and three RoPE runs (blue), with their means; RoPE is lower at every checkpoint, ending at 3.552 against 3.674. Right: the gap between the two means at each checkpoint, from −0.443 at 0.13B tokens shrinking to −0.122 at 1.0B, every bar far below the dashed line at −0.012 that marks the bar for a real difference.
Tokens seenLearned positionsRoPEGap
0.13B5.57145.1280−0.4434
0.26B4.71464.3357−0.3789
0.39B4.17263.9628−0.2098
0.52B3.94683.7859−0.1609
0.66B3.82113.6811−0.1399
0.79B3.73893.6089−0.1300
0.92B3.69183.5670−0.1248
1.00B3.67433.5519−0.1224

(Means of three runs. The two series' runs never overlapped: at every checkpoint the highest RoPE run was below the lowest learned-position run.)

Two things are in the table. The difference shows up from the first checkpoint, at 0.13B tokens, which answers the plan's second question: it does not open up late in training, it is there from the start. And it shrinks steadily, quickly at first and then slowly: by 0.28 between the first checkpoint and the halfway point (0.52B tokens), and by 0.04 over the second half.

One reading fits that shape. A learned position table starts as random numbers (normal, standard deviation 0.02) and has to learn from the data what each of the 1,024 positions means; rotary embeddings give attention the relative distance between tokens from the first step, with nothing to learn. If that is the mechanism, the learned table is catching up, and a longer run would narrow the gap further. I did not test this. The gap was still shrinking at the last checkpoint (by 0.0024 over the final 82M tokens), and these runs cannot say whether it closes, settles or stops.

The cost

Learned positionsRoPE
Training throughput (mean of 3 runs)361,190 tokens/s342,060 tokens/s (−5.3%)
Wall time per run, including evaluation48.7 min51.7 min

I did not profile where the time goes. The rotation is extra element-wise work on the queries and keys in all 12 layers, written in plain PyTorch and left to torch.compile; a fused kernel would likely cost less.

The cost does not change the comparison. The RoPE runs were already below the learned-position runs' final loss (3.674) at the 0.79B-token checkpoint (3.609) and still above it at 0.66B (3.681). At RoPE's throughput, 1,500 steps take 38 minutes of training time, against 46 for all 1,907 steps of a learned-position run (both computed from the logged throughput, without evaluation and compile time).

What this does not show

  • One learning rate. Both sides used GPT-2's 6e-4. The speedrun changed the learning rate together with RoPE; a learning rate tuned for each side could move both numbers, and I have not separated the two. Post 5 is planned to sweep the learning rate.
  • One budget, one size. GPT-2 small at 1B tokens. The gap was still shrinking at the end.
  • No long-context test. RoPE is often chosen for how it extends past the training length. Every run here trained and was evaluated at 1,024 tokens.
  • One RoPE setting. Base 10,000, applied to all dimensions. Partial RoPE and other bases were not tried.
  • The throughput number belongs to this implementation. Another kernel would give another number.

Next

The next post replaces LayerNorm with RMSNorm and then adds QK-norm on top, each against the same three GPT-2 runs. Those runs are still in progress.

The plan, the scripts and the raw logs of all six runs are in the reproduction package below, no login needed. Its README has the commands that recompute every number in this post from those logs, without a GPU.

Reproduction package

The harness, raw logs and result CSVs behind this post. No account needed; the README has a command that recomputes the tables without a GPU.

gpt2-rope-repro.zip · 155 KB

Download

Stay Updated

Follow us for the latest posts and tutorials

Subscribe to Newsletter

Related Posts