Pretraining Without Backprop, Rerun: What Happens to Dust When Backprop Gets Adam or More Epochs?
We reran Q Labs' Dust at 1M tokens with its own code. Under the paper's SGD setup our numbers matched the paper's, and with 20 backprop seeds the gap between Dust at 1,024 draws and backprop was under 0.02 (5.938 vs 5.941), smaller than the paper's 0.025. Backprop with the paper's Adam settings reached 5.384 in the same single epoch, and 16 epochs over the same 1M tokens reached 5.209 with SGD while using about 5% of Dust's compute. Our first Adam run came out worse than SGD; the paper's own settings fixed it, and we note what to check in your own runs.

Pretraining Without Backprop, Rerun: What Happens to Dust When Backprop Gets Adam or More Epochs?
Q Labs' Dust (Dahal, Mandal, Gülbahar and Vegesna) trains a transformer language model without a backward pass. It perturbs activations at every token, measures how the loss changes, and turns those changes into an update. The paper calls it "the first zeroth-order method that is competitive with backprop at pretraining transformer language models," and reports that "at 100k and 1M tokens Dust ends below backprop, from a few hundred draws at 100k and from a thousand at 1M."
The authors are also clear about what they are not claiming: "We do not attempt to make it compute-efficient enough to replace backprop today." Dust at 1,024 draws spends about 350 times backprop's compute per step in our measurement. The question is whether its gradient estimate can match backprop's, not whether it is cheaper.
We wanted to know three things, and fixed the questions, the decision rule and the sentence for each outcome before training anything:
- Rerun. At 1M tokens, with the authors' public code, does Dust at 1,024 draws end below backprop under the same SGD setup?
- Baseline. If backprop gets a tuned Adam-style optimizer instead, does the order hold within one epoch?
- Compute. If backprop may see the same 1M tokens several times, using no more compute than Dust, does the order change?
Questions 2 and 3 are not in the paper. A different answer there would not make the paper wrong, because the paper fixed one epoch and SGD on purpose, to compare the two gradient estimates under the same optimizer.
Setup
- Code: the authors' repository (qlabs-eng/dust at
b20f7c0, MIT), unmodified. Its README calls it a "minimal implementation ... omitting execution optimizations used in the full experiments." - Data and model: the repository's 1M-token FineWeb split (file hashes match its manifest), a 4,096-token BPE vocabulary, 8 layers of width 512 (37.7M parameters), batches of 16k tokens, so 61 steps per epoch.
- Measure: test loss at the checkpoint with the lowest validation loss, as in the paper. Seeds 0, 1 and 2 for every cell.
- Rule: three seeds per condition (SGD backprop later extended to 20, see below). A difference counts only if it exceeds twice its standard error. Otherwise we write that we can't tell.
- Compute: FLOPs per step measured with PyTorch's FLOP counter. One backprop step costs 4.6 × 10¹²; Dust costs 91 times that at 256 draws and 347 times at 1,024.
- Hardware: one A100 80GB, shared with another training job, so we report losses and FLOPs but not wall-clock time.

1. The rerun: the paper's numbers, and a gap under 0.02
| 1M tokens, one epoch | Paper | Ours | Compute vs backprop |
|---|---|---|---|
| Backprop, SGD | 5.959 | 5.941 (20 seeds) | 1× |
| Dust, 256 draws | 6.067 | 6.078 (3 seeds) | 91× |
| Dust, 1,024 draws | 5.934 | 5.938 (3 seeds) | 347× |
The paper's values come from the data behind its Figure 2. Our numbers land within about 0.02 of each.
With three seeds each, we couldn't tell which method was lower: Dust at 1,024 draws was 0.004 below backprop, and seed changes alone could move the gap by 0.026. Almost all of that spread came from backprop, whose result varied a lot between seeds while Dust's barely moved. So we added backprop seeds, which take about three minutes each, and left Dust at three, since each Dust run takes over 80 minutes and varies little. We fixed the decision rule before adding them.
With 20 seeds, backprop ranged from 5.905 to 6.035 and averaged 5.941. Dust was 0.003 lower, and the range consistent with our data runs from Dust 0.018 lower to 0.011 higher. At 1M tokens, any gap between the two is under about 0.02.
The paper's gap, Dust 0.025 lower, is outside that range, so we did not see a gap that large. Our Dust value matched the paper's (5.938 vs 5.934); the difference is on the backprop side. With backprop varying this much between seeds, two three-seed averages can easily differ by 0.02.
At 256 draws Dust was 0.136 higher than backprop, well beyond what seed changes produce, which is the same direction the paper reports ("from a thousand at 1M").
2. Backprop with Adam: the first result looked wrong
Our first Adam-style run used AdamW the common way: one learning rate for every weight, weight decay 0.1, a short warmup and a cosine decay, with the rate picked on validation (0.01).
It reached 6.427, worse than plain SGD backprop (5.942) and worse than Dust. Adam usually trains at least as well as SGD, so we read this as a setup problem rather than a property of Adam.
Our suspect was the embeddings, the layer that turns tokens into vectors. One epoch here is only 61 updates. With a learning rate sized for the other weights, the embeddings may barely move before training ends. The paper treats them that way too. Its SGD recipe gives the token embedding a separate, much larger rate, and its Adam backprop setting (Table 7) uses three rate groups: 0.002 for weight matrices and the output head, 0.3 for token and value embeddings, and 0.2 for the residual scalars, with β₁ 0.8 and no weight decay.
So we reran backprop with exactly those settings. The paper gives no schedule, so we kept the rate constant as in its SGD setup. Three seeds averaged 5.384, against the paper's 5.361. With the setup fixed, Adam behaved as expected. We changed the rate groups, β₁, weight decay and the schedule together, so we can't pin the whole difference on the embedding rate, and we did not measure how far the embeddings moved. It is the most likely cause, not a proven one.
This is worth checking in your own runs. With few updates or little data, a single learning rate for every weight can leave the embeddings undertrained. If Adam comes out worse than SGD, look at per-layer learning rates before blaming the optimizer.
The corrected backprop was 0.554 below Dust at 1,024 draws (5.938), far beyond what seed changes alone produce here (0.015). It is a comparison across optimizers, though, Dust with SGD against backprop with Adam. The paper makes the like-for-like comparison itself: under Adam at 1M tokens, Dust at 1,024 draws reaches 5.594 against backprop's 5.361, and only the curve fitted through larger populations, extrapolated to infinitely many draws, reaches 5.248, below backprop. We did not run Dust under Adam.
3. More epochs on the same data
The paper fixes one pass over the data. If the data is limited to 1M tokens but compute is not, backprop can simply see the data again.
We gave backprop 4, 16 or 32 epochs and chose the count on validation loss for seed 0, never on test loss. Both optimizers picked 16. With the paper's SGD recipe, 16 epochs reached 5.209, which is 0.729 below Dust at 1,024 draws, far beyond the noise (twice the standard error is 0.009). That used 16 times one epoch of backprop, about 4.6% of the compute of one Dust run at 1,024 draws.
The single-rate AdamW from section 2 reached 4.905 at 16 epochs. Sixteen epochs means 976 updates, so the same rate has far more chances to train every layer. Treat it as a lower bound on what Adam does here, not as a tuned result.
What this says about Dust
- The paper's numbers reproduced in its own setting. With the public minimal code, Dust at 1,024 draws landed within 0.004 of the paper's value, and within 0.02 of SGD backprop. That is notable for a method with no backward pass.
- At 1M tokens, Dust and backprop are within 0.02 of each other. With 20 backprop seeds, Dust was at most 0.018 lower or 0.011 higher; the paper's 0.025 gap did not show up.
- Change the baseline and backprop is ahead. With the paper's own Adam settings, or with repeated epochs on the same data at a few percent of Dust's compute, backprop's loss was much lower. Neither condition is what the paper set out to compare, and the paper says Dust is not meant to replace backprop today.
- Dust belongs to an old family. The paper describes it as node perturbation and cites Widrow and Lehr (1990) and Werfel et al. (2003). Its claim is about making that family work for transformer pretraining, not about inventing it.
Limits
- Only 1M tokens. The paper's more interesting claims are at 10M and 20M, where the gap narrows with population, and we did not run those.
- The public code is a minimal implementation. The paper's runs used optimized code, which could matter at larger populations.
- SGD backprop has 20 seeds; every other condition has three, and differences smaller than about 0.03 are invisible in those comparisons.
- The Adam run with the paper's settings was added after the first result. Its values came from the paper, not from our tuning.
- We did not run Dust under Adam, so we have no like-for-like Adam comparison of our own.
- No wall-clock times: the GPU was shared with another job throughout.
The plan, scripts, per-run logs and results are in the reproduction package below. The newsletter covers the next rerun when it is published.
Files for this post
Reproduction package
The harness, raw logs and result CSVs behind this post. No account needed; the README has a command that recomputes the tables without a GPU.
dust-backprop-free-rerun-repro.zip · 1,221 KB