An interactive research essay

The best optimizer depends on batch size.

Same data budget. Different batch size. A different winner. Discover why, one experiment at a time.

Noisy quadratic · 4,096 samples
SGDNewtonNewton leads

Retuned learning rates. A toy model, not language-model data. Surface height: log(1 + loss).

The question

If you change the batch size,
should you change the optimizer, too?

A leaderboard with a moving winner.

We often ask which optimizer trains a model best. But the answer usually comes from a benchmark at one batch size. Increase the batch, and the same ranking need not survive.

In the paper’s language-model experiments, each optimizer is independently retuned at every batch size. The model sees roughly 1.7 billion training tokens either way. Even after that retuning, the winner changes.

Measured in the paper

Move the batch. Watch the ranking.

FineWeb validation loss ↓ lower is better
Validation loss versus batch size
At 128K tokens / stepLoss

Read the crossover At 128K, SOAP leads. At 2M, Shampoo leads. This is a comparison at a fixed token budget, not a wall-clock speed contest. A better optimizer at one batch is not automatically better at another.

Take the noisy way down.

Imagine a valley: steep across, gentle along. An optimizer follows a gradient toward the bottom, but each minibatch gives it a noisy estimate. Larger batches give cleaner estimates, and fewer updates for the same amount of data.

Newton’s method rescales the directions using curvature. That helps it move along the valley, but also scales the noise. SGD leaves that geometry alone. Play with the balance between fast progress and noisy steps.

Live simulation

Two optimizers. One sample budget.

Click the valley to move the startSEED 7
SGD Newtonx: flat & noisy · y: sharp & noiseless
0 / 4,096 samples
Expected final loss

Solid color: remaining initialization error. Hatched color: noise contribution. Calculated exactly for the displayed learning rates; the animated paths are single random draws.

Loss along this random run
Open the hood: the model and the tuning

The objective is L(w) = ½(w₁² + h₂w₂²). Gradients are ĝᵢ = hᵢwᵢ + √(cᵢ/B) zᵢ, with independent standard Gaussian noise at every step. Here h₁ = 1, c₁ is the noise slider, and c₂ = 0. Both methods start at the same point and use the same underlying noise draws.

SGD uses wᵢ ← wᵢ − ηĝᵢ. Newton uses wᵢ ← wᵢ − (η/hᵢ)ĝᵢ. The learning rate is constant within each run. The canvas shows original parameter coordinates.

For aᵢ = η (SGD) or η/hᵢ (Newton), qᵢ = (1 − aᵢhᵢ)² and S = 4096/B, expected final loss is

½ ∑ᵢ hᵢ [w₀,ᵢ² qᵢˢ + (aᵢ²cᵢ/B) ∑ⱼ₌₀ˢ⁻¹ qᵢʲ].

We search 701 logarithmically spaced learning rates in the stable interval and refine locally. This is a numerical illustration of the mechanism in Section 5, not the paper’s theorem construction or a proof that every landscape has a crossover.

Preconditioning changes both the progress and the noise.

One dial. Two different needs.

Why not just scale the learning rate with the batch size? The paper tested 216 rules in language modeling and 648 in image classification. The best common rules differed: fixed learning rates with increasing matrix weight decay in the language model, square-root learning-rate scaling in CIFAR-5M.

The noisy quadratic gives an intuition. In SignSGD with momentum, a noisy direction benefits when a larger batch makes its estimated sign more reliable. A direction with an already reliable sign mostly gets fewer steps. The learning-rate adjustment that compensates for one need not compensate for the other.

Analytic illustration

Can one exponent preserve both directions?

η′ = η × (B′/B)α
High CNR (1) Low CNR (0.001)Dashed: same movement as reference
Fixed (0)Square-root (½)Linear (1)

CNR = curvature² / single-example noise variance. Both directions are evaluated at w = 1 with momentum 0.9. These are local expected movements, not final training losses.

Where do these curves come from?

At a fixed parameter value, the stationary momentum has an expected sign A(B) = erf(w√[CNR · B(1 + μ)/(2(1 − μ))]). Relative movement per processed sample is κα−1A(κB)/A(B). We use reference B = 1. Noise-dominated response grows approximately as √B; signal-dominated response saturates. Preserving this local movement motivates square-root or linear scaling in the respective limits. The approximation freezes the parameter and assumes stationary momentum; it does not guarantee full-trajectory invariance. See Section 5 and Appendix C of the paper.

What the model explains Curvature alone is not enough: noise differs across directions too. The paper’s direction-dependent scaling results concern SignSGD with momentum. They do not establish one universal formula for every optimizer.

A tiny subspace. A surprisingly large effect.

Now switch a language model from 128K to 2M tokens per matrix update; a 16× increase. Keep the small-batch updates in just a few selected directions. Use the large-batch updates everywhere else.

Choosing the sharpest 768 out of roughly 85 million hidden-matrix directions removes up to 59.5% of the local loss penalty. A random subspace of the same size changes it by at most 3%. Which directions you preserve matters.

Measured in the paper

Choose the directions you keep small.

0.0009%768 / ~85 million directions
1,000 · early12,000 · late
Local branch penalty ↓ lower is better10⁻³ nats

Penalty = branch validation loss − small-batch control loss, after 1,024 base steps (~134M tokens). Control loss is the zero reference.

Local penalty removed59.5%

A local intervention, not an end-to-end training speedup. Missing branches remain marked “not run.”

Download endpoints ↓
Protocol, controls, and what this does not establish

The held subspace is fixed at each anchor and comes from a Muon-preconditioned Gauss-Newton curvature proxy. Both Muon copies receive the same stream of base-batch gradients; the large-batch copy accumulates 16 of them before an update. Matrix weight decay and auxiliary AdamW updates stay on the base clock. This is a diagnostic intervention, not an ordinary large-batch training run.

Endpoints use 10.5M held-out tokens. Most branches are single runs, so these percentages are descriptive measurements, not significance estimates. The experiment selects directions by curvature and does not measure the full direction-dependent noise covariance. Its effect largely disappears near the end of training. We therefore cannot infer a 59.5% improvement in final loss, total training cost, or wall-clock speed.

Benchmark the optimizer.
And the batch size.

Evaluate

Compare across several batch sizes, with hyperparameters retuned at each one.

Explain

Look at how preconditioning changes both useful motion and gradient noise.

Investigate

Explore scaling that accounts for different directions. The intervention is a clue, not a finished optimizer.

More writing by Xingyu Dang ↗

Evidence & further reading

This essay accompanies The Best Optimizer Depends on Batch Size, ICLR 2027 manuscript, source revision 462dc51 (September 28, 2026). It presents measured results from the paper alongside two clearly labeled mathematical illustrations.

Simulation: two-dimensional diagonal quadratic, fixed 4,096-sample budget, exact moment-based learning-rate search. The animations are explanatory experiments, not new evidence about language models. Data provenance ↗