tiny-llm / Training studies / 2026-09-17
How small models learn.
Across schedules and workers.
The latest hyperparameter tuning for 20.4M-parameter language models on C4. Cosine-to-zero and WSD, synchronous and decentralized training, measured across token budgets.
Synchronous training has the lowest selected mean loss at each shared horizon. The decentralized cosine results draw closer as the training budget grows.
Explore all results ↗Cosine-to-zero meets WSD.
At the shared 20, 40, and 80-token horizons, the selected cosine recipes have lower mean final validation loss for both measured training modes.
Cosine starts each horizon afresh. WSD continues earlier training and prunes learning rates. Their search grids and trainer versions differ, so these results describe the measured recipes.
| Tokens/parameter | WSD − cosine · sync | WSD − cosine · eight workers |
|---|---|---|
| 20 | +0.006881 | +0.007539 |
| 40 | +0.022921 | +0.028831 |
| 80 | +0.027347 | +0.039969 |
Four-worker WSD is unavailable. WSD at 120 and 160 tokens per parameter is available in the extended results.
One global budget. One, four, or eight workers.
Decentralized workers divide the global token budget. Their parameters are averaged before validation. These cosine experiments compare the best tested recipe for each mode.
Best tested hyperparameters
| Schedule | Training mode | Tokens/parameter | LR | β₁ | β₂ | Final loss ± sample SD |
|---|---|---|---|---|---|---|
| Cosine-to-zero | Synchronous | 20 | 0.01 | 0.9 | 0.99 | 3.549818 ± 0.003284 |
| Cosine-to-zero | Four workers | 20 | 0.01 | 0.95 | 0.99 | 3.570302 ± 0.004893 |
| Cosine-to-zero | Eight workers | 20 | 0.014 | 0.95 | 0.99 | 3.585589 ± 0.003106 |
| WSD | Synchronous | 20 | 0.004 | 0.95 | 0.98 | 3.556699 ± 0.002045 |
| WSD | Eight workers | 20 | 0.006 | 0.95 | 0.999 | 3.593128 ± 0.004363 |
| Cosine-to-zero | Synchronous | 40 | 0.01 | 0.9 | 0.98 | 3.429879 ± 0.003875 |
| Cosine-to-zero | Four workers | 40 | 0.01 | 0.95 | 0.99 | 3.441328 ± 0.001791 |
| Cosine-to-zero | Eight workers | 40 | 0.012 | 0.95 | 0.999 | 3.447451 ± 0.003544 |
| WSD | Synchronous | 40 | 0.004 | 0.95 | 0.98 | 3.452800 ± 0.003136 |
| WSD | Eight workers | 40 | 0.006 | 0.95 | 0.999 | 3.476282 ± 0.002687 |
| Cosine-to-zero | Synchronous | 80 | 0.01 | 0.9 | 0.99 | 3.347422 ± 0.003371 |
| Cosine-to-zero | Four workers | 80 | 0.01 | 0.974 | 0.999 | 3.351158 ± 0.004694 |
| Cosine-to-zero | Eight workers | 80 | 0.012 | 0.974 | 0.999 | 3.355351 ± 0.001988 |
| WSD | Synchronous | 80 | 0.003 | 0.95 | 0.999 | 3.374770 ± 0.003080 |
| WSD | Eight workers | 80 | 0.003 | 0.95 | 0.999 | 3.395320 ± 0.004652 |
Three seeds per configuration. SD describes seed variation, not a confidence interval. Validation is used for tuning. Inspect rankings and matched comparisons ↗
What stays shared.
What changes.
The same model size, token-cache identity, global batches, and validation target are used across the current campaigns. Schedules, search grids, and continuation policies differ.
Read the experimental protocol ↗- Model
- 20,403,520 parameters · context 1,024
- Global batch
- 131,072 prediction targets
- Warmup
- 312 optimizer updates
- Selection
- Mean final full-validation loss · seeds 42–44
- Validation
- 197,411,295 prediction targets
- Execution
- All workers packed on one GH200
From a cache to an experiment.
Prepare data ↗
Create the C4 cache with room for training and unseen analysis.
Train locally ↗
Choose a schedule, launch sync or packed workers, and retain checkpoints.
Submit Slurm jobs ↗
Request a GPU, launch a run, and monitor its progress.
Analyze the model ↗
Measure Hessian and gradient-noise alignment from saved checkpoints.
The published cosine sweeps saved no checkpoints. Retain weights in a new run when you want to perform Hessian analysis.
Measured on a single GH200.
Observed steady training throughput at 80 tokens per parameter, calculated from complete logging windows after warmup. Every completed run contributes.
Synchronous
1.745M tokens/sIQR 1.730–1.758 · 72 runs
39.48 GiB median peak allocation
Four workers
1.599M tokens/sIQR 1.587–1.610 · 144 runs
40.78 GiB median peak allocation
Eight workers
1.522M tokens/sIQR 1.508–1.535 · 180 runs
42.29 GiB median peak allocation
All horizons, timing boundaries, and memory measurements ↗
BF16 autocast · FP32 parameters · compiled SDPA · global batch 131,072. These are single-GPU measurements, with all workers sharing the device.