Research archive

AWC packed-8 20M tuning at 128K tokens per batch

Archived reference. Read the current results →

Historical report; retained as an experimental record.

Historical report. Settings and conclusions describe the recorded experiment. See the current documentation for maintained guidance.

Measured on 2026-09-12: 240 successful runs covering the complete 80-configuration LR/beta1/beta2 grid, with runtime seeds 42, 43, and 44 for every configuration. Each run trains eight local 20M models using AWC.

The best tested recipe is LR 0.0056, beta1 0.95, beta2 0.999, with mean final full-validation cross-entropy 3.609816 ± 0.005523 nats. All ± values and error bars are sample standard deviations across the three seeds, not confidence intervals. Lower loss is better.

LR 0.008 with the same betas measures 3.610121 ± 0.003780, only 0.000305 nats higher. This gap is small relative to seed variation. The winner is at the lowest tested LR and highest tested beta2; the search does not locate an interior optimum in those dimensions. Selection uses validation data, with no independent test estimate. These results do not establish statistical significance, equivalence, or global optimality.

Search space

ParameterValues
Learning rate0.0056, 0.008, 0.010, 0.0112, 0.016
AdamW beta10.9, 0.925, 0.95, 0.975
AdamW beta20.95, 0.98, 0.99, 0.999
Runtime seeds42, 43, 44

The repeated 0.010 in the requested LR list is represented once. Every unique combination has three runs: 5 × 4 × 4 × 3 = 240. All finite outcomes are included. Ranking minimizes mean final full-validation loss, breaking exact ties by lower LR, then beta1, then beta2. There are no missing or failed configurations.

Beta1/beta2 heatmaps

Complete beta1/beta2 heatmaps at all five learning rates

PDF · PNG. Each panel contains 16 configurations and 48 seeded runs. All five panels share a color scale; cells show mean ± sample SD. Orange outlines mark the lowest mean within each LR, so the outlines are conditional selections.

The best configurations at each learning rate are:

LRBeta1Beta2Mean loss ± sample SD
0.00560.950.9993.609816 ± 0.005523
0.0080.950.9993.610121 ± 0.003780
0.010.950.9993.612587 ± 0.005239
0.01120.9250.993.614815 ± 0.006469
0.0160.950.983.683340 ± 0.017137

Beta1 0.95 and beta2 0.999 lead at LR 0.0056, 0.008, and 0.010. At LR 0.0112, the lowest mean uses beta1 0.925 and beta2 0.99. At LR 0.016, beta1 0.95 and beta2 0.98 perform best, but even that selected mean is 0.073524 nats above the overall minimum. The preferred beta2 depends on LR; a single marginal average would obscure that relationship.

Learning-rate response

Learning-rate response for all four beta1 values

PDF · PNG. Each panel fixes beta1 and shows all four beta2 curves across the five LRs. Error bars show sample SD; all panels share the same loss scale. Lines connect measured points. The star identifies the best tested configuration.

The three leading configurations use beta1 0.95 and beta2 0.999 across LRs 0.0056–0.010. Their closely spaced means support considering this region together when interpreting the sweep. Higher LR does not improve the best achievable mean within the measured beta grid; the selected minimum at LR 0.016 is clearly higher in these measurements.

Complete ranking

All 80 configurations are included below and in the summary CSV.

RankLRBeta1Beta2Mean loss (nats)Sample SD
10.00560.950.9993.6098160.005523
20.0080.950.9993.6101210.003780
30.010.950.9993.6125870.005239
40.01120.9250.993.6148150.006469
50.00560.9250.9993.6155850.004953
60.010.950.993.6164930.005410
70.0080.9250.993.6168540.008050
80.0080.9250.9993.6181100.007121
90.0080.950.993.6196030.005069
100.010.9250.993.6205300.009368
110.010.950.983.6209460.009662
120.00560.950.993.6225430.003616
130.01120.950.993.6228090.017091
140.00560.9250.993.6231260.005860
150.00560.90.9993.6232350.003125
160.010.90.993.6237310.003187
170.0080.90.9993.6237850.014239
180.01120.950.983.6240880.002247
190.0080.950.983.6295510.003996
200.010.9250.9993.6300010.005558
210.01120.90.983.6308610.001472
220.01120.9250.983.6314950.012256
230.01120.950.9993.6328720.021452
240.0080.9250.983.6333030.003561
250.00560.90.993.6334990.002481
260.010.9250.983.6341250.002443
270.0080.90.993.6346990.009648
280.00560.950.983.6367750.003345
290.01120.90.9993.6383340.003647
300.010.90.983.6392490.011188
310.00560.9250.983.6407310.006696
320.0080.90.983.6414970.016445
330.00560.90.983.6485330.007157
340.010.90.9993.6497520.022700
350.010.9750.9993.6571420.006067
360.01120.90.993.6614060.022119
370.0080.950.953.6628510.006363
380.010.9750.993.6636850.015993
390.00560.9750.9993.6652910.007835
400.01120.9250.9993.6653680.060654
410.0080.9750.9993.6661750.004378
420.01120.9750.9993.6668700.027670
430.0080.9250.953.6673500.002067
440.0080.9750.993.6675540.006194
450.01120.9250.953.6716700.007558
460.00560.9750.993.6717730.003918
470.01120.950.953.6748550.018216
480.01120.9750.993.6758320.023010
490.00560.950.953.6761500.001768
500.010.950.953.6774390.011847
510.010.9250.953.6793690.006412
520.0080.9750.983.6804230.010460
530.010.9750.983.6818350.025892
540.00560.9250.953.6819810.001346
550.0160.950.983.6833400.017137
560.01120.9750.983.6867840.015275
570.0080.90.953.6881870.003637
580.010.90.953.6907360.003331
590.0160.950.993.6930780.017020
600.01120.90.953.6932480.007327
610.00560.9750.983.6939420.012194
620.00560.90.953.6953940.004024
630.0160.9250.993.7049630.011708
640.0160.9750.993.7053370.020623
650.0160.9750.983.7075140.019757
660.0160.9250.983.7119790.004783
670.0160.950.953.7224260.017962
680.0080.9750.953.7228340.007294
690.00560.9750.953.7298510.005455
700.0160.9750.953.7351070.015670
710.01120.9750.953.7369990.012860
720.0160.90.983.7394400.027980
730.0160.9250.953.7405620.004757
740.010.9750.953.7425660.006432
750.0160.90.953.7494910.025530
760.0160.90.993.7757550.023854
770.0160.9750.9993.8067130.038797
780.0160.950.9993.8892720.045119
790.0160.9250.9994.0637240.025521
800.0160.90.9994.1149620.072276

The selected recipe’s individual final full-validation losses are:

Runtime seedLoss (nats)
423.603487168
433.612296254
443.613663870

Comparison with four-worker AWC

The four-worker AWC study selects LR 0.008, beta1 0.95, beta2 0.99, with 3.595559 ± 0.005025. The best eight-worker mean is 0.014257 nats higher. At that same four-worker recipe, eight workers measure 3.619603 ± 0.005069, a difference of 0.024045 nats.

There are 28 shared LR/beta1/beta2 configurations, with the same runtime seeds. Eight workers have lower means in 13 and higher means in 15. The complete matched comparison CSV contains means, sample SDs, and signed differences (eight minus four). The eight-worker winner has no exact four-worker counterpart in the measured grid.

Both studies use the same model architecture, global batch and training-token budget, cache identity, topology family, and averaged-model validation. Eight workers use microbatch 16 and 51,008,512 targets per local model; four workers use microbatch 32 and 102,017,024 targets per local model. Changing worker count also changes the mixing schedule and number of local optimizer states. The separately selected best recipes differ in LR and beta2, and the search spaces are unequal. This comparison is descriptive, not a controlled estimate of a single mechanism or a distributed scaling benchmark.

Shared training protocol

AWC computes forward and backward passes and clips local gradients, mixes parameters, then applies local AdamW updates. Optimizer moments remain local; adaptive consensus is disabled. Validation evaluates the globally averaged model.

SettingValue
Local models8 × 20,403,520 parameters; 163,228,160 total local parameters
Architecture per model8 layers, width 320, 5 heads, FFN 896, vocabulary 32,000
Topology / schemeone_peer_exponential / AWC
Context / local microbatch1,024 tokens / 16 sequences
Global batch8 × 16 × 1,024 = 131,072 targets; no accumulation
Global / per-model training budget408,068,096 / 51,008,512 targets
Epochs / updates40 / 3,119, including shortened epoch-ending steps
Other AdamW settingsWeight decay 0.1, epsilon 1e-8, local gradient clipping 1.0
Schedule5% token-based linear warmup; cosine decay to 10% of peak LR
Epoch evaluationFixed 1,024-block validation subset
Final evaluationFull C4 validation split: 197,411,295 targets
Preprocessing / validation seeds42 / 12345
ExecutionGH200 120GB, BF16 autocast, FP32 weights, compiled SDPA, fused AdamW, 8 CPU threads
RuntimePython 3.12.13, PyTorch 2.14.0, CUDA 13.0
Checkpoint policynone; no model, optimizer, or weight files saved

Each job simulates eight local workers on one GPU. Physical inter-node communication costs are not measured. Runtime seeds vary initialization and sample order; nondeterministic execution does not guarantee bitwise replay.

Median run time, including validation, was 445.8 seconds (range 432.3–590.7 seconds). Runs used at most 24 concurrent GPUs with a 30-minute allocation limit. No training checkpoints were saved; configuration, environment, logs, metrics, and final validation results are retained.

Audit and artifacts

All 240 Slurm tasks in array 2318595 completed with exit code 0. The audit verified immutable source/config hashes, resolved configurations, runtime source digests and versions, worker counts, training-token and update budgets, complete full validation, finite losses, and absence of checkpoint files. All 80 groups contain exactly runtime seeds 42, 43, and 44; means and sample SDs are recomputed from individual results.

Frozen inputs and run artifacts are under runs/packed8-awc-20m-batch128k-20260912. The source is based on commit d0d0d00 plus the archived checkpoint-policy change. The full source digest is 86648f97bc5f1db1f2815351e7f9b512549cf0d04519528eacb6112d157c7dd1. The results JSON records the frozen input hashes, per-run result/config/environment hashes, and the four-worker reference hash.

Reproduction

The winning recipe is configs/packed8-20m-awc.yaml. With the prepared C4 cache, run:

uv run tiny-llm train --config configs/packed8-20m-awc.yaml

The recipe saves a final checkpoint for subsequent evaluation or analysis. Add --set training.checkpoint_policy=none to reproduce the sweep’s checkpoint-free retention. Repeat with runtime seeds 43 and 44 in separate output directories. The preset explicitly selects eight AWC workers and supplies the measured optimizer settings, architecture, global batch, and training budget. For exact frozen configurations, use the archived config files and source. Checkpoint-free runs retain evaluation metrics but cannot be resumed or used for checkpoint analysis.

Regenerate figures and this report from the audited data:

uv run python doc/data/recipe_sweep_packed8_awc_20m_128k/plot.py
uv run python doc/data/recipe_sweep_packed8_awc_20m_128k/write_report.py
View Markdown source ↗
Search current documentation

Historical reports are available in the archive.