Historical evidence
Keep the record.
Follow the current work.
These reports preserve earlier experiments and detailed campaign narratives. Their settings, commands, and implementation status describe the time of measurement.
For maintained conclusions, start with the results overview andcurrent explorer.
- Historical cosine explorer ↗
- Original WSD explorer ↗
- GH200 checkpoint-analysis performance ↗
- C4 campaign recipe selection ↗
- GH200 training performance ↗
- Implementation validation — 2026-09-08 ↗
- Packed training: GH200 validation and benchmarks ↗
- Earlier experiment protocols ↗
- 20M training-recipe benchmark ↗
- Synchronous 20M tuning at 128K tokens per batch ↗
- Packed-4 AWC 20M tuning at 128K tokens per batch ↗
- ATC packed-4 20M tuning at 128K tokens per batch ↗
- AWC beta2 = 0.99 analysis for packed-4 20M at 128K tokens per batch ↗
- Four-worker AWC without gradient clipping: 20M tuning at 128K tokens ↗
- AWC packed-8 20M tuning at 128K tokens per batch ↗
- Earlier results overview ↗
- Earlier training performance ↗
- Benchmarking, profiling, and sweeps ↗
- Adaptive-consensus investigation ↗
- Algorithms and estimators ↗
- Gradient-noise and Hessian analysis ↗
- Training implementation and artifacts ↗
- Training recipe and evidence ↗
- Checkpoint-analysis performance ↗
- Clipping & worker-count comparisons ↗
- Optimizer tuning ↗
- WSD tuning across training horizons ↗