guides
Running checkpoint analysis
Analysis
The current cosine tuning runs retained metrics and results but saved no checkpoints. Their migrated logs cannot supply Hessian-vector products. Train a new run with the default final checkpoint, or retain epoch checkpoints as described below. Start with data preparation and local training.
Unseen analysis starts after the entire final training shuffle group, including
candidate sequences beyond the nominal training budget. Prepare enough additional
training data to cover an unseen epoch and its full final shuffle group; set
data.prepare_train_tokens when preparing a new cache. Preparation rounds this
minimum capacity up to include complete groups and sequence lookahead.
For epoch-by-epoch analysis, train with --set training.checkpoint_policy=interval.
Analyze all checkpoints or select filenames relative to the run directory.
Use a prepared cache with enough data for both seen and unseen cases; the tool
checks capacity before computing. Packed runs use root, averaged checkpoints.
Packed runs also analyze every worker’s consensus error x_i - x_bar, using
the Hessian at the averaged model. Worker-level scalars are saved, and aggregate
alignment and norm curves join the existing figures. Epoch exports require all
matching node-NNN/ weight files; packed .pt states already contain them.
Use --no-consensus with analyze or submit-analysis to disable this diagnostic.
uv run tiny-llm analyze --run runs/20m --checkpoints all \
--output runs/20m/analysis
# Analyze a selection with the measured BF16/HVP batch settings:
uv run tiny-llm analyze --run runs/20m \
--checkpoints epoch-001.safetensors epoch-020.safetensors final.pt \
--amp --hvp-batch-size 64 --output runs/20m/analysis-bf16
Defaults are both data cases, 32 noise samples, 32 random directions, seed 42,
FP32 without AMP, TF32 enabled on CUDA, and compiled Pearlmutter HVPs with
reference attention. Gradient microbatch size always comes from resolved.yaml;
HVP batch size defaults to that size. Useful options:
--data-case seen|unseen|bothselects data cases.--noise-samples N|alland--random-samples Ncontrol sampling.--amp --hvp-batch-size 64enables BF16 forward AMP with FP32 parameters.--device cpu --dtype float64runs eager derivative checks on CPU.--no-tf32,--no-compile-hvp, and--no-plotsdisable those features.
Results include per-sample JSON, a manifest, summary.csv, and three PDF/PNG
figures: normalized alignment, unnormalized noise alignment, and gradient/noise
norms. Seen and unseen curves share axes against training tokens. Repeat the
same command to reuse completed compatible results; use a new output directory
when changing source or numerical settings. Regenerate figures with:
uv run tiny-llm plot-analysis --output runs/20m/analysis
Add --training-norms to include logged training norms. See the
analysis reference for formulas, data selection, and precision,
and measurements for BF16 accuracy and runtime.