2026 1 report
Cluster '26
LayerCheck: Adaptive Layer-wise Checkpointing for Large Language Model Post-training
Minqiu Sun, Xin Huang, Luanzheng Guo, Nathan R. Tallent, Kento Sato, Dong Dai
Save only the layers that changed, and rebuild the rest at recovery — 22.6× smaller checkpoints without changing what the model learns.
22.6× smaller total checkpoint storage than LowDiff0.54% max loss deviation after recovery, within training noise3.4× faster recovery at matched restart freshness
2025 1 report
SC '25
STELLAR: Storage Tuning Engine Leveraging LLM Autonomous Reasoning for HPC Parallel File Systems
Chris Egersdoerfer, Philip Carns, Shane Snyder, Robert Ross, Dong Dai
An agentic LLM tuner that reaches near-expert Lustre configurations in five runs instead of thousands.
7.8× peak speedup over default Lustre settings≤ 5 configurations tried per tuning run13 parameters kept, out of 159+ Lustre exposes