Paper Reports
Plain-language walkthroughs of papers from our group — what problem the paper attacks, how the system works, and what the numbers actually say. Each report links back to the paper itself.
2026 1 report

Cluster '26
LayerCheck: Adaptive Layer-wise Checkpointing for Large Language Model Post-training
Save only the layers that changed, and rebuild the rest at recovery — 22.6× smaller checkpoints without changing what the model learns.
22.6× smaller total checkpoint storage than LowDiff0.54% max loss deviation after recovery, within training noise3.4× faster recovery at matched restart freshness
2025 1 report

SC '25
STELLAR: Storage Tuning Engine Leveraging LLM Autonomous Reasoning for HPC Parallel File Systems
An agentic LLM tuner that reaches near-expert Lustre configurations in five runs instead of thousands.
7.8× peak speedup over default Lustre settings≤ 5 configurations tried per tuning run13 parameters kept, out of 159+ Lustre exposes