← All paper reports
SC '26

ZOCheck: CPU-Shadow Checkpointing for Zeroth-Order LLM Fine-Tuning

A zeroth-order training step is a seed and a scalar. Log those, let a CPU shadow replay them, and checkpointing costs almost nothing while recovery is bit-for-bit exact.

Minqiu Sun1*Xin Huang2Luanzheng Guo3,4Nathan R. Tallent3Kento Sato2Dong Dai1✉
1 University of Delaware·2 RIKEN Center for Computational Science·3 Pacific Northwest National Laboratory·4 University of Washington
* Ph.D. student mentored  ·  ✉ corresponding author
219.7×
lower checkpoint overhead than asynchronous full-state checkpointing
21.3×
less end-to-end wasted time on Qwen3-8B across MTBFs of 3–15 h
1.55×
faster recovery than async checkpointing every 7 steps; 586× vs. log-only
0
parameter difference after recovery — bitwise identical to the uninterrupted run
ZOCheck overview
Top: naive zeroth-order checkpointing pauses GPU training to write a full checkpoint and recovers it from disk. Bottom: ZOCheck's GPU emits only a seed and a scalar per step; a CPU shadow replays each step to maintain a replica, snapshots it at clean boundaries, and persists snapshots asynchronously. After a failure, training resumes from the freshest CPU-side image and replays only a short suffix of the log.

Key findings

  • A zeroth-order (forward-only) training step is fully determined by a random seed and one scalar, so its log entry is a few tens of bytes regardless of model size. But log-only recovery must replay every step from the start, and shortcut replay that skips MeZO's in-place perturbation sequence drifts from the executed floating-point trajectory.
  • ZOCheck keeps the unmodified training loop on the GPU and runs a CPU shadow process that replays the log — regenerating perturbations from seeds and doing only element-wise arithmetic, no forward passes — to maintain a near-current replica in host memory. The shadow snapshots the replica at clean step boundaries and persists snapshots asynchronously, so no model-sized copy ever touches the GPU critical path. A cost model picks the snapshot interval from measured GPU, replay, snapshot, and persistence times.
  • Across Qwen3, Llama3, and OPT models, ZOCheck cuts checkpoint overhead by up to 219.7× (0.48 s of logging over a whole run), recovers 1.55× faster on average than asynchronous full-state checkpointing at its most aggressive interval and 586× faster than log-only replay, and reduces end-to-end wasted time by 14.0–21.3× on Qwen3-8B. Recovered ZO-SGD runs match uninterrupted training bitwise after one and after three failures. The design depends on the CPU keeping pace: BF16/FP16, small batches, short sequences, or ZO-Adam can push it into a lag regime where recovery slows sharply.

Abstract

Zeroth-order (ZO) optimization is an attractive option for memory-efficient LLM fine-tuning, but its fault tolerance remains underexplored. Unlike first-order training, ZO progress can be represented by lightweight seed-and-scalar step logs, yet naive log-only recovery still incurs replay cost that grows with training progress, and shortcut replay does not preserve the executed floating-point trajectory. We present ZOCheck, a fault-tolerant ZO training system that exploits this replayable structure through a CPU shadow process that continuously replays logged updates, materializes consistent recovery images off the GPU critical path, and persists them asynchronously. ZOCheck therefore combines non-blocking checkpointing during training with fast recovery from a near-current state. We also develop a cost model for choosing the snapshot policy under realistic failure rates. Experiments show that ZOCheck reduces checkpoint overhead by up to 219.7× and recovery latency by 1.55× on average compared with asynchronous full-state checkpointing, translating into up to 21.3× lower end-to-end wasted time across the evaluated failure rates, while preserving exact recovery behavior.

The problem: checkpointing a job whose steps are almost free to describe

Fine-tuning is where a pre-trained language model becomes useful, and as models and runs grow, memory is the first wall a fine-tuning job hits. Zeroth-order (ZO) optimization is one way around it. Instead of backpropagation, a ZO method estimates the gradient from forward passes alone, so a model can be fine-tuned at close to inference-time memory. MeZO showed this can approach first-order quality, and a line of follow-up work has improved convergence, efficiency, and low-precision support since.

Long runs fail, and ZO runs are no exception. During OPT-175B training, hardware failures caused at least 35 manual restarts and an estimated 70 or more automatic ones over two months. During BLOOM training, one GPU crash erased 7.3 hours of work. Post-training is not cheap either: a single OpenThinker3-7B fine-tuning run took 25,000 A100 GPU-hours. The standard defense is checkpoint/restart, and its standard cost is well documented: prior work reports checkpoint overheads averaging 12% of training time and reaching 43%.

A decade of checkpointing systems has attacked that cost with asynchronous persistence, differential checkpoints, multi-level staging, and in-memory replicas. All of them are built for first-order training, where every step changes the whole model and the optimizer state, so the thing being saved is always model-sized. They make the copy cheaper. They cannot make it small.

ZOCheck starts from a property that first-order training does not have: a ZO step can be written down in a few tens of bytes. The paper builds a checkpointing system around that fact, and shows that the obvious ways to use it are wrong.

The observation: a ZO step is a seed and a scalar

A ZO training step, as MeZO runs it, has three stages. First, a pseudorandom seed defines a random perturbation tensor z with the same shape as the model. Second, the same mini-batch is evaluated twice, once with the parameters nudged to θ + εz and once to θ − εz. The difference between the two losses, divided by 2ε, is a scalar that says whether to move along z or against it, and by how much. Third, the optimizer applies that scalar times z as the update.

The middle stage, two full forward passes, dominates the cost of a step on the GPU. But the step's outcome depends only on the seed, the scalar, and the current parameters. Given those, anyone can regenerate z and reapply the update without touching the training data. The paper's log entry per step is exactly that: the seed, the projected-gradient scalar, and the step's learning rate and perturbation scale, a constant number of scalars regardless of model size.

There is a catch, and the paper is careful about it. MeZO never stores z. It perturbs the live parameter buffer in place, adding +εz, then −2εz, then +εz to nominally return to where it started before applying the descent step. In exact arithmetic that sequence cancels. In floating point, each addition is rounded, so the "restored" buffer is not quite the original. The realized state transition is the ordered sequence of in-place operations, and any replay that skips or fuses them lands on a different floating-point trajectory.

Figure 1. What happens if replay takes the shortcut. Starting from the same state in FP16, one replay re-executes MeZO's in-place perturbation sequence before each update; the other applies only the final update. The maximum and mean per-parameter differences both grow with the number of replayed steps, from 0.1172 and 0.0241 at 200 steps to 0.2656 and 0.0491 at 800.
Figure 1. What happens if replay takes the shortcut. Starting from the same state in FP16, one replay re-executes MeZO's in-place perturbation sequence before each update; the other applies only the final update. The maximum and mean per-parameter differences both grow with the number of replayed steps, from 0.1172 and 0.0241 at 200 steps to 0.2656 and 0.0491 at 800.

The figure quantifies the consequence. Over 200 to 800 replayed steps in FP16, direct accumulation of the projected updates drifts steadily away from the state MeZO actually produced. Recovery therefore has to reconstruct the realized parameters, not an idealized update direction, and that constrains how the log can be used.

Why the obvious designs fail

The compact log leaves two natural checkpointing strategies, and the paper's overhead analysis shows that neither is good enough.

Log only. Record the per-step entries and take no full checkpoints. Checkpoint overhead is essentially zero, but after a failure the system must load the initial model and replay every logged step in order. Recovery cost grows linearly with training progress, and even under a uniform failure time it replays about half of the mean-time-between-failures window on average, a floor that does not shrink as hardware gets more reliable.

Log plus anchors. Take a full checkpoint (an anchor) every K steps while still logging, so recovery replays at most about K steps. This is the familiar Young–Daly trade-off with a square-root optimum, and it bounds replay, but each anchor is a blocking, model-sized copy on the GPU's critical path. Overlapping the disk write asynchronously helps, but the device-to-host copy still stalls training.

There is also a ZO-specific reason full checkpoints are awkward. In first-order training, parameters sit unchanged for the whole forward and backward pass, a wide window in which a background thread can copy them. In MeZO the live buffer is perturbed during both forward evaluations. The only clean moment to snapshot is the narrow boundary after the descent update and before the next perturbation begins.

So the design goal is specific: keep the compact log, bound replay after a failure, preserve the exact operation order, and get model-sized copies off the GPU entirely.

How ZOCheck works

Figure 2. ZOCheck compared with periodic full-state ZO checkpointing. Top: a naive design pauses GPU training to write a checkpoint, and recovery reloads it from disk. Bottom: the GPU emits only a seed and a scalar per step; a CPU shadow replays each step to maintain a replica, snapshots it at clean boundaries, and persists snapshots asynchronously. After a failure, training resumes from the freshest CPU-side image and replays only a short suffix of the log.
Figure 2. ZOCheck compared with periodic full-state ZO checkpointing. Top: a naive design pauses GPU training to write a checkpoint, and recovery reloads it from disk. Bottom: the GPU emits only a seed and a scalar per step; a CPU shadow replays each step to maintain a replica, snapshots it at clean boundaries, and persists snapshots asynchronously. After a failure, training resumes from the freshest CPU-side image and replays only a short suffix of the log.

The idea is to split the work between the two processors that a fine-tuning server already has. The GPU runs the unmodified ZO training loop and, after each completed step, publishes one log entry. A separate CPU process, the shadow, consumes the log and re-executes each step's parameter transition on a replica in host memory. Because the replica is on the CPU, it can be snapshotted and written to disk without the GPU noticing. Four stages implement this path.

Step log. After each step the GPU appends the entry (seed, projected gradient, learning rate, perturbation scale) to the update history and pushes it through a shared queue. The step index lets the shadow apply entries in order even when perturbation generation on the GPU is pipelined.

Replay. The shadow regenerates z from the seed and applies the same in-place sequence the GPU did: +εz, −2εz, +εz, then the descent update. It does not collapse the perturbations into a no-op, for the reason Figure 1 shows. What it also does not do is run a forward pass. CPU replay consists of random-tensor generation and element-wise arithmetic over the parameters, so its cost depends on the model's size and nothing else: not batch size, not sequence length, not the data. That is what makes it plausible for a CPU to keep up with a GPU.

Snapshot. Mid-replay, the replica is in a perturbed, half-updated state, just as the GPU's buffer is mid-step. Every N replayed steps the shadow pauses at a clean boundary and copies the replica into a second host-memory buffer. A valid snapshot carries parameters, any optimizer state, and the index of the last fully replayed entry, so that resume starts from exactly that boundary and replays a strict suffix of the log. Because the snapshot lives in a separate process's memory, it survives a crash of the training process and can be loaded straight back without going to disk.

Persistence. A background thread flushes the frozen snapshot buffer to durable storage, with at most one write in flight at a time so the newest durable image is always unambiguous after a crash. This is what protects against a host failure, and it is what bounds replay distance: recovery loads the latest durable snapshot and replays only what came after it.

The two-buffer arrangement is why the shadow needs two copies of the state in host DRAM. The live replica keeps replaying while the frozen snapshot is being written; otherwise the asynchronous write could race with replay and persist a mixture of old and new parameters. For ZO-SGD, where the state is just the parameters, that is 2× model size in host memory. For ZO-Adam, which adds two moment tensors per parameter, it rises to 6×.

The design keeps two invariants. The GPU path is unchanged: no extra replica, no replay worker, and no model-sized copy on the critical path, so ZOCheck adds no GPU memory. And recovery images come only from the shadow at clean post-update boundaries, which is what makes the resume point an exact floating-point boundary rather than an approximation of one.

Keeping the CPU in step with the GPU

Everything above rests on the shadow staying close to the GPU. If replay is slower than training, lag accumulates, and recovery has to replay a longer suffix. The paper treats replay throughput as a first-class systems target.

Replay has two stages that can be pipelined: a producer that generates the perturbation tensor from the seed, and a consumer that applies the ordered in-place arithmetic to the replica. The arithmetic order within a tensor partition cannot be changed, but partitions are independent, so generation and update overlap across partitions and each partition is handed to the consumer as soon as it is ready.

The remaining knob is how to split P CPU threads between the two stages. Replay time is set by the slower stage, so ZOCheck profiles a range of splits at startup and picks the minimum with a linear scan.

Figure 3. Profiling the producer/consumer thread split on Qwen3-8B. Each point is the measured CPU replay time for one allocation of consumer threads; the star marks the selected split of 68 generator and 59 consumer threads, at 6.586 s per step. The dashed line is the GPU step time of 8.650 s.
Figure 3. Profiling the producer/consumer thread split on Qwen3-8B. Each point is the measured CPU replay time for one allocation of consumer threads; the star marks the selected split of 68 generator and 59 consumer threads, at 6.586 s per step. The dashed line is the GPU step time of 8.650 s.

The scan on Qwen3-8B shows a broad flat minimum below the GPU budget, so the choice is not brittle to a single integer. Across the models in the main evaluation, replay stays under the GPU step time:

Model CPU replay step (s) GPU training step (s) Keeps pace Snapshot interval
Qwen3-0.6B 0.62 0.63 yes 23
Qwen3-1.7B 1.48 1.61 yes 5
Qwen3-4B 2.86 4.21 yes 5
Qwen3-8B 6.59 8.65 yes 3
OPT-6.7B 3.43 6.31 yes 5
Llama3.2-3B 0.97 3.66 yes 9
Llama3.1-8B 2.48 7.87 yes 9

The headroom is largest on the bigger models, because GPU time grows with two forward passes while CPU replay grows only with parameter count. The narrowest margin is Qwen3-0.6B, where replay and training are within a few milliseconds of each other. ZOCheck remains feasible there, but it has to choose a much longer snapshot interval to stay out of the lag regime.

Choosing the snapshot interval

The one policy decision left is N, how many replayed steps go between snapshots. Snapshot too often and the copy time slows replay; too rarely and recovery has more to replay. The paper's cost model resolves this with two hardware constraints and one objective.

Pacing. The shadow keeps pace only if one replay step plus its amortized share of a snapshot fits inside one GPU step. That puts a lower bound on N: the snapshot cost divided by the slack between GPU and CPU step time.

Persistence. Each asynchronous write must finish before the next snapshot is produced, or snapshots pile up and eventually stall the shadow. That is a second lower bound: the durable-write time divided by the replay time.

The objective is expected wasted time per failure window, which combines the per-step logging cost with the expected recovery cost, and recovery cost depends on the expected replay distance. That distance has two regimes. When the shadow keeps pace, it is about 1 + N/2 steps. When it does not, it also includes the lag accumulated over the failure window, a term that grows with mean time between failures.

The optimal N follows. In the keep-pace regime it is simply the smallest interval that satisfies both constraints, and it does not depend on the failure rate at all, which makes the configuration robust to a wrong MTBF estimate. In the lag regime, N balances the snapshot term against residual replay and recovers the same square-root scaling as the classical anchor interval. The workloads in the paper are almost all in the first regime. As a fallback when the CPU cannot keep up, ZOCheck can periodically resynchronize the shadow with one blocking device-to-host copy, capping lag at the resync interval.

In deployment, all of this reduces to a short startup profile: measure the GPU step time, profile CPU replay to find the best thread split, measure snapshot and persistence cost, and compute N. The profile completes within tens of steps.

How fast is recovery?

The experiments run on two servers: an NVIDIA L40S (48 GB) with an AMD EPYC 9554 (128 hardware threads, 1.5 TB of host memory), and an NVIDIA A100 (40 GB) with an EPYC 7763 (256 threads, 2 TB). Snapshots persist to SSD. Models are Qwen3 at 1.7B, 4B, and 8B, Llama-3.2-3B, Llama-3.1-8B, and OPT-6.7B, plus Qwen3-0.6B as a pacing microbenchmark. The default configuration is full-parameter FP32 fine-tuning on SST-2 with batch size 16 and sequence length 512, using ZO-SGD at a learning rate of 10⁻⁷.

The baselines are the two designs from the overhead analysis. MeZO (log only) logs every step and replays from the initial model. Log + anchor writes a full checkpoint synchronously every K steps. Log + asynchronous anchor blocks training only for the device-to-host copy and persists to SSD in the background. Both anchor baselines are run at K = 7, 300, and 1000. Each recovery experiment trains for 20,000 steps and injects one failure at step 16,000, in one of two modes: a software failure that kills the training process but leaves the shadow alive, and a host failure in which only the latest durable snapshot survives.

Figure 4. Recovery time after a single injected failure, on a log scale. "From CPU" is software recovery from the near-current shadow; "From Disk" is recovery from the latest durable snapshot. Synchronous and asynchronous anchors at the same K have identical recovery time, since the write strategy does not change the recovery point.
Figure 4. Recovery time after a single injected failure, on a log scale. "From CPU" is software recovery from the near-current shadow; "From Disk" is recovery from the latest durable snapshot. Synchronous and asynchronous anchors at the same K have identical recovery time, since the write strategy does not change the recovery point.

ZOCheck recovers fastest on every model and in both failure modes. Averaged across them, it is 1.55× faster than asynchronous anchoring at the most aggressive interval tested (K = 7), and 586× faster than log-only replay. The comparison against K = 7 is the harsh one: that baseline checkpoints every seven steps to keep its recovery point fresh, and pays for it during training, which is the next result.

What it costs during training

Figure 5. Cumulative checkpoint overhead over a failure-free run, on a log scale. Anchors at K = 7 cost the most because each one stalls the GPU; asynchronous persistence reduces but does not remove the stall. ZOCheck's cost is the logging alone.
Figure 5. Cumulative checkpoint overhead over a failure-free run, on a log scale. Anchors at K = 7 cost the most because each one stalls the GPU; asynchronous persistence reduces but does not remove the stall. ZOCheck's cost is the logging alone.

Over a failure-free run, ZOCheck adds 0.48 s of cumulative logging overhead in total. Every anchor baseline is visibly above it, because each full checkpoint blocks the GPU for at least the device-to-host copy. Against asynchronous anchoring at its cheapest setting (K = 1000), ZOCheck reduces checkpoint overhead by up to 219.7×.

That pairing is the paper's central point. The anchor baselines trade the two quantities against each other: fresh recovery points cost training time, and cheap training costs replay after a failure. ZOCheck sits at the fast-recovery end of one axis and the low-overhead end of the other, because the work that produces a fresh recovery point happens on a processor that was otherwise idle.

End to end

Figure 6. End-to-end wasted time on Qwen3-8B, with failures at a mean time between failures of 3, 9, and 15 hours. Each bar stacks checkpoint stalls, load time, and replay. The boxed numbers are ZOCheck's totals in minutes: from disk and from CPU memory.
Figure 6. End-to-end wasted time on Qwen3-8B, with failures at a mean time between failures of 3, 9, and 15 hours. Each bar stacks checkpoint stalls, load time, and replay. The boxed numbers are ZOCheck's totals in minutes: from disk and from CPU memory.

Putting the two together on Qwen3-8B, across mean-time-between-failures values of 3 to 15 hours, ZOCheck reduces end-to-end wasted time by 14.0–21.3× relative to the lowest-waste asynchronous anchor baseline at each failure rate. After a host failure, where recovery has to start from disk, the reduction is still 6.5–10.0× over the same range. ZOCheck's own wasted time is under half a minute at every failure rate in the figure.

Is recovery exact?

Speed would be beside the point if the resumed run were a different run. The paper checks this two ways.

Figure 7. Training loss over a 1,000-step window with three injected failures and recoveries at steps 200, 400, and 600, for ZO-SGD and for Sparse-MeZO, each against its uninterrupted run. Recovered and uninterrupted curves lie on top of each other; the inset zooms on steps 1000–1060.
Figure 7. Training loss over a 1,000-step window with three injected failures and recoveries at steps 200, 400, and 600, for ZO-SGD and for Sparse-MeZO, each against its uninterrupted run. Recovered and uninterrupted curves lie on top of each other; the inset zooms on steps 1000–1060.

The loss curves for the three-failure runs coincide with their uninterrupted counterparts through every recovery boundary. But overlapping loss curves are necessary, not sufficient: two nearby trajectories can look identical on a loss plot. So the paper hashes the full parameter tensor after every step and compares recovered runs against uninterrupted references. For ZO-SGD, both the single-failure and the triple-failure run match the reference bitwise over the entire run. The maximum element-wise parameter difference is exactly zero, and so is the loss difference. This is the payoff of replaying the in-place sequence faithfully: the recovered state is the same floating-point state, not a close one.

The same experiment on Sparse-MeZO, a variant that perturbs and updates only a magnitude-selected subset of parameters each step, also recovers onto its uninterrupted trajectory. The architecture applies to any ZO optimizer whose step is determined by a seed, a scalar, and the current optimizer state, and the paper evaluates it on ZO-SGD, ZO-Adam, and Sparse-MeZO.

Where the design stops working

The paper is direct about the boundary. ZOCheck's near-instant recovery holds while the CPU keeps pace, and the sensitivity study maps where that stops. All of these runs are on Qwen3-8B unless noted.

Faster GPUs. Lower precision speeds up both sides, but not equally:

Setting CPU replay step (s) GPU training step (s) Recovery time (s)
FP32, L40S + EPYC 9554 (default) 6.59 8.65 13.56
BF16 3.68 2.85 1364.08
FP16 3.62 2.78 1313.07
FP32, A100 + EPYC 7763 2.53 9.99 13.18

Under BF16 and FP16 the GPU step drops by about 3× while CPU replay drops by less than 2×, replay falls behind, and recovery time jumps from 14 seconds to over 20 minutes even though training itself got faster. The A100 server moves the other way: replay speeds up more than training, so the margin widens and recovery stays at 13 seconds.

Smaller batches, shorter sequences. These shrink the GPU step but leave replay untouched, which is exactly the pacing model's prediction:

Setting GPU training step (s) Keeps pace Recovery time (s)
Batch 16, sequence 512 (default) 8.65 yes 13.56
Batch 8 4.58 no 2025.91
Batch 32 16.57 yes 11.56
Sequence 256 4.52 no 1988.39
Sequence 1024 16.94 yes 11.56

Halving the batch or the sequence length pushes the system into the lag regime and recovery climbs past 30 minutes. Doubling either preserves pacing. Across six datasets (SQuAD, DROP, BoolQ, WSC, WiC, MultiRC) at the default batch and sequence length, GPU step time stays within 8.41–8.81 s and the shadow keeps pace on all of them.

Heavier optimizers. ZO-Adam adds moment-tensor arithmetic to every replayed step. On Qwen3-1.7B (the 8B model does not fit with Adam in this setup), replay goes from 1.48 s to 3.41 s per step while the GPU step barely changes (1.61 s to 1.70 s), and recovery time rises from 8.41 s to 241.07 s.

In every one of these cases the fallback still applies: periodic resynchronization bounds the lag, and any replay progress the shadow makes still reduces recovery work relative to log-only replay.

What it consumes

ZOCheck moves checkpointing state off the GPU and onto host resources. On Qwen3-8B with the default configuration:

Strategy Peak VRAM (GB) Peak DRAM (GB) CPU peak utilization (%)
MeZO (log only) 31.616 36.854 187.8
Log + anchor 31.794 57.272 205.7
ZOCheck 31.94 62.669 12,772.6

GPU memory is essentially unchanged. Host memory rises for the replica and the snapshot buffer. CPU utilization is the number to notice: 12,772% is 127.72 cores' worth, and it is the whole point. The design spends abundant, otherwise idle CPU parallelism to take checkpoint work off the processor that is actually the bottleneck.

Limits, and what comes next

The evaluation is single-node and feasibility-focused, and the authors list what it does not cover. Distributed ZO training is future work. So are adaptive ZO optimizers whose updates cannot be reconstructed from lightweight metadata (the paper names DeepZero-style coordinate-wise schemes as needing case-specific extensions), and runtime policies that co-optimize replay lag, snapshot frequency, and storage placement across memory tiers. Several system-level issues are named but not evaluated: NUMA-aware placement and thread pinning, contention on the persistence path, crash consistency for the asynchronous log, and floating-point reproducibility across instruction-set architectures. Finally, specialized low-precision CPU replay kernels could close the BF16/FP16 gap that the sensitivity study exposed.

Why this result matters

The reusable idea is that fault tolerance should follow the optimizer's execution semantics rather than a generic notion of "training state." First-order checkpointing systems have spent years making a model-sized copy cheaper because, for first-order training, there is nothing smaller to save. Zeroth-order training breaks that assumption: a step is a seed and a scalar, and the model-sized work of turning that into a recovery image can be done by a different processor at its own pace.

Two details separate this from a trick. Replay must reproduce the executed floating-point path, not the idealized update, or the recovered run is a different run. And the CPU must actually keep up, which is a pacing condition the paper models, measures, and shows both sides of. Within that condition, ZOCheck turns checkpointing for ZO fine-tuning from a trade-off between overhead and recovery into something close to free on both counts, with a recovered state that is bit-for-bit the one training would have reached.

Cite this paper

Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (SC 2026), Chicago, IL, USA, November 15–20, 2026

@inproceedings{sun2026zocheck,
  author    = {Minqiu Sun and Xin Huang and Luanzheng Guo and Nathan R.
               Tallent and Kento Sato and Dong Dai},
  title     = {{ZOCheck:} {CPU}-Shadow Checkpointing for Zeroth-Order {LLM}
               Fine-Tuning},
  booktitle = {Proceedings of the International Conference for High
               Performance Computing, Networking, Storage and Analysis, {SC}
               2026, Chicago, IL, USA, November 15-20, 2026},
  publisher = {{ACM}},
  year      = {2026},
  note      = {To appear}
}

This report summarizes the paper for a general systems audience. All numbers and quotes are taken from the paper itself; figures are reproduced from it. For the full method and evaluation, read the original.