The hero run — the 535-billion-parameter mixture-of-experts model this page tracks — is taking steps at its normal pace again after a stall this morning. It stopped at 9:38 a.m. (ET) today with every GPU busy and no data moving between them: the silent hang the team tracks as #8870, which has now stopped the run three times since it moved to the NCCL 2.30.7 build of the library that carries data between GPUs on September 9. The team has not found the cause.
A watchdog that kills the job when steps stop did so with no one stepping in. A replacement process attached at 9:55 a.m. (ET), restored the checkpoint saved at 9:33 a.m. (ET), and passed the step it hung on by 10:05 a.m. (ET) — twenty-seven minutes from the stall to the first new step, and the few steps after that checkpoint were redone, so no training progress was lost. No one has written the stall up in the status log yet.
The climb in gradient norm that Larry Dial flagged on the night of September 12 (ET) — the size of the push each update makes on the weights, which would warn of instability if it ran away — has stayed under this morning's peak since, including the readings logged after the restart, so it still reads as one spike rather than a new level.