Marin 535B-A23B MoE

18T tokens · launched Aug 20 updated · auto-generated via LLM, using mumwelt and data from GitHub, Discord, and WandB
26.4% · step 103,037 of 390,251
Aug 20 4.75T of 18T tokens Nov 27

The hero run — the 535-billion-parameter mixture-of-experts model this page tracks — is taking steps at its normal pace again after a stall this morning. It stopped at 9:38 a.m. (ET) today with every GPU busy and no data moving between them: the silent hang the team tracks as #8870, which has now stopped the run three times since it moved to the NCCL 2.30.7 build of the library that carries data between GPUs on September 9. The team has not found the cause.

A watchdog that kills the job when steps stop did so with no one stepping in. A replacement process attached at 9:55 a.m. (ET), restored the checkpoint saved at 9:33 a.m. (ET), and passed the step it hung on by 10:05 a.m. (ET) — twenty-seven minutes from the stall to the first new step, and the few steps after that checkpoint were redone, so no training progress was lost. No one has written the stall up in the status log yet.

The climb in gradient norm that Larry Dial flagged on the night of September 12 (ET) — the size of the push each update makes on the weights, which would warn of instability if it ran away — has stayed under this morning's peak since, including the readings logged after the restart, so it still reads as one spike rather than a new level.

Sep 14
hero-ragged_a2a-nccl2307-ep-step81k
herohangrestart
hero-ragged_a2a-nccl2307-ep-step81k
herograd-norm
hero-ragged_a2a-nccl2307-ep-step81k
herograd-norm
hero-ragged_a2a-nccl2307-ep-step81k
herograd-norm
hero-ragged_a2a-nccl2307-ep-step81k
herograd-norm
hero-ragged_a2a-nccl2307-ep-step81k
herograd-norm
hero-ragged_a2a-nccl2307-ep-step81k
herograd-norm
AUG-LIN0-2L-DENSE-MUONH-X600-HIGH-2P2-d512-600x-lr2.2
lr-sweep
hero-ragged_a2a-nccl2307-ep-step81k
herograd-norm
AUG-LIN0-2L-DENSE-MUONH-X600-HIGH-2P2-d512-600x-lr2.2
lr-sweep
hero-ragged_a2a-nccl2307-ep-step81k
herograd-norm
AUG-LIN0-2L-DENSE-MUONH-X600-HIGH-2P2-d512-600x-lr2.2
lr-sweep
hero-ragged_a2a-nccl2307-ep-step81k
herograd-norm
AUG-LIN0-2L-DENSE-MUONH-X600-HIGH-2P2-d512-600x-lr2.2
lr-sweep
hero-ragged_a2a-nccl2307-ep-step81k
herograd-norm
AUG-LIN0-2L-DENSE-MUONH-X600-HIGH-2P2-d512-600x-lr2.2
lr-sweep
hero-ragged_a2a-nccl2307-ep-step81k
herograd-norm
AUG-LIN0-2L-DENSE-MUONH-X600-HIGH-2P2-d512-600x-lr2.2
lr-sweeprestart
hero-ragged_a2a-nccl2307-ep-step81k
herograd-normeval
AUG-LIN0-2L-DENSE-MUONH-X600-HIGH-2P2-d512-600x-lr2.2
lr-sweep
hero-ragged_a2a-nccl2307-ep-step81k
herograd-norm
lr-sweep
hero-ragged_a2a-nccl2307-ep-step81k
[The gradient-norm rise Larry Dial flagged](https://discord.com/channels/1354881461060243556/1548520475397718068/1548520790742274108) has fallen well back from the record it set an hour ago. The open 250-step block, steps 101,750-101,999, holds a median of 0.556654 on the output projection at seventeen of its twenty-five readings — fifteenth of the run's eighty closed blocks measured at the same reading count, against the 0.597616 the record block 101,500-101,749 held there and the 0.587258 of 100,750-100,999. All twelve closed blocks that read within 0.010 of it at seventeen readings went on to close between 0.538465 and 0.577332, none of them above the 0.597565 that steps 100,750-100,999 set on Sunday evening, so the record block reads as a spike rather than a new level. Train loss has not moved: the 500-step block 101,500-101,999 reads 1.21259 across 203 readings, inside the 1.2111-1.2164 band the last seven closed blocks held. The hero cleared its thirty-first hourly save since Saturday's restart at 08:01:06Z, a single step of 20.7 seconds against an 18.2-second median, and reached step 101,920 at 08:05:07Z with no step slower than the 23.0 seconds of the 3:01 a.m. (ET) save since step 101,500. Its next reading on held-out text is due at step 101,999 near 4:30 a.m. (ET), the first since the peak. Note that W&B's sampled history, which feeds this page's progress bar and chart, has been pinned at step 101,734 for three reads while the run itself is at 101,920 — the counter trails the run by about 190 steps.
herograd-norm
lr-sweep
gradnorm-101316-pooled-20260914
A third reading of the gradient norm on the frozen step-101,316 checkpoint clears the ragged all-to-all backend of causing [the rise Larry Dial flagged](https://discord.com/channels/1354881461060243556/1548520475397718068/1548520790742274108). `gradnorm-101316-pooled-20260914` finished at 07:12:07Z on the pooled all-to-all path and returned 0.578236 on the output projection, against 0.584267 and 0.584210 from [the two ragged readings](https://wandb.ai/marin-community/marin_moe/runs/gradnorm-101316-ragged-v2-20260914) — a 1.0% spread, about a fifth of the 0.03 the run's block medians have climbed. `grad/norm/total` agrees to 0.5%; the two backends do diverge by several times on the `sconv_*` tensors, but those sit at 1e-4 and are dominated by rounding. The hero's own open 250-step block, steps 101,750-101,999, has meanwhile come off the record: at ten of its twenty-five readings it holds a median of 0.566696 — twelfth of the eighty closed blocks measured at ten readings, against 0.606334 for the record block 101,500-101,749 and 0.605051 for 100,750-100,999. The twelve closed blocks that read within 0.010 of it at ten readings all closed between 0.5385 and 0.5798, below the 0.602308 record. Train loss has not moved: the 500-step block 101,500-101,999 reads 1.21336 across 168 readings, inside the 1.2112-1.2164 band the last seven closed blocks held. The hero reached step 101,848 at 07:43:46Z at a median of 18.46 seconds a step, with no gap over 26.4 seconds in four hours, and [its next reading on held-out text](https://wandb.ai/marin-community/marin_moe/runs/hero-ragged_a2a-nccl2307-ep-step81k) is due at step 101,999 near 4:30 a.m. (ET).
herograd-norm
lr-sweep
hero-ragged_a2a-nccl2307-ep-step81k
herograd-norm
lr-sweep
lr-sweep
hero-ragged_a2a-nccl2307-ep-step81k
herograd-norm
lr-sweep
herograd-norm
herograd-norm
herograd-norm
lr-sweep
herograd-norm
lr-sweep
hero-ragged_a2a-nccl2307-ep-step81k
herotraining-dynamicsgradient-norm
lr-sweep
lr-sweep
hero-ragged_a2a-nccl2307-ep-step81k
herotraining-dynamicsgradient-norm
hero-ragged_a2a-nccl2307-ep-step81k
herotraining-dynamicsgradient-norm
lr-sweep
lr-sweep
hero-ragged_a2a-nccl2307-ep-step81k
herotraining-dynamicsgradient-norm
hero-ragged_a2a-nccl2307-ep-step81k
herocheckpoint
hero-ragged_a2a-nccl2307-ep-step81k
herotraining-dynamicsgradient-norm
hero-ragged_a2a-nccl2307-ep-step81k
herotraining-dynamicsgradient-norm
hero-ragged_a2a-nccl2307-ep-step81k
herocheckpoint
hero-ragged_a2a-nccl2307-ep-step81k
[The gradient-norm rise Larry flagged](https://discord.com/channels/1354881461060243556/1548520475397718068/1548520790742274108) closed a third block running without a new high, and the block after it has opened at the highest half-filled reading the run has produced. The 250-step block covering steps 100,500-100,749 closed at 10:04 p.m. (ET) at a median of 0.5649 on the output projection — ninth of the run's seventy-six closed blocks, under the 0.5798 run high set over steps 99,750-99,999, and close to the 0.5620 this page projected for it at twenty-four samples. The block covering steps 100,750-100,999 is twelve of its twenty-five samples in at 0.6051, above that run high: no closed block in the run has read 0.60 at twelve samples, and the nearest, 0.5884, belongs to the block that just closed at 0.5649. Its samples swing from 0.3467 to 0.7990, and the three highest eleven-sample readings the run has given — 0.6047, 0.6044 and 0.6013 — closed at 0.5591, 0.5649 and 0.5738, so a high half-block has not once carried through to a high close. Thirteen samples are still out and the block closes near 11:23 p.m. (ET). Neither the training loss nor [the run's reading on held-out text](https://wandb.ai/marin-community/marin_moe/runs/hero-ragged_a2a-nccl2307-ep-step81k) has turned: the 500-step block covering steps 100,500-100,999 reads 1.2159 at 191 samples, inside the 1.2113-1.2182 band the last twelve closed blocks held, and the held-out reading is unchanged at step 98,999 with the next due at step 101,999 near 4:28 a.m. (ET). [The hero](https://wandb.ai/marin-community/marin_moe/runs/hero-ragged_a2a-nccl2307-ep-step81k) has stepped without a break for more than twenty-five hours, and its twenty-sixth save since the restart falls near 10:59 p.m. (ET).
herotraining-dynamicsgradient-norm
hero-ragged_a2a-nccl2307-ep-step81k
herocheckpointtraining-dynamicsgradient-norm
The two-layer wave of [the learning-rate sweep](https://github.com/marin-community/marin/issues/7856) is complete, all twenty of its rungs finished, and the highest of its four rates took the fifth and last training length. [The 1.4 rate closed its 21,150 steps at 10:03 p.m. (ET) at 4.0032 on held-out text](https://wandb.ai/marin-community/marin_moe/runs/AUG-LIN0-2L-DENSE-MUONH-019-d512-600x-lr1.4), 0.0010 behind [the 1.7 rate's 4.0022](https://wandb.ai/marin-community/marin_moe/runs/AUG-LIN0-2L-DENSE-MUONH-020-d512-600x-lr1.7), after leading at step 20,000 and reading 4.0038 at step 21,000. It gained 0.0137 over its closing 1,149 steps against the 0.0147 a win needed — close to the 0.0141 that the three earlier finishers' wind-down trend projected, and the shortfall this page called. Across the wave 1.4 wins the four shorter lengths, at 4.5915, 4.3191, 4.1208 and 4.0490, and 1.7 the longest, and 1.4's margin over 1.7 shrank at every step up in length — 0.0099, 0.0026, 0.0013, 0.0015 — before turning over to 1.7's 0.0010, so the best rate drifts toward the top of the grid as the training length grows. A winner at the top of the tested grid is the condition that opens a rung above the range. [The six-layer wave closed at the top of its own grid too](https://wandb.ai/marin-community/marin_moe/runs/AUG-LIN0-6L-DENSE-MUONH-020-d512-600x-lr1.4), one rate lower, and [the probe it drew above that range](https://wandb.ai/marin-community/marin_moe/runs/AUG-LIN0-6L-DENSE-MUONH-X600-HIGH-1P7-d512-600x-lr1.7) lost at all three lengths it ran, which left the grid standing. No two-layer probe has appeared as of 10:06 p.m. (ET).
lr-sweep
The two-layer wave of [the learning-rate sweep](https://github.com/marin-community/marin/issues/7856) has its stalled rung back and a new leader at the newest matched step. [The 1.0 rung](https://wandb.ai/marin-community/marin_moe/runs/AUG-LIN0-2L-DENSE-MUONH-018-d512-600x-lr1) resumed at 9:09 p.m. (ET) after 58 minutes without a logged step at 19,879, took its step-20,000 reading at 4.0247 on held-out text, and has about 600 of its 21,150 steps left, closing near 9:36 p.m. (ET). It sits second at that step, behind [1.7's 4.0189](https://wandb.ai/marin-community/marin_moe/runs/AUG-LIN0-2L-DENSE-MUONH-020-d512-600x-lr1.7) and ahead of [0.7's 4.0326](https://wandb.ai/marin-community/marin_moe/runs/AUG-LIN0-2L-DENSE-MUONH-017-d512-600x-lr0.7), and needs 0.0225 over its last 1,149 steps to take the length from [1.7's finishing 4.0022](https://wandb.ai/marin-community/marin_moe/runs/AUG-LIN0-2L-DENSE-MUONH-020-d512-600x-lr1.7) — more than the 0.0167 that 1.7 found over the same closing stretch and the 0.0065 that 0.7 found, so second is the likelier place. At step 18,000, the newest step all four rungs have reached, [the 1.4 rate that won the four shorter lengths](https://wandb.ai/marin-community/marin_moe/runs/AUG-LIN0-2L-DENSE-MUONH-019-d512-600x-lr1.4) now leads at 4.0431, ahead of 1.0's 4.0448, 1.7's 4.0494 and 0.7's 4.0502, after reading third of the four at step 17,000. It has about 2,150 steps left and closes near 10:05 p.m. (ET), and needs 0.0409 from step 18,000 to win the length — between the 0.0472 that 1.7 gained from there and the 0.0241 that 0.7 gained, so the length still cannot be called.
lr-sweep
hero-ragged_a2a-nccl2307-ep-step81k
herotraining-dynamicsgradient-norm
hero-ragged_a2a-nccl2307-ep-step81k
heromilestonecheckpoint
The two-layer wave of [the learning-rate sweep](https://github.com/marin-community/marin/issues/7856) has a second finisher at its fifth and last training length, and the order at matched steps has inverted again. [The 0.7 rate closed its 21,150 steps at 8:52 p.m. (ET) at 4.0261 on held-out text](https://wandb.ai/marin-community/marin_moe/runs/AUG-LIN0-2L-DENSE-MUONH-017-d512-600x-lr0.7), 0.0239 behind [the 1.7 rate that closed at 8:34 p.m. (ET) at 4.0022](https://wandb.ai/marin-community/marin_moe/runs/AUG-LIN0-2L-DENSE-MUONH-020-d512-600x-lr1.7) — and yet 0.7 led 1.7 at every step the two can be compared at, by 0.0179 at step 17,000. The whole gap turned over in the schedule's wind-down to zero learning rate, where 1.7 gained 0.0849 from step 17,000 against 0.7's 0.0431. [The 1.4 rate that won the four shorter lengths](https://wandb.ai/marin-community/marin_moe/runs/AUG-LIN0-2L-DENSE-MUONH-019-d512-600x-lr1.4) reads 4.0732 at step 17,000, so it needs 0.0710 from there to take this one — more than 0.7 found in its wind-down, less than 1.7 found in its own. It has about 3,200 steps left and closes near 10:00 p.m. (ET), and the length cannot be called before then. [The 1.0 rung](https://wandb.ai/marin-community/marin_moe/runs/AUG-LIN0-2L-DENSE-MUONH-018-d512-600x-lr1) has logged no step since 8:11 p.m. (ET), a stall approaching an hour with a live heartbeat and 1,270 steps left, about twenty minutes of work. If 1.7 holds the length it wins at the top of the tested grid, the case that earned [the six-layer wave a probe above its range](https://wandb.ai/marin-community/marin_moe/runs/AUG-LIN0-6L-DENSE-MUONH-X600-HIGH-1P7-d512-600x-lr1.7).
lr-sweep
hero-ragged_a2a-nccl2307-ep-step81k
herotraining-dynamicsgradient-normcheckpoint
The two-layer wave of [the learning-rate sweep](https://github.com/marin-community/marin/issues/7856) has its first finisher at the fifth and last training length, and it is the highest of the four rates. [The 1.7 rung closed its 21,150 steps at 8:35 p.m. (ET) at 4.0022 on held-out text](https://wandb.ai/marin-community/marin_moe/runs/AUG-LIN0-2L-DENSE-MUONH-020-d512-600x-lr1.7) after reading last of the four at every matched step up to 16,000 — 4.1031 there, against [0.7's 4.0814](https://wandb.ai/marin-community/marin_moe/runs/AUG-LIN0-2L-DENSE-MUONH-017-d512-600x-lr0.7), [1.0's 4.0851](https://wandb.ai/marin-community/marin_moe/runs/AUG-LIN0-2L-DENSE-MUONH-018-d512-600x-lr1) and [1.4's 4.0933](https://wandb.ai/marin-community/marin_moe/runs/AUG-LIN0-2L-DENSE-MUONH-019-d512-600x-lr1.4). It made all of that back in the schedule's wind-down to zero learning rate, gaining 0.1009 from step 16,000 to the finish where 0.7 has gained 0.0551 so far. [The 1.4 rate that won the four shorter lengths](https://wandb.ai/marin-community/marin_moe/runs/AUG-LIN0-2L-DENSE-MUONH-019-d512-600x-lr1.4) is 0.0098 ahead of 1.7 at step 16,000 and so must gain 0.0911 from there to take this one; it is about 4,300 steps back and closes near 10:02 p.m. (ET), so the length cannot be called before then. Neither rate in between is stepping: [1.0 has logged nothing since 8:11 p.m. (ET)](https://wandb.ai/marin-community/marin_moe/runs/AUG-LIN0-2L-DENSE-MUONH-018-d512-600x-lr1), thirty-three minutes with 1,270 steps left, and [0.7 nothing since 8:32 p.m. (ET)](https://wandb.ai/marin-community/marin_moe/runs/AUG-LIN0-2L-DENSE-MUONH-017-d512-600x-lr0.7) with 74 steps left — a pause many times the eighty seconds those steps take. Both keep a live heartbeat, the same shape as the eleven stalls this wave has taken since it launched at 1:55 p.m. (ET).
lr-sweep
lr-sweep
hero-ragged_a2a-nccl2307-ep-step81k
herotraining-dynamicsgradient-norm
hero-ragged_a2a-nccl2307-ep-step81k
herocheckpoint
Sep 13
lr-sweep
hero-ragged_a2a-nccl2307-ep-step81k
herotraining-dynamicsgradient-norm
hero-ragged_a2a-nccl2307-ep-step81k
herotraining-dynamicsgradient-norm
lr-sweep
lr-sweep
hero-ragged_a2a-nccl2307-ep-step81k
herocheckpoint
lr-sweep
hero-ragged_a2a-nccl2307-ep-step81k
herotraining-dynamicsgradient-norm
lr-sweep
hero-ragged_a2a-nccl2307-ep-step81k
herotraining-dynamicsgradient-norm
hero-ragged_a2a-nccl2307-ep-step81k
herocheckpoint
lr-sweep
lr-sweep
hero-ragged_a2a-nccl2307-ep-step81k
herotraining-dynamicsgradient-norm
lr-sweep
lr-sweep
lr-sweep
lr-sweep
hero-ragged_a2a-nccl2307-ep-step81k
herotraining-dynamicsgradient-norm
lr-sweep
lr-sweep
hero-ragged_a2a-nccl2307-ep-step81k
herocheckpointmilestone
hero-ragged_a2a-nccl2307-ep-step81k
herotraining-dynamicsgradient-norm
lr-sweep
lr-sweep
lr-sweep
lr-sweep
lr-sweep
lr-sweep
hero-ragged_a2a-nccl2307-ep-step81k
herocheckpointmilestone
hero-ragged_a2a-nccl2307-ep-step81k
herotraining-dynamicsgradient-norm
hero-ragged_a2a-nccl2307-ep-step81k
herotraining-dynamicsgradient-norm
hero-ragged_a2a-nccl2307-ep-step81k
herocheckpointmilestone
hero-ragged_a2a-nccl2307-ep-step81k
herotraining-dynamicsgradient-norm
lr-sweep
lr-sweep
hero-ragged_a2a-nccl2307-ep-step81k
herotraining-dynamicsgradient-norm
hero-ragged_a2a-nccl2307-ep-step81k
heroevaltraining-dynamicsmilestone
lr-sweep
hero-ragged_a2a-nccl2307-ep-step81k
herocheckpointmilestone
hero-ragged_a2a-nccl2307-ep-step81k
herotraining-dynamicsgradient-norm
lr-sweep
hero-ragged_a2a-nccl2307-ep-step81k
herotraining-dynamicsgradient-norm
lr-sweep
hero-ragged_a2a-nccl2307-ep-step81k
herotraining-dynamicsgradient-norm
hero-ragged_a2a-nccl2307-ep-step81k
herocheckpointmilestone
lr-sweep
hero-ragged_a2a-nccl2307-ep-step81k
herotraining-dynamicsgradient-norm
lr-sweep
hero-ragged_a2a-nccl2307-ep-step81k
herocheckpointmilestone
hero-ragged_a2a-nccl2307-ep-step81k
herotraining-dynamicsgradient-norm
mcwitt hero-ragged_a2a-nccl2307-ep-step81k
herocheckpointmilestone
hero-ragged_a2a-nccl2307-ep-step81k
herotraining-dynamicsgradient-norm
Helw150 h100-mix25-20260912-d1536
ladderdata-mixturemilestone
mcwitt hero-ragged_a2a-nccl2307-ep-step81k
herocheckpointtraining-dynamicsgradient-norm
Helw150 h100-mix25-20260912-d1536
ladderdata-mixture
hero-ragged_a2a-nccl2307-ep-step81k
herotraining-dynamicsgradient-norm
WhenWen AUG-LIN0-6L-DENSE-MUONH-X600-HIGH-1P7-d512-600x-lr1.7
ladderlr-sweep
WhenWen AUG-LIN0-6L-DENSE-MUONH-020 / -005
ladderlr-sweepmilestone
mcwitt hero-ragged_a2a-nccl2307-ep-step81k
herocheckpointmilestone
mcwitt hero-ragged_a2a-nccl2307-ep-step81k
herocheckpointmilestone
WhenWen AUG-LIN0-6L-DENSE-MUONH-019 / -005
ladderlr-sweep
WhenWen AUG-LIN0-6L-DENSE-MUONH-020 / -005
ladderlr-sweep
WhenWen AUG-LIN0-6L-DENSE-MUONH-019 / -005
ladderlr-sweepmilestone
WhenWen AUG-LIN0-6L-DENSE-MUONH-018-d512-600x-lr0.45
ladderlr-sweepmilestone
mcwitt hero-ragged_a2a-nccl2307-ep-step81k
herocheckpointmilestone
WhenWen AUG-LIN0-6L-DENSE-MUONH-X150-HIGH-1P7-d512-150x-lr1.7
ladderlr-sweepmilestone
mcwitt hero-ragged_a2a-nccl2307-ep-step81k
herocheckpointmilestone
mcwitt hero-ragged_a2a-nccl2307-ep-step81k
herotraining-dynamicsgradient-norm
mcwitt hero-ragged_a2a-nccl2307-ep-step81k
herocheckpointmilestone
WhenWen AUG-LIN0-6L-DENSE-MUONH-X150-HIGH-1P7
ladderlr-sweepmilestone
mcwitt hero-ragged_a2a-nccl2307-ep-step81k
herotraining-dynamicsgradient-norm
mcwitt hero-ragged_a2a-nccl2307-ep-step81k
heromilestone
mcwitt hero-ragged_a2a-nccl2307-ep-step81k
herocheckpointmilestone
Jeff H #hero-run-2026
herotraining-dynamicsgradient-norm
mcwitt hero-ragged_a2a-nccl2307-ep-step81k
herocheckpointmilestone
mcwitt hero-ragged_a2a-nccl2307-ep-step81k
herocheckpointmilestone
WhenWen AUG-LIN0-6L-DENSE-MUONH-005
ladderlr-sweepmilestone
WhenWen AUG-LIN0-6L-DENSE-MUONH-017 / -004
ladderlr-sweepmilestone
WhenWen AUG-LIN0-6L-DENSE-MUONH-017
ladderlr-sweepmilestone
WhenWen AUG-LIN0-6L-DENSE-MUONH-019 / -005
ladderlr-sweeprestart
WhenWen moe-lgr-9110-larry-d768-muonh
dialroutingmilestone
mcwitt hero-ragged_a2a-nccl2307-ep-step81k
herocheckpointmilestone
WhenWen AUG-LIN0-6L-DENSE-MUONH-019
ladderlr-sweeprestart
WhenWen AUG-LIN0-6L-DENSE-MUONH-019
ladderlr-sweepcrash
WhenWen moe-lgr-9110-larry-d768-muonh
experimentmoeroutingmilestone
WhenWen moe-lgr-9110-larry-d768-muonh
dialroutingrestart
WhenWen AUG-LIN0-6L-DENSE-MUONH-016
ladderlr-sweepmilestone
mcwitt hero-ragged_a2a-nccl2307-ep-step81k
herocheckpointrestart
WhenWen AUG-LIN0-6L-DENSE-MUONH-005 / -019
ladderlr-sweeprestart
WhenWen AUG-LIN0-6L-DENSE-MUONH-015 / -020
ladderlr-sweepmilestone
Larry #hero-run-2026
herotraining-dynamicsgradient-norm
mcwitt hero-ragged_a2a-nccl2307-ep-step81k
herorestarteval
mcwitt #8870
herohang
mcwitt hero-ragged_a2a-nccl2307-ep-step81k
herorestarthang
mcwitt hero-ragged_a2a-nccl2307-ep-step81k
herocrashhang
WhenWen AUG-LIN0-6L-DENSE-MUONH-018
ladderlr-sweeprestart
WhenWen AUG-LIN0-6L-DENSE-MUONH-018
ladderlr-sweeprestart
WhenWen AUG-LIN0-6L-DENSE-MUONH-015 … -019
ladderlr-sweepcrash
Sep 12
WhenWen AUG-LIN0-6L-DENSE-MUONH-019
ladderlr-sweepcrash
WhenWen AUG-LIN0-6L-DENSE-MUONH-004 / -019
ladderlr-sweep
WhenWen moe-lgr-9110-larry-d768-muonh
experimentmoerouting
WhenWen AUG-LIN0-6L-DENSE-MUONH-014
ladderlr-sweepmilestone
WhenWen moe-lgr-9110-larry-d768-muonh
experimentmoerouting
WhenWen AUG-LIN0-6L-DENSE-MUONH-017 / -018
ladderlr-sweep
WhenWen AUG-LIN0-6L-DENSE-MUONH-012 / -013
ladderlr-sweepmilestone
WhenWen moe-lgr-9110-larry-d768-muonh
experimentmoerouting
WhenWen moe-lgr-9110-larry-d768-muonh
experimentmoerouting
mcwitt hero-ragged_a2a-nccl2307-ep-step81k
herocheckpointmilestone
WhenWen moe-lgr-9110-larry-d768-muonh
experimentmoerouting
WhenWen moe-lgr-9110-larry-d768-muonh
experimentmoerouting
WhenWen AUG-LIN0-6L-DENSE-MUONH-016
ladderlr-sweep
WhenWen AUG-LIN0-6L-DENSE-MUONH-011
ladderlr-sweepmilestone
WhenWen AUG-LIN0-6L-DENSE-MUONH-X30-HIGH-1P7
ladderlr-sweep
WhenWen moe-lgr-9110-larry-d768-muonh
experimentmoerouting
mcwitt hero-ragged_a2a-nccl2307-ep-step81k
herocheckpointmilestone
WhenWen AUG-LIN0-6L-DENSE-MUONH-003
ladderlr-sweepmilestone
WhenWen AUG-LIN0-6L-DENSE-MUONH-014
ladderlr-sweep
WhenWen moe-lgr-9110-larry-d512-muonh
experimentmoeroutingmilestone
WhenWen AUG-LIN0-6L-DENSE-MUONH-011
ladderlr-sweep
WhenWen AUG-LIN0-6L-DENSE-MUONH-008
ladderlr-sweepmilestone
mcwitt hero-ragged_a2a-nccl2307-ep-step81k
herocheckpointmilestone
WhenWen AUG-LIN0-6L-DENSE-MUONH-006
ladderlr-sweep
WhenWen AUG-LIN0-6L-DENSE-MUONH-002
ladderlr-sweepmilestone
WhenWen moe-lgr-9110-larry-d512-muonh
experimentmoeroutingmilestone
mcwitt hero-ragged_a2a-nccl2307-ep-step81k
herocheckpointmilestone
Helw150 9126
ladderdata-mixhero
WhenWen AUG-LIN0-6L-DENSE-MUONH-001
ladderlr-sweepmilestone
WhenWen AUG-LIN0-6L-DENSE-MUONH
ladderlr-sweepmilestone
WhenWen moe-lgr-9110-larry-d512-muonh
experimentmoeroutingrestart
WhenWen AUG-LIN0-1L-DENSE-SGDH-X600-HIGH-4-d512-600x-lr4.3
ladderlr-sweepcrash
WhenWen moe-lgr-9110-larry-d512
experimentmoeroutingcrash
WhenWen AUG-LIN0-1L-DENSE-SGDH-X600-HIGH-4-d512-600x-lr4.3
ladderlr-sweeprestart
WhenWen moe-lgr-9110-d768-rmsgated
experimentmoeroutingmilestone
WhenWen AUG-LIN0-1L-DENSE-SGDH-X600-HIGH-4-d512-600x-lr4.3
ladderlr-sweep
mcwitt hero-ragged_a2a-nccl2307-ep-step81k
herocheckpointmilestone
WhenWen AUG-LIN0-1L-DENSE-SGDH-X600-HIGH-4-d512-600x-lr4.3
ladderlr-sweep
WhenWen AUG-LIN0-1L-DENSE-SGDH-X600-HIGH-4-d512-600x-lr4.3
ladderlr-sweep
mcwitt hero-ragged_a2a-nccl2307-ep-step81k
herocheckpointmilestone
held h100-mix25-20260912-d1024/d1536
ladderdata-mix
held h100-mix25-20260912-d768
Correction: the new four-width scaling ladder is not behind the August one. Read the way [the hero chart on this page](https://wandb.ai/marin-community/marin_moe/runs/hero-ragged_a2a-nccl2307-ep-step81k) reads its own eval curve — with the router's per-expert capacity limit switched off, so no token is dropped — both finished widths beat their August counterparts: [the 0.6-billion-parameter rung](https://wandb.ai/marin-community/marin_moe/runs/h100-mix25-20260912-d512) scored 3.3063 against [August's 3.3234](https://wandb.ai/marin-community/marin_moe/runs/h100-ladder-d512-ep8-bs1024-791tpp-20260824-rno2a), and [the 1.6-billion one](https://wandb.ai/marin-community/marin_moe/runs/h100-mix25-20260912-d768) 3.0012 against [August's 3.0181](https://wandb.ai/marin-community/marin_moe/runs/h100-ladder-d768-2xep8-bs1024-791tpp-20260824). The two entries that reported these rungs quoted the other reading, taken with the limit on, where the same two runs trail by about 0.05 — the new models lose about 0.066 more to dropping at both widths. The no-drop reading is also the quieter of the two: [one mixture's four random seeds](https://wandb.ai/marin-community/marin_moe/runs/h100-d512-mixprior-996f489106c7b922-seed3-from10pct-20260829) span 0.0031 on it against 0.0126 with the limit on, so the 0.017 lead is about five times its seed noise. The earlier caveat stands: three weeks of code changes and an extra mixture-stage switch separate the two ladders, so neither reading isolates the mixture the new ladder exists to test.
ladderdata-mixcorrection
held h100-mix25-20260912-d768
ladderdata-mixmilestone
mcwitt hero-ragged_a2a-nccl2307-ep-step81k
herocheckpointmilestone
WhenWen AUG-LIN0-1L-DENSE-SGDMH-X600-HIGH-1-d512-600x-lr1.1
ladderlr-sweepmilestone
WhenWen AUG-LIN0-1L-DENSE-SGDH-X600-HIGH-4-d512-600x-lr4.3
ladderlr-sweep
WhenWen AUG-LIN0-1L-DENSE-SGDH-X600-HIGH-4-d512-600x-lr4.3
ladderlr-sweep
WhenWen AUG-LIN0-1L-DENSE-MUONH-X600-HIGH-1P4-d512-600x-lr1.4
ladderlr-sweepmilestone
mcwitt hero-ragged_a2a-nccl2307-ep-step81k
herocheckpointmilestone
WhenWen AUG-LIN0-1L-DENSE-SGDMH-X600-HIGH-1P4-d512-600x-lr1.4
ladderlr-sweepmilestone
mcwitt hero-ragged_a2a-nccl2307-ep-step81k
herocheckpoint
WhenWen AUG-LIN0-1L-DENSE-MUONH-X600-HIGH-1-d512-600x-lr1.1
ladderlr-sweepmilestone
mcwitt hero-ragged_a2a-nccl2307-ep-step81k
herocheckpoint
WhenWen AUG-LRC-1L-SGDH-X600-HIGH-2-d512-600x-lr1.7
ladderlr-sweepmilestone
mcwitt hero-ragged_a2a-nccl2307-ep-step81k
herocheckpointmilestone
mcwitt hero-ragged_a2a-nccl2307-ep-step81k
herocheckpointmilestone
WhenWen AUG-LIN0-1L-DENSE-SGDH-X600-HIGH-4-d512-600x-lr4.3
ladderlr-sweep
WhenWen AUG-LIN0-1L-DENSE-SGDH-X600-HIGH-3-d512-600x-lr2.7
ladderlr-sweepmilestone
WhenWen AUG-LIN0-1L-DENSE-SGDH-X300-HIGH-4-d512-300x-lr4.3
ladderlr-sweepmilestone
held h100-mix25-20260912-d512
ladderdata-mixmilestone
WhenWen #9110 — first gate fails at d512 open
experimentmoeroutingcorrection
WhenWen moe-lgr-9110-d512-rmsgated
experimentmoeroutingmilestone
WhenWen #7856 LR sweep — 600x above-grid probes
ladderlr-sweepmilestone
mcwitt hero-ragged_a2a-nccl2307-ep-step81k
herocheckpointmilestone
WhenWen AUG-LIN0-1L-DENSE-SGDH-X150-HIGH-4-d512-150x-lr4.3
ladderlr-sweepmilestone
mcwitt hero-ragged_a2a-nccl2307-ep-step81k
heroevalstraining-dynamics
WhenWen AUG-LIN0-1L-DENSE-SGDMH-X300-HIGH-2-d512-300x-lr1.7
ladderlr-sweep
WhenWen moe-lgr-9110-d512-rmsgated / d768-rmsgated
experimentmoeroutingrestart
mcwitt hero-ragged_a2a-nccl2307-ep-step81k
herocheckpointmilestone
larrydial dynact-latent-in-double-d512-fa4fix
ladderablationcrash
WhenWen AUG-LIN0-1L-DENSE-MUONH-X300-HIGH-2-d512-300x-lr1.7
ladderlr-sweep
WhenWen AUG-LIN0-1L-DENSE-SGDMH-X600-HIGH-*
ladderlr-sweepmilestone
WhenWen AUG-LIN0-1L-DENSE-SGDH-X150/X300-HIGH-4
ladderlr-sweep
held h100-mix25-20260912-d512/d768/d1024/d1536
ladderdata-mixscaling-ladder
larrydial dynact-latent-in-double-d512-fa4fix
ladderablationcrash
WhenWen h100-d512-mixprior wave 5
ladderdata-mixmilestone
WhenWen AUG-LIN0-1L-DENSE-MUONH-X600-HIGH-*
ladderlr-sweep
larrydial dynact-latent-in-double-d512-fa4fix
ladderablationcrash
WhenWen h100-d512-mixprior wave 5
ladderdata-mix
WhenWen h100-d512-mixprior wave 5
ladderdata-mixcrash
WhenWen AUG-LRC-1L-SGDH-X600-HIGH-2-d512-600x-lr1.7
ladderlr-sweep
larrydial dynact-latent-* d512 fa4fix
ladderablation
WhenWen #7856 LR sweep — 600x grids
ladderlr-sweep
WhenWen h100-d512-mixprior-5f6e920ed52cd6d0-seed0-from10pct-20260829
ladderdata-mix
larrydial dynact-latent-in-double / out-half d512 fa4fix
ladderablationcrash
WhenWen h100-d512-mixprior-ba9dd4219b35ff93-seed0-from10pct-20260829
ladderdata-mixcrash
larrydial dynact-latent-in-half-d512-fa4fix
ladderarchitecture
WhenWen h100-d512-mixprior-df90e8b7f0951a11-seed0-from10pct-20260829
ladderdata-mix
larrydial dynact-latent-out-half-d512-fa4fix
ladderarchitecturecrash
kaiyuew #7856 LR sweep
ladderlrdensemilestone
kaiyuew h100-d512-mixprior wave 5
ladderdata-mixrestart
WhenWen AUG-LRC-1L-SGDH-X600-HIGH-2-d512-600x-lr1.7
ladderlr-sweepcrash
WhenWen AUG-LIN0-1L-DENSE-SGDH-X600-HIGH-3-d512-600x-lr2.7
ladderlrdensemilestone
held h100-d512-mixprior wave 5
ladderdata-mixmilestone
larrydial dynact-latent-* d512 fa4fix
The remaining three attention-shape probes came back, on the same rebuilt image as the first. [The arm that halves what enters the attention latent](https://wandb.ai/marin-community/marin_moe/runs/dynact-latent-in-half-d512-fa4fix), [the one that doubles what leaves it](https://wandb.ai/marin-community/marin_moe/runs/dynact-latent-out-double-d512-fa4fix) and [the one that halves what leaves it](https://wandb.ai/marin-community/marin_moe/runs/dynact-latent-out-half-d512-fa4fix) all restarted from step 0 between 2:13 and 2:14 a.m. (ET), each carrying the identical package change as [the arm that went first at 2:00 a.m. (ET)](https://wandb.ai/marin-community/marin_moe/runs/dynact-latent-in-double-d512-fa4fix): the fused attention kernel `flash-attn-4` moves from build b16 to b28 and the communication library steps back from 2.31.2 to 2.30.7. A whole cohort rebuilt together confirms this was an operator withdrawal, not four separate faults. Two corrections to earlier entries. This wave does not test expert design: all four arms keep the standard expert block and change the width of the attention latent — the narrow buffer each layer's attention reads from and writes to — while [the first wave](https://wandb.ai/marin-community/marin_moe/runs/dynact-dual-updown-d512) did the reverse. And the four are not matched on size: they carry 446M, 522M, 749M and 900M parameters against [the baseline](https://wandb.ai/marin-community/marin_moe/runs/dynact-baseline-fsdp-d512)'s 598M, because the latent width is itself what sets their parameter count. A win here would not separate a better shape from more parameters.
ladderarchitecturerestartcorrection
kaiyuew h100-d512-mixprior wave 5
ladderdata-mixrestart
kaiyuew #7856 LR sweep
ladderlr-sweep
kaiyuew #7856 LR sweep
ladderlr-sweep
larrydial dynact expert-design test
laddermoerestart
kaiyuew h100-d512-mixprior wave 5
ladderdata-mixrestart
kaiyuew h100-d512-mixprior-2d02e594203d6b7c-seed0-from10pct-20260829
ladderdata-mixmilestonecorrection
larrydial dynact-latent-* d512
ladderarchitecturecrash
WhenWen AUG-LIN0-1L-DENSE-SGDH-X300-HIGH-1-d512-300x-lr1.1
ladderlrdensemilestone
held h100-d512-mixprior wave 5
ladderdata-mixcrash
mcwitt hero-ragged_a2a-nccl2307-ep-step81k
herocheckpointmilestone
WhenWen AUG-LRC-1L-DENSE-SGDMH-021 … 025
ladderlrdensemilestone
larrydial dynact-latent-* d512
ladderarchitecture
held h100-d512-mixprior-ee3799698c71ea8f-seed0-from10pct-20260829
ladderdata-mixmilestone
kaiyuew h100-d512-mixprior-8d5f0aed232ce1b7-seed0-from10pct-20260829
ladderdata-mixmilestone
kaiyuew h100-d512-mixprior-5f6e920ed52cd6d0-seed0-from10pct-20260829
ladderdata-mixcrash
WhenWen AUG-LIN0-1L-DENSE-MUONH-X60-HIGH-2-d512-60x-lr1.7
ladderlrdensemilestone
kaiyuew h100-d512-mixprior-996f489106c7b922-seed3-from10pct-20260829
ladderdata-mixcrash
kaiyuew h100-d512-mixprior-dd6c1760c4691b33-seed0-from10pct-20260829
ladderdata-mixrestart
WhenWen AUG-LIN0-1L-DENSE-SGDH-X600-HIGH-1-d512-600x-lr1.1
ladderlrdensemilestone
WhenWen AUG-LIN0-1L-DENSE-SGDH-X150-HIGH-3-d512-150x-lr2.7
ladderlrdense
WhenWen AUG-LIN0-1L-DENSE-SGDMH-X150-HIGH-1-d512-150x-lr1.1
ladderlrdensemilestone
mcwitt hero-ragged_a2a-nccl2307-ep-step81k
herocheckpointmilestone
WhenWen h100-d512-mixprior wave 5
ladderdata-mixh100
WhenWen AUG-LIN0-1L-DENSE-SGDH-X150-HIGH-2-d512-150x-lr1.7
ladderlrdensemilestone
WhenWen AUG-LIN0-1L-DENSE-MUONH-X30/X60-HIGH-2
ladderlrdensemilestone
WhenWen h100-d512-mixprior wave 5
ladderdata-mixh100
WhenWen AUG-LIN0-1L-DENSE-MUONH-X30/X60-HIGH
ladderlrdensemilestone
WhenWen AUG-LIN0-1L-DENSE-SGDH-X60-HIGH-3-d512-60x-lr2.7
ladderlrdensemilestone
WhenWen AUG-LIN0-1L-DENSE-SGDMH-X150/X300-HIGH-1
ladderlrdensemilestone
mcwitt hero-ragged_a2a-nccl2307-ep-step81k
herocheckpointmilestone
WhenWen AUG-*-HIGH-* boundary extensions
ladderlrdensemilestone
WhenWen AUG-LIN0-1L-DENSE-SGDMH-020-d512-300x-lr0.7
ladderlrdense
WhenWen AUG-LIN0-1L-DENSE-SGDH-X60-HIGH-3-d512-60x-lr2.7
ladderlrdense
WhenWen AUG-LRC-1L-SGDH-X600-HIGH-1-d512-600x-lr1.1
ladderlrmoemilestone
larrydial dynact-dual-updown-d512
laddermoeh100milestone
larrydial dynact-coupled-d512-v3
laddermoeh100
WhenWen h100-d512-mixprior wave 5
ladderdata-mixh100
WhenWen h100-d512-mixprior wave 4
ladderdata-mixh100milestone
larrydial dynact-coupled-d512-v3
laddermoeh100
larrydial dynact-dual-updown-d512
laddermoeh100milestone
larrydial dynact-double-neurons-d512
laddermoeh100milestone
larrydial dynact-dual-updown-d512
laddermoeh100
larrydial dynact-coupled-d512-v3
laddermoeh100
mcwitt hero-ragged_a2a-nccl2307-ep-step81k
herocheckpointmilestone
larrydial dynact-coupled-d512-v3
laddermoeh100
larrydial dynact-dual-updown-d512
laddermoeh100
WhenWen AUG-*-d512-600x
ladderlrdense
larrydial dynact-coupled-d512-v3
laddermoeh100
larrydial dynact-dual-updown-d512
laddermoeh100
WhenWen AUG-LIN0-1L-DENSE-MUONH-016 … 020
ladderlrdensemilestone
WhenWen AUG-*-d512-600x
ladderlrdensemilestone
larrydial dynact-coupled-d512-v2
laddermoeh100
larrydial dynact-baseline-fsdp-d512
laddermoeh100
mcwitt hero-ragged_a2a-nccl2307-ep-step81k
herocheckpointmilestone
larrydial dynact-coupled-d512-v2 / -double-neurons / -dual-updown
laddermoeh100
WhenWen AUG-LIN0-1L-DENSE-SGDMH-016 … 020
ladderlrdensemilestone
larrydial dynact-coupled-d512
laddermoeh100
WhenWen AUG-*-d512-600x
ladderlrdensemilestone
WhenWen AUG-LIN0-1L-DENSE-SGDH-016 … 020
ladderlrdensemilestone
Sep 11
WhenWen AUG-LRC-1L-DENSE-SGDMH-016 … 024
ladderlrdensemilestone
ravwojdyla #9120 open
monitoring
larrydial dynact-baseline-fsdp-d512
laddermoeh100
mcwitt hero-ragged_a2a-nccl2307-ep-step81k
herocheckpointmilestone
WhenWen AUG-LIN0-1L-DENSE-SGDH-X150-HIGH-1 / X60-HIGH-2
ladderlrdensemilestone
WhenWen AUG-LRC-1L-DENSE-SGDH-X150-HIGH-1
ladderlrdensemilestone
mcwitt lb0911moebared3
herobenchone-rackmilestone
mcwitt #9118 open
long-context
WhenWen AUG-LIN0-1L-DENSE-SGDH-X30-HIGH-2
ladderlrdensemilestone
mcwitt lb0911moebared3
herobenchone-rack
WhenWen AUG-LIN0-1L-DENSE-SGDH-X30-HIGH-2 / X60-HIGH-2
ladderlrdensemilestone
mcwitt hero-ragged_a2a-nccl2307-ep-step81k
herocheckpointmilestone
mcwitt lb0911moecurrentd3
herobenchone-rackmilestone
WhenWen AUG-LIN0-1L-DENSE-SGDH-X30-HIGH-1 / X60-HIGH-1
ladderlrdensemilestone
mcwitt lb0911moecurrentd3
herobenchone-rack
WhenWen h100-d512-mixprior — fourth wave
ladderdata-mix
WhenWen moe-lgr-9110-d512-rmsgated / d768-rmsgated
experimentmoeroutingcrash
WhenWen AUG-LIN0-1L-DENSE-SGDH-X30-HIGH-1 / X150-HIGH-1
ladderlrdensemilestone
mcwitt lb0911moecurrentd2
herobenchone-rackmilestone
WhenWen h100-d512-mixprior — fourth wave
ladderdata-mix
mcwitt #9119 open
long-context
WhenWen AUG-LIN0-1L-DENSE-SGDH-X60-HIGH-1
ladderlrdensemilestone
WhenWen h100-d512-mixprior — seed repeats
ladderdata-mixmilestone
mcwitt hero-ragged_a2a-nccl2307-ep-step81k
herocheckpointmilestone
mcwitt hero-ragged_a2a-nccl2307-ep-step81k
herocheckpoint
WhenWen AUG-LRC-1L-DENSE-SGDH-X150-HIGH-1
ladderlrdensemilestone
WhenWen AUG-LIN0 / AUG-LRC-1L-DENSE-011 … 020
ladderlrdense
mcwitt lb0911moebared2
herobenchone-rackmilestone
WhenWen h100-d512-mixprior — fourth wave
ladderdata-mix
WhenWen moe-lgr-9110-d512-rmsgated / d768-rmsgated
experimentmoerouting
mcwitt hero-ragged_a2a-nccl2307-ep-step81k
herocheckpointmilestone
ClassicLarry #8827
heroeval
WhenWen AUG-LIN0 / AUG-LRC-1L-DENSE-006 … 015
ladderlrdensemilestone
mcwitt lb0911moebared1
herobenchone-rackmilestone
mcwitt hero-ragged_a2a-nccl2307-ep-step81k
herocheckpoint
mcwitt lb0911moecurrentd1
herobenchone-rackmilestone
WhenWen AUG-LIN0-1L-DENSE-SGDH / MUONH / SGDMH-006 … 015
ladderlrdense
WhenWen AUG-LRC-1L-DENSE-SGDMH-010 … 015
ladderlrdense
mcwitt lb0911moecurrentd1
herobenchone-rack
WhenWen AUG-LRC-1L-DENSE-SGDMH-006 … 014
ladderlrdense
WhenWen AUG-LIN0-1L-DENSE-SGDH / MUONH / SGDMH-001 … 010
ladderlrdense
mcwitt hero-ragged_a2a-nccl2307-ep-step81k
heroeval
WhenWen AUG-LRC-1L-DENSE-SGDMH-001 … 010
ladderlrdense
WhenWen AUG-LIN0-1L-DENSE-SGDH / MUONH / SGDMH-001 … 010
ladderlrdense
mcwitt hero-ragged_a2a-nccl2307-ep-step81k
herocheckpoint
mcwitt lc-main-ragged-262k-0911
long-contextheromilestone
WhenWen moe-lgr-9110-d512-gated-smoke
ladderrouter
WhenWen AUG-LIN-1L-DENSE-SGDH / MUONH-001 … 005
ladderlrdensecrash
WhenWen AUG-LIN0-1L-DENSE-SGDH / MUONH / SGDMH-001 … 005
ladderlrdense
WhenWen #7856
ladderlrcorrection
WhenWen AUG-LRC-1L-DENSE-SGDMH-001 … 005
ladderlrdense
WhenWen moe-lgr-9110-d512-gated-smoke
ladderrouter
WhenWen #9110
ladderrouter
held h100-d512-mixprior × 15
ladderdata
WhenWen AUG-LIN-1L-DENSE-SGDH / MUONH-001 … 005
ladderlrdense
mcwitt lc-main-ragged-262k-0911
long-contexthero
WhenWen #7856
ladderlrdense
mcwitt hero-ragged_a2a-nccl2307-ep-step81k
herocheckpoint
mcwitt lc-main-ragged-262k-0911
long-contexthero
WhenWen AUG-LRC-1L-SGDH-025-d512-600x-lr0.7
ladderlrmilestone
mcwitt lc-main-ragged-262k-0911
long-contexthero
mcwitt hero-ragged_a2a-nccl2307-ep-step81k
heromilestone
WhenWen AUG-LRC-1L-DENSE-MUONH-024 / 025
ladderlrdense
WhenWen AUG-LRC-1L-DENSE-MUONH-019 / 020
ladderlrdense
mcwitt hero-ragged_a2a-nccl2307-ep-step81k
herocheckpoint
mcwitt hero-ragged_a2a-nccl2307-ep-step81k
herocheckpoint
WhenWen AUG-LRC-1L-025, AUG-LRC-1L-SGDH-024
ladderlr
WhenWen AUG-LRC-1L-DENSE-SGDH-019 / 020, MUONH-016 / 017 / 018
ladderlrdense
WhenWen AUG-LRC-1L-DENSE-SGDH-024 / 025, MUONH-021 / 022 / 023
ladderlrdense
mcwitt hero-ragged_a2a-nccl2307-ep-step81k
herocheckpoint
WhenWen AUG-LRC-1L-SGDH-X300-HIGH-1-d512-300x-lr1.1
ladderlr
WhenWen AUG-LRC-1L-DENSE-SGDH-018 / 022 / 023
ladderlrdense
WhenWen AUG-LRC-1L-DENSE-SGDH-016 / 017
ladderlrdense
WhenWen AUG-LRC-1L-DENSE-SGDH-021-d512-600x-lr0.1
ladderlrdense
mcwitt hero-ragged_a2a-nccl2307-ep-step81k
heromilestone
WhenWen AUG-LRC-1L-022 / 023 / 024, AUG-LRC-1L-SGDH-021 / 022
ladderlr
mcwitt hero-ragged_a2a-nccl2307-ep-step81k
herocheckpoint
WhenWen AUG-LRC-1L-SGDH-023, AUG-LRC-1L-021
ladderlr
WhenWen AUG-LRC-1L-DENSE-MUONH-014 / 015
ladderlrdense
WhenWen AUG-LRC-1L-DENSE-MUONH-019 / 020
ladderlrdense
mcwitt hero-ragged_a2a-nccl2307-ep-step81k
herocheckpoint
held h100-d512-mixprior × 40
ladderdata
WhenWen AUG-LRC-1L-DENSE-SGDH-014 / MUONH-011 / 012 / 013
ladderlrdense
WhenWen AUG-LRC-1L-DENSE-SGDH-019 / 020, MUONH-016 / 017 / 018
ladderlrdense
WhenWen AUG-LRC-1L-DENSE-SGDH-015-d512-150x-lr0.7
ladderlrdense
mcwitt hero-ragged_a2a-nccl2307-ep-step81k
herocheckpoint
held h100-d512-mixprior × 40
ladderdata
WhenWen AUG-LRC-1L-DENSE-SGDH-014-d512-150x-lr0.45
ladderlrdenserestart
WhenWen AUG-LRC-1L-DENSE-SGDH-017 / 018
ladderlrdense
WhenWen AUG-LRC-1L-DENSE-SGDH-011 / 012 / 013
ladderlrdense
WhenWen AUG-LRC-1L-DENSE-SGDH-016-d512-300x-lr0.1
ladderlrdense
WhenWen AUG-LRC-1L-DENSE-MUONH-010 / 015
ladderlrdense
held h100-d512-mixprior × 40
ladderdatarestart
mcwitt hero-ragged_a2a-nccl2307-ep-step81k
herocheckpoint
WhenWen AUG-LRC-1L-DENSE-MUONH-009 / 014
ladderlrdense
held h100-d512-mixprior × 40
ladderdatarestart
WhenWen AUG-LRC-1L-DENSE-MUONH-010, DENSE-SGDH-014
ladderlrdenserestart
WhenWen AUG-LRC-1L-SGDH-X300-HIGH-1-d512-300x-lr1.1
ladderlr
WhenWen AUG-LRC-1L-DENSE-MUONH-011 / 012 / 013
ladderlrdense
held h100-d512-mixprior × 40
ladderdatarestart
mcwitt hero-ragged_a2a-nccl2307-ep-step81k
herocheckpoint
held h100-d512-mixprior × 40
ladderdatacrash
WhenWen AUG-LRC-1L-DENSE-SGDH-009 / 015
ladderlrdense
WhenWen AUG-LRC-1L-SGDH-X150-HIGH-1-d512-150x-lr1.1
ladderlr
mcwitt hero-ragged_a2a-nccl2307-ep-step81k
herocheckpoint
held h100-d512-mixprior × 40
ladderdatarestart
WhenWen AUG-LRC-1L-DENSE-MUONH-009 / 010, SGDH-013 / 014
ladderlrdense
held h100-d512-mixprior × 40
ladderdatacrash
WhenWen AUG-LRC-1L-DENSE-SGDH-006 / 007 / 008
ladderlrdense
WhenWen AUG-LRC-1L-DENSE-SGDH-009-d512-60x-lr0.45
ladderlrdenserestart
WhenWen AUG-LRC-1L-DENSE-SGDH-011 / 012
ladderlrdense
WhenWen AUG-LRC-1L-DENSE-MUONH-006 / 007 / 008
ladderlrdense
held h100-d512-mixprior × 40
ladderdatacrash
WhenWen AUG-LRC-1L-DENSE-MUONH-004 / 005
ladderlrdense
WhenWen AUG-LRC-1L-DENSE-SGDH-009-d512-60x-lr0.45
ladderlrdensecrash
held h100-d512-mixprior × 40
ladderdatarestart
held h100-d512-mixprior × 40
ladderdata
WhenWen AUG-LRC-1L-DENSE-SGDH × 5
ladderlrdense
mcwitt hero-ragged_a2a-nccl2307-ep-step81k
herocheckpoint
WhenWen AUG-LRC-1L-DENSE-MUONH-002 / 003
ladderlrdense
WhenWen AUG-LRC-1L-DENSE-SGDH-006-d512-60x-lr0.1
ladderlrdensecrash
WhenWen AUG-LRC-1L-DENSE-SGDH × 5
ladderlrdense
held h100-d512-mixprior × 40
ladderdatacrash
held h100-d512-mixprior × 40
ladderdata
WhenWen AUG-LRC-1L-023 / SGDH-021
ladderlrrestart
WhenWen AUG-LRC-1L-DENSE-MUONH-001-d512-30x-lr0.1
ladderlrdense
WhenWen AUG-LRC-1L-SGDH-X150-HIGH-1-d512-150x-lr1.1
ladderlr
held h100-d512-mixprior × 40
ladderdata
WhenWen AUG-LRC-1L-DENSE-SGDH × 5
ladderlrdense
WhenWen AUG-LRC-1L-SGDH-025-d512-600x-lr0.7
ladderlr
WhenWen AUG-LRC-1L-023 / SGDH-021
ladderlrcrash
mcwitt hero-ragged_a2a-nccl2307-ep-step81k
herocheckpoint
held h100-d512-mixprior × 40
ladderdatarestart
WhenWen AUG-LRC-1L-SGDH-024-d512-600x-lr0.45
ladderlr
held h100-d512-mixprior × 40
ladderdatarestart
hero-ragged_a2a-nccl2307-ep-step81k
infracheckpointing
held h100-d512-mixprior × 40
ladderdatahang
held h100-d512-mixprior × 40
ladderdatahang
WhenWen AUG-LRC-1L-025-d512-600x-lr0.7
ladderlr
WhenWen AUG-LRC-1L-024-d512-600x-lr0.45
ladderlr
hero-ragged_a2a-nccl2307-ep-step81k
infracheckpointing
hero-ragged_a2a-nccl2307-ep-step81k
infracheckpointing
WhenWen AUG-LRC-1L × 6
ladderlr
held h100-d512-mixprior × 40
ladderdata
hero-ragged_a2a-nccl2307-ep-step81k
infracheckpointingeval
Larry #hero-run-2026
routerstability
WhenWen AUG-LRC-1L × 40
ladderlr
hero-ragged_a2a-nccl2307-ep-step81k
infracheckpointing
hero-ragged_a2a-nccl2307-ep-step81k
infracheckpointing
held h100-d512-mixprior × 20
ladderdata
Sheng Zha #hero-run-2026
routerstability
held h100-d512-mixprior-8b014e07
ladderhang
Sep 10
held h100-d512-mixprior × 20
ladderdata
held h100-d512-mixprior-f297c5b7
ladder
mcwitt lc-main-ragged-4k-0910
seqlenep
held h100-d512-mixprior-09c11815
ladderhang
held h100-d512-mixprior-09c11815
ladderhang
held h100-d512-mixprior × 20
ladderhang
hero-ragged_a2a-nccl2307-ep-step81k
infracheckpointing
held h100-d512-mixprior-8b014e07
ladderhang
held h100-d512-mixprior-a6b2cadc
ladderhang
held h100-d512-mixprior-8b014e07
ladderhang
hero-ragged_a2a-nccl2307-ep-step81k
infracheckpointing
held h100-d512-mixprior-7026cf51
ladderhang
mcwitt #8934
incidenthardware
held h100-d512-mixprior × 20
ladderdata
held h100-d512-mixprior-a6b2cadc
ladderhang
mcwitt lc-ragged-cp2-smoke-0910
seqlenep
mcwitt #9086
incidenthardwaremonitoring
hero-ragged_a2a-nccl2307-ep-step81k
infracheckpointing
hero-ragged_a2a-nccl2307-ep-step81k
infracheckpointing
held h100-d512-mixprior × 20
ladderdata
mcwitt #8870
hangdebugging
mcwitt #9082 open
monitoring
hero-ragged_a2a-nccl2307-ep-step81k
infracheckpointing
mcwitt #8870
hangdebugging
mcwitt #9077
kernelshang
held h100-d512-mixprior × 20
ladderdata
hero-ragged_a2a-nccl2307-ep-step81k
infracheckpointing
hero-ragged_a2a-nccl2307-ep-step81k
infraep
hero-ragged_a2a-nccl2307-ep-step81k
infracheckpointing
hero-ragged_a2a-nccl2307-ep-step81k
infracheckpointing
hero-ragged_a2a-nccl2307-ep-step81k
infracheckpointing
hero-ragged_a2a-nccl2307-ep-step81k
memoryeval
hero-ragged_a2a-nccl2307-ep-step81k
infracheckpointing
held h100-d512-mixprior × 20
ladderdata
hero-ragged_a2a-nccl2307-ep-step81k
infracheckpointing
hero-ragged_a2a-nccl2307-ep-step81k
evalmemory
held h100-d512-mixprior × 19
ladderdata
hero-ragged_a2a-nccl2307-ep-step81k
infraepcheckpointing
hero-ragged_a2a-nccl2307-ep-step81k
infraepcheckpointing
held h100-d512-mixprior × 17
ladderdata
mcwitt #8506
infraepincident
hero-ragged_a2a-nccl2307-ep-step81k
infraepincident
hero-ragged_a2a-nccl2307-ep-step81k
infraepincident
loom-oa-dev #8506
infraepincident
held h100-d512-mixprior × 6
ladderdata
mcwitt #8506
infraepincident
mcwitt #8506
infraepincident
hero-ragged_a2a-nccl2307-ep-step81k
infraepincident
hero-ragged_a2a-nccl2307-ep-step81k
infraepincident
Larry #hero-run-2026
training-dynamics
mcwitt #8506
infracheckpointsep
mcwitt restore-nccl2307-82906
infraep
hero-ragged_a2a-nccl2307-ep-step81k
infraep
held h100-d512-mixprior × 20
ladderdata
hero-ragged_a2a-nccl2307-ep-step81k
The replay is finished, it matched, and the run is now past the ground its predecessor covered. Step 81,916 — the last of the two hundred steps the relaunch set out to repeat — was logged at 7:56:34 p.m. (ET), and the run passed 81,919, the final step the [pooled-wave control](https://wandb.ai/marin-community/marin_moe/runs/hero-wd-gate-router-p02-step58k) ever reached, at 7:57:30 p.m. (ET). It stood at step 81,935 at 8:02:35 p.m. (ET), sixteen steps beyond that mark, so everything from here is training the project did not have. The run has stepped without a break for sixty-six minutes since its first step at 6:56:15 p.m. (ET) — 220 steps at a median 17.1 seconds, no gap longer than 54 seconds — and W&B holds it running with a heartbeat current to the minute. Its GPUs read the profile of training rather than of a stall: the memory subsystem is busy at 22-25% of capacity with tens of gigabytes a second moving on the links between GPUs, the opposite of the pegged-but-idle pattern the day's four hangs showed. The check the relaunch exists to make has now passed in full: over all 201 steps of the repeated window, loss against the [pre-merge trial run](https://wandb.ai/marin-community/marin_moe/runs/hero-ragged_a2a-ep-step81k) agrees to a mean difference of 0.00008 nats, in a band from −0.00046 to +0.00030. The released build [#9062](https://github.com/marin-community/marin/pull/9062) therefore reproduces the private one the trial ran, and the [data-loader fix](https://github.com/marin-community/marin/pull/9040) that landed on main in between has not changed the order the training data arrives in or what the model learns from it. Set against the control on the 26 steps the two share, loss is 0.0065 nats lower; MFU averages 21.9% with a median of 23.2% against the control's 21.5%; the share of tokens dropped for want of room at their expert averages 7.3 in 100,000 against the control's 3.4 in 100; routing entropy is unchanged at 5.949; and throughput averages 2.57 million tokens a second.
infraep
Sep 9
mcwitt #8506
infraep
loom-oa-dev #8945
infraincident
hero-ragged_a2a-nccl2307-ep-step81k
The trial run was stopped and relaunched under a new name that records the communication library it was built against. `hero-ragged_a2a-ep-step81k` logged its last step, 82,025, at 6:49:53 p.m. (ET). Over the next twenty seconds its GPUs wound down — utilisation, memory traffic and traffic on the links between GPUs all falling toward zero — and at 6:50:10 p.m. (ET) they stopped reporting at all. That is the profile of processes exiting, and it is the opposite of the hangs earlier today, which held the GPUs at full utilisation and kept reporting throughout. W&B has since moved the run to the crashed state on its missing heartbeat. Two minutes after the shutdown, at 6:52:17 p.m. (ET), [a new run](https://wandb.ai/marin-community/marin_moe/runs/hero-ragged_a2a-nccl2307-ep-step81k) started on the same machine, restored the same pinned step-81716 checkpoint, and logged step 81,716 at 6:56:15 p.m. (ET). It has stepped every 19 seconds since, reaching 81,744 at 7:05:34 p.m. (ET), and its GPUs read the profile of training: memory subsystem busy near a quarter of capacity, tens of gigabytes a second on the links between GPUs, and the matrix units active close to half the time. Its recorded software differs from the stopped attempt in one line: the layer that runs the compiled model on the GPUs is build `b040736f2c4a` rather than `b183c591142c`. Both are compiled against the newer NVIDIA communication library the job loads, which is what the new name reports. A fresh run id is the project's [standing convention](https://github.com/marin-community/marin/pull/8868) for a relaunch that repeats steps, because W&B refuses to log a step below a run's last one. The cost is that the 309 steps the stopped attempt logged — 106 of them past anything the [old run](https://wandb.ai/marin-community/marin_moe/runs/hero-wd-gate-router-p02-step58k) reached — are being run again. Nothing on the [status log](https://github.com/marin-community/marin/issues/8506) covered the change when this was written, because the mirror this page reads lags GitHub by about an hour. The [entry explaining it](https://github.com/marin-community/marin/issues/8506#issuecomment-5609817090) had in fact been posted at 6:51 p.m. (ET).
infraep
mcwitt #8506
mcwitt explains the swap this page could not account for: the hero was redeployed from the main branch. The wheel that carried the trial sat on a local branch behind a signed URL, so the run was not reproducible from main. [#9062](https://github.com/marin-community/marin/pull/9062) merged the same change as a released build, landing on main as `04fb348456`. The trial coordinator was cancelled at 6:50:30 p.m. (ET) at about step 82,020 after 300 clean steps; its tree is kept as the trial artifact and is not a lineage source. The [new run](https://wandb.ai/marin-community/marin_moe/runs/hero-ragged_a2a-nccl2307-ep-step81k) was [launched](https://github.com/marin-community/marin/issues/8506#issuecomment-5609812430) at 6:50:42 p.m. (ET) from a pristine worktree at that commit, restoring the old hero's pinned step-81716 checkpoint so the lineage stays on main commits. Its first 200 steps replay the validated window against the pooled-wave control, which also checks the released binary against the pre-merge build the trial ran. Main also took [#9040](https://github.com/marin-community/marin/pull/9040), a fix to two dataset loaders, between the two commits; the paired replay will expose it if it changes data order. A caution closes the note: the regional telemetry service for this cluster has refused queries since about 6:22 p.m. (ET) and the hub has nothing newer from it, so a `TrainingTelemetryGone` alarm fired at 6:36 p.m. (ET) for the trial run while W&B showed it stepping normally. Liveness is being read from W&B until that service recovers.
infraepincident
hero-ragged_a2a-ep-step81k
The trial window is complete and the run has moved past the ground the old run covered. Step 81,916 — the two-hundredth and last step of the window opened at the cutover — was logged at 6:16 p.m. (ET), and the run carried on through the control's final step, 81,919, at 6:17 p.m. (ET), so every step from 81,920 is new training rather than a replay. It stood at step 81,956 at 6:28 p.m. (ET), and W&B reports it running with a current heartbeat. The attempt has now stepped for 59 minutes without a break: 199 steps at a median 17.1 seconds each, with no gap longer than 62 seconds. Scored over the whole window against the [control](https://wandb.ai/marin-community/marin_moe/runs/hero-wd-gate-router-p02-step58k) on the same batches, three of the four conditions in mcwitt's [gate](https://github.com/marin-community/marin/issues/8506#issuecomment-5606002395) are met. Loss ran 0.0064 nats below the control across the 26 steps W&B retained for it in this range, in a narrow band from 0.0060 to 0.0067, against the ~0.006 the dropless routing was expected to earn — inside the gate's bf16-noise allowance. The share of tokens dropped for want of room at their expert averaged 7.7 in 100,000 and peaked at 1.4 in 10,000, against a ceiling of 1 in 1,000 and the control's 3.4 in 100. MFU averaged 23.1% with a median of 23.5%, against the control's 21.5% and a floor of about 21%. Routing entropy was unchanged at 5.949. The fourth condition — no hang, no retry, no watchdog kill — is not met, but every hang fell on the build that the sixth attempt replaced, so the window was carried by software the gate was not written for. Whether that counted as a pass was mcwitt's call; he [ruled it a pass](https://github.com/marin-community/marin/issues/8506#issuecomment-5609573154) at 6:24 p.m. (ET), a ruling this page could not yet see when this was written.
infraepthroughput
mcwitt #8506
infraepthroughput
mcwitt #8870
mcwitt reviews the evidence on the silent hang and withdraws part of it. A table of the four eleven-rack ragged all-to-all runs that ran clean shows two of them used the older compiled headers, so his earlier statement that every clean reproduction used the newer headers was wrong: verified exposure on the newer headers is 402 completed steps across two runs, against 237 clean steps on the old ones. The earlier clean windows are also not comparable — the first two omitted the checkpointer's per-step broadcast, and the third restored it but added bounded barriers, wait instrumentation, initialisation waits and post-collective device-to-host copies, so its success cannot isolate the header change. Today's production attempt is the stronger evidence: 704 GPU processes on the released kernels, with the source difference from the failing build reduced to a [two-line build-default change](https://github.com/marin-community/xla/pull/22), and without the September 4 diagnostic barriers and copies — which argues those are not needed for the improvement. He is explicit that this is one uninstrumented success with build, allocation and timing differences, not a randomised wheel A/B or a tested restart sequence, and that a hardware fabric fault stays credible given the [September 5 NVLink error](https://github.com/marin-community/marin/issues/8870#issuecomment-5554893088) on the pooled-wave hero. Preliminary paired numbers over matched steps 81,726–81,758 show loss differing by at most 0.000110 nats, with MFU 0.48 points lower on the new wheel — observational, and not enough to establish a regression either way. Sustained training and restart validation remain open, tracked on the [incident record](https://echo.oa.dev/wiki/378).
infraepkernelsincident
mcwitt #9062 merged
infrakernelsep
hero-ragged_a2a-ep-step81k
infraep
#8506
mcwitt's account of the fourth hang corrects this page's count of the attempts and names the change that followed. Attempt 4 (`coord-f48b4a48`) restored the checkpoint, cleared its loader at 4:57 p.m. (ET), then took no step for ten minutes, with every task on rack index 2 running 3.6–4.4 CPU cores against 7.6–9.0 elsewhere — the same per-rack asymmetry that caught the two attempts before it. An NCCL RAS sample at 5:01 p.m. (ET) put every cross-rack AllReduce three operations apart, and py-spy dumps from the same minute placed every rank, stalled and healthy alike, inside the compiled training step rather than in the data loader or the compiler. The tally on main `9ccc1bd5e4` is four hangs in four attempts, on rack indices 7, 9, 1 and 2, after 43, 0, 0 and 0 completed steps. Attempt 5 (`coord-c6b2ede7`) was cancelled two minutes after submit, before it restored anything, in order to change wheels. Attempt 6 — the one now training — launched at 5:11 p.m. (ET) as `coord-72aad6e7` from an unpushed local branch: `9ccc1bd5e4` plus a single change, the PJRT wheel `0.11.1+marin.b183c591142c`, whose device kernels are compiled against NCCL 2.30.7 headers, in place of `3a8f3392723e`, compiled against 2.29.7. The run's own package set installs `nvidia-nccl-cu13==2.30.7`, so the new kernels are built against the same version of the collective library the job actually loads. Everything else, including the step-81716 handoff, is unchanged, and the wheel sits under a seven-day temporary prefix. This is the second time a fault in this backend has turned on how the PJRT device kernels were built: [#8313](https://github.com/marin-community/marin/issues/8313) traced a first-step crash at this expert-parallel width to a 32-peer bound inside a barrier kernel, and fixed it with a rebuilt wheel rather than a code change in the model.
infraepkernelsincident
hero-ragged_a2a-ep-step81k
The replacement run has hung twice more, and neither attempt reached a training step. The attempt that attached at 4:41 p.m. (ET) followed the launch profile in order — GPU memory dropped to the level held before any model is loaded, climbed through a checkpoint restore, and produced one burst of full-power compute at 4:44 p.m. (ET) — and then stopped. From that burst onward its GPUs read full utilisation with their memory subsystem idle and no traffic on the links between them, and their resident memory sat at exactly the figure the first attempt held while hung. That process was killed about four minutes later. A third attempt attached at 4:51 p.m. (ET), restored a checkpoint by 4:54 p.m. (ET), reached the same resident-memory figure at 4:56 p.m. (ET), and has not moved since: full utilisation, no memory traffic, no link traffic, and — unlike the second attempt — not one burst of real work at any point. For scale, the first attempt went from restore to full-rate training in half a minute and logged its first step two minutes after that; this one has been past that mark for seven minutes. Its heartbeat to W&B stopped at 4:53 p.m. (ET), two minutes into the attempt, although the machine monitor on the same job keeps reporting and W&B still lists the run as running. The last step the run logged was at 3:59 p.m. (ET), an hour ago. This is the signature of the open fault [#8870](https://github.com/marin-community/marin/issues/8870): a job spinning inside an exchange between machines that never completes. Nobody has claimed any of the restarts. Nothing has been posted to the [status log](https://github.com/marin-community/marin/issues/8506) since the 3:49 p.m. (ET) [first-steps note](https://github.com/marin-community/marin/issues/8506#issuecomment-5607788308), and no rollback run has appeared in the project. The [plan](https://github.com/marin-community/marin/issues/8506#issuecomment-5606002395) mcwitt wrote before the cutover called for an immediate rollback on the first hang — relaunching the old configuration and resuming [the run this one replaced](https://wandb.ai/marin-community/marin_moe/runs/hero-wd-gate-router-p02-step58k) from that tree's own checkpoint — and the go/no-go it set for about 4:45 p.m. (ET) has passed.
infraep
mcwitt #8506
infraep
hero-ragged_a2a-ep-step81k
The hung process was killed and the run is trying again. After thirty-nine minutes in which [the replacement run](https://wandb.ai/marin-community/marin_moe/runs/hero-ragged_a2a-ep-step81k) held its GPUs at full utilisation while their memory and the links between them stayed idle, that process ended between 4:38 and 4:41 p.m. (ET), and a new one attached to the same run name at 4:41:09 p.m. (ET). Its opening minutes repeat the 3:37 p.m. (ET) launch in order — GPU memory dropped back to the level held before any model is loaded, then climbed through a checkpoint restore, then produced one burst of full-power compute at 4:44 p.m. (ET) — but it reached that point in three minutes rather than eleven, the executable no longer needing a cold compile. Memory is still rising toward the level the previous attempt held, so the attempt is still starting; no step has been logged, and it is too early to say whether this one will run. Nobody has claimed the restart. Nothing has been posted to the [status log](https://github.com/marin-community/marin/issues/8506) since the 3:49 p.m. (ET) [first-steps note](https://github.com/marin-community/marin/issues/8506#issuecomment-5607788308), and when the [September 3 attempt](https://wandb.ai/marin-community/marin_moe/runs/hero-ragged_a2a-ep-step54k) hung on this same routing backend it was the progress watchdog that killed and relaunched it. That precedent is not encouraging: the September 3 relaunch hung again after a single step, and the hero was rolled back to the pooled-wave configuration within the hour. The restart also decides the trial. The [gate](https://github.com/marin-community/marin/issues/8506#issuecomment-5606002395) mcwitt wrote before the cutover required no hang, no retry and no watchdog kill across the two hundred steps that were to determine whether the new routing stays, and the run has now produced at least the first two of the three, forty-two steps in. The go/no-go the gate named, about 4:45 p.m. (ET), has passed with the run restarting rather than stepping.
infraep
mcwitt #8870
infraepkernels
hero-ragged_a2a-ep-step81k
The replacement run has stopped taking steps, and the way it stopped matches the one fault the trial was gated against. Its last step, 81,757, was logged at 3:59 p.m. (ET). Twenty-seven minutes later nothing has followed, against a measured pace of one step every 18.3 seconds, so roughly ninety steps are missing. Nothing has failed in the ordinary sense: W&B holds the run in the running state, its heartbeat is current to within a minute, and its GPU memory stays pinned at 89.9 per cent, so the processes are alive and the model is still resident. What has stopped is the work. The GPUs report full utilisation while their memory subsystem sits idle — DRAM activity fell from about 23 per cent to under a thousandth of a per cent at 3:59:50 p.m. (ET) and has not moved since — and NVLink traffic between them has been zero over the same span. A GPU that is fully busy while moving no data is spinning inside a collective exchange that never completes. The [September 3 attempt](https://wandb.ai/marin-community/marin_moe/runs/hero-ragged_a2a-ep-step54k) on this same routing backend entered the identical state, full utilisation with no memory or link traffic, and never left it; that attempt reached about seventy steps of its window, and this one reached forty-two of two hundred. Every gated measure was passing up to the stop. One gate was that the run must not hang, so on its own terms the trial has now failed, and the [plan](https://github.com/marin-community/marin/issues/8506#issuecomment-5606002395) calls for an immediate rollback: relaunch `trigger_hero.sh` at commit `0908920c86`, which resumes the run it replaced from that tree's own newest checkpoint, saved at step 81,900 at 3:28 p.m. (ET). The root cause of [#8870](https://github.com/marin-community/marin/issues/8870) is still open, and its strongest lead — a fabric fault on the racks that land at replica indices 8 to 10 — is one the trial could not design around. Nothing has been posted to the [status log](https://github.com/marin-community/marin/issues/8506) since the 3:49 p.m. (ET) first-steps note, which set the decision point at about 4:45 p.m. (ET).
infraep
mcwitt #8506
infraep
mcwitt #8506
infraep
mcwitt hero-ragged_a2a-ep-step81k
The replacement hero run started at 3:37 p.m. (ET), four minutes after the old one stopped. It restores the checkpoint [forced and pinned](https://github.com/marin-community/marin/issues/8506#issuecomment-5606812293) at step 81,716 — the last of the four checkpoint paths it is given points into the old run's storage at that step — and it holds the same shape of machine: eleven racks, 176 tasks of four GB200s each, the same data mixture, the same optimizer, and the same 390,251-step finish line. Two things differ, and they are the two the cutover was for. It uses the [routing backend](https://github.com/marin-community/marin/pull/9043) withdrawn on September 3, which moves tokens between experts in one ragged exchange instead of fixed waves, and it keeps its master copy of the weights in full precision on the GPU. It is also built against the [upgraded kernels](https://github.com/marin-community/marin/pull/8959): the package list W&B stores with the run names QuACK 0.6.4, where the old run's names 0.6.1. So both changes are in the image that launched, whatever the pull requests themselves say. The run has logged no training step yet, eight minutes in. That is normal rather than a warning — the [September 3 attempt](https://wandb.ai/marin-community/marin_moe/runs/hero-ragged_a2a-ep-step54k) on this same backend took about thirty-four minutes from launch to its first step. W&B holds it in the running state and its heartbeat is current. The [trial report](https://wandb.ai/marin-community/marin_moe/reports/ep_change:-hero-ragged-a2a-swap,-200-step-trial-%28step-81716%29--VmlldzoxNzkwMjQyNQ==) that will place it beside the old run over the window is live.
infraep
mcwitt #8506
infraep
mcwitt hero-wd-gate-router-p02-step58k
infraep
mcwitt 8959-expert-capture
infrakernels
mcwitt 8959-gradient-control
The frozen-weight gradient check is finished, and the two kernel versions produce the same arithmetic. The [control arm](https://wandb.ai/marin-community/marin_moe/runs/8959-gradient-control-20260909) stopped at 2:51 p.m. (ET) at step 80,920, the last of the nine steps it was set to take, matching the [upgraded arm](https://wandb.ai/marin-community/marin_moe/runs/8959-gradient-treatment-20260909) step for step from the same step-80912 checkpoint with every weight held fixed. Their stored settings differ only in name and storage path. Which kernel version each job carries — the thing this page could not confirm earlier today — is in fact recorded, in the package list W&B stores with every run: the upgraded arm was built against QuACK 0.6.4 and the control against 0.6.1, the version the hero itself is running. On loss the arms are identical: the same value at all nine steps, to every digit W&B stores. On gradients they are close enough to settle the question. The total gradient norm differs between the arms by two parts in ten thousand on average and by ten parts at worst, against a spread of twelve per cent across the nine steps themselves. Taken group by group — the model has 1,380 of them, so 11,979 comparisons in all — a little over half agree within a thousandth and more than nine in ten within a hundredth. The widest gap is seven per cent and only one other exceeds five; both fall on parameter groups whose norms are a hundredth or less of the total, where the arithmetic has the fewest digits to work with. Neither arm sits above the other. This is what the morning's matched pair could not show: those two started training, so their weights parted after the first step and a kernel difference could not be told from ordinary drift. Speed cannot be read from these arms, because each step writes 1,380 norms and so takes about three times the morning pair's. The [reference arm](https://wandb.ai/marin-community/marin_moe/runs/8959-gradient-reference-20260909) that started between the two is now in the crashed state, having logged no step. mcwitt has posted none of this to [#8959](https://github.com/marin-community/marin/pull/8959), which is still a draft. The hero trained through all of it on the older kernels, on its own eleven racks, and W&B holds it in the running state.
infrakernels
mcwitt 8959-gradient-control
infrakernels
mcwitt #8506
The cutover has started: the checkpoint the replacement run would begin from has been written and pinned. At 2:26 p.m. (ET) mcwitt forced a save on the live run through the training-control endpoint, taking step 81,716 out of turn rather than waiting for the next scheduled permanent one. It committed two minutes later — 5.36 TB across 26,795 objects — and at 2:30 p.m. (ET) its metadata was rewritten so the hourly cleanup treats it as permanent and cannot delete it. This is the step the [plan](https://github.com/marin-community/marin/issues/8506#issuecomment-5606002395) posted an hour earlier called for, and it costs about one step of training. The old run keeps going from there: it now trains the two hundred steps that serve as the matched control, ending at step 81,916, which mcwitt put near 3:35 p.m. (ET). The launcher change is not a separate pull request after all — it went into [#9043](https://github.com/marin-community/marin/pull/9043), which now points the trigger script at this checkpoint under the run id `hero-ragged_a2a-ep-step81k`. So the cutover is blocked on two merges and nothing else: [#8959](https://github.com/marin-community/marin/pull/8959) and [#9043](https://github.com/marin-community/marin/pull/9043) both have to land on the main branch. mcwitt also linked the [report](https://wandb.ai/marin-community/marin_moe/reports/ep_change:-hero-ragged-a2a-swap,-200-step-trial-%28step-81716%29--VmlldzoxNzkwMjQyNQ==) that will show the old and new runs side by side over the window.
infracheckpointing
mcwitt 8959-gradient-treatment
infrakernels
mcwitt #8959
kernels
mcwitt #8506
Sets out a plan to stop this run today and relaunch it on the expert-routing backend that was withdrawn a week ago. At 1:26 p.m. (ET) mcwitt posted the sequence to the status log. Merge the kernel-library upgrade ([#8959](https://github.com/marin-community/marin/pull/8959)), then the backend change ([#9043](https://github.com/marin-community/marin/pull/9043)). Force a checkpoint on the live run through the training-control endpoint about seventy minutes before the cutover and mark it so the hourly cleanup cannot delete it. Open a launcher pull request pointing the trigger script at that checkpoint under a new run id. Let the current run train two hundred steps past it, so the old and new runs can be read against each other on the same batches. Then cancel the old coordinator and launch from a clean checkout. Waiting for the next scheduled permanent checkpoint was rejected: it lands about 3 a.m. (ET) tomorrow and the one after on Thursday, while forcing a checkpoint costs a single step, which is what the September 4 pause cost. Beyond the backend, the new run would carry one train-step executable per rank ([#8911](https://github.com/marin-community/marin/pull/8911)), the rebuilt routing kernel wheel ([#8976](https://github.com/marin-community/marin/pull/8976)), and two fixes to attention and the loss calculation ([#8960](https://github.com/marin-community/marin/pull/8960), [#8961](https://github.com/marin-community/marin/pull/8961)). The trial passes only if the new run's loss matches the old one within the noise of its own arithmetic once a known small offset is allowed for, dropped tokens stay below a thousandth of assignments, speed is at least the old run's, and nothing hangs, retries or trips the progress watchdog; anything short of that rolls back at once to the commit the run is on now, which resumes its own newest checkpoint. mcwitt says the runbook scripts are being dry-run read-only against the cluster, and that a W&B report and each transition — forced checkpoint, cancel, launch, first steps, go/no-go — will be posted here before launch. The reason the backend was withdrawn on September 3, two silent hangs on one rack ([#8870](https://github.com/marin-community/marin/issues/8870)), is still unexplained; the plan names a fabric fault on the racks that land at three particular positions as the strongest lead and says the trial cannot avoid them.
infraincidentthroughput
mcwitt #9043 open
Puts up the change that would return this run to the way of moving tokens between experts that it lost on September 3. mcwitt opened it at 1:25 p.m. (ET). It makes the ragged all-to-all transport the hero's default again, with fp32 master weights held on the GPU, undoing the fallback taken in [#8884](https://github.com/marin-community/marin/pull/8884) after two attempts hung silently on one rack at eleven racks ([#8870](https://github.com/marin-community/marin/issues/8870)). The case for trying again is three changes since: one train-step executable per rank removed a duplicate registration ([#8911](https://github.com/marin-community/marin/pull/8911)), the kernel wheel moved to a rebuilt routing kernel ([#8976](https://github.com/marin-community/marin/pull/8976)), and the kernel-library upgrade replaces the scheduler that could leave a multiply's output unwritten under contention ([#8959](https://github.com/marin-community/marin/pull/8959), which merges first). mcwitt writes plainly that the hang's cause is still open, and names a fabric fault on particular NVL72 racks as the strongest lead — the same pattern as the September 5 NVLink failure on this run ([#8934](https://github.com/marin-community/marin/issues/8934)), which the cluster operator has not yet answered. One step is one-way: this run's checkpoints keep their fp32 master copy in host memory and the new mode reads them in and moves it to the GPU, but the reverse conversion is refused, so a rollback has to resume this run's own checkpoint tree rather than anything the replacement writes. Automated review has failed on the pull request twice, each time reporting that the commit it was asked to read does not exist.
infrathroughput
mcwitt 8959-hero-treatment
The matched benchmark that #8959 names as its own merge condition has finished, and the two arms came out the same on both halves of it. The treatment arm stopped at 11:21 a.m. (ET) at step 81,011 — the control's own stop, from the control's own step-80912 checkpoint, after the same 339,788,955,648 tokens. Both arms logged the same 100 steps, 80,912 through 81,011. On loss they are the same to the fourth decimal: the first step after the restore is identical on both at 1.234894, and across the 100 steps the mean absolute difference is 0.0002, against a median move of 0.0205 from one step to the next — a hundredth of the curve's own movement. It does not grow: the treatment sits 0.00017 above the control on average over the first fifty steps and 0.00004 below over the second fifty, and the largest single gap, 0.00075, falls two steps from the end. That is what reordered floating-point arithmetic produces, not a change in how the model learns. On speed the arms are level: the median interval between logged steps is 16.14 seconds on the control and 16.17 on the treatment, and the whole span from first to last logged step is 1,607.5 seconds against 1,618.3. The treatment's slow tail is a little longer — 90th-percentile step 17.08 seconds against 16.56 — which is the same few-slow-steps difference this page reported when the arm was a third of the way through. So the evidence [#8959](https://github.com/marin-community/marin/pull/8959) asks for is in hand and it is clean. mcwitt has posted nothing to the pull request, which is still in draft, so the merge and the restart that would carry the upgrade into the hero stay held. A second treatment arm, [8959-d768-treatment](https://wandb.ai/marin-community/marin_moe/runs/8959-d768-treatment-20260909), started a minute later on a 1.6-billion-parameter model — a small shape, not the hero — and has logged no step yet. The hero kept training on its own eleven racks throughout.
infrakernels
hero-wd-gate-router-p02-step58k
evals
mcwitt rdx-split-latest-a
mcwitt rdx-cudnn-e
epablation
mcwitt rdx-cudnn-d
A second cudnn arm reached the control's stop and landed on the other side of it, so the pair is level — and the 5% slowdown this page reported for the first arm was the cost of a profiler, not of the code being tested. rdx-cudnn-d started at 5:35 a.m. (ET) and finished clean at 5:57 a.m., at step 54,059 after 226,744,074,240 tokens: the same stop and the same tokens as rdx-ctl-a and rdx-cudnn-b. Final train losses are 1.27988 for the control, 1.27996 for rdx-cudnn-b and 1.27983 for rdx-cudnn-d, so the two cudnn arms sit 0.00008 above and 0.00005 below the control — a spread of 0.00013 on a loss near 1.28. Wall clock is 1,390.7 seconds for the control, 1,462.5 for rdx-cudnn-b and 1,341.0 for rdx-cudnn-d. W&B's stored configuration explains that split: rdx-cudnn-b runs the profiler — which records where a step's time goes, and costs time to do it — for three steps from step 54,003, and neither the control nor rdx-cudnn-d does. Correcting what this page said earlier: the record does say what the cudnn name marks. Every cudnn job installs three packages the control does not — a custom build of the JAX GPU plugin, taken from a branch of mcwitt's named ragged-dot after the grouped matrix multiply this model's expert layers use, plus NVIDIA's cuDNN and cuBLAS libraries pinned to versions 9.25.1.1 and 13.6.1.10 — while the control installs nothing extra and runs the deployed image. Two builds of that plugin appear: one in rdx-cudnn-a and -b, another in -c, -d and -e. The control also records an optimiser option, use_syrk, that no cudnn job records at all, which is what a different build would produce. Two of the remaining cudnn jobs failed: rdx-cudnn-a crashed after eight minutes, and rdx-cudnn-c, still running at the last read, ended in W&B's crashed state having logged no training step at all. rdx-cudnn-e started at 5:59 a.m. (ET) with the profiler on and is still running. mcwitt has posted nothing about any of these to the status log, to [#8870](https://github.com/marin-community/marin/issues/8870) or to Discord. The hero trained through all of it on its own eleven racks.
epablation
markhart0034 #hero-run-2026
evals
mcwitt rdx-cudnn-b
ablationep
Percy Liang #hero-run-2026
evals
mcwitt rdx-ctl-a
mcwitt put the expert-routing backend that was rolled back on September 3 back on a machine, and it ran clean. A job named rdx-ctl-a started at 2:20 a.m. (ET), one minute after the last kernel benchmark released the rack, and finished at 2:43 a.m. It holds the whole 535-billion-parameter model on 64 GPUs — one rack, against the eleven the hero occupies — restores the same step-54000 checkpoint the September 2 backend swap started from, and carries the ragged all-to-all expert transport — the code path that moves each token to the expert chosen for it — that hung twice on that checkpoint. Its optimiser settings are the ones that checkpoint was trained with, without the gate and router weight decay the hero added on September 4, so this tests the transport rather than the recipe. It was configured to stop after 60 steps, and it stopped at step 54,059 in W&B's finished state, holding about 24.8% model FLOPs utilization throughout. That is not evidence the hang is fixed: both hangs took eleven racks, and the standup already records that later one-rack and eleven-rack attempts failed to reproduce them ([#8970](https://github.com/marin-community/marin/issues/8970)). The name marks this a control arm; no treatment run to compare it against has appeared, and nobody has written about it in the issues or in Discord. Every job on this rack since 5:52 p.m. (ET) on September 8, this one included, places its first task on s1wvxs64 — the node whose fabric link went down on September 7 and stopped the hero ([#8934](https://github.com/marin-community/marin/issues/8934)). The node is back in service and has carried about nine hours of work since. The hero was not touched.
epablation
mcwitt 8870-quack-main-static-r1
The overnight kernel-upgrade test ran its third arm, and on matched steps the three builds are the same on both halves of the gate. The arm is labelled static — QuACK 0.6.4, the kernel-library version #8959 upgrades to, run with a fixed work schedule rather than the dynamic one whose fix [#8959](https://github.com/marin-community/marin/pull/8959) is for. It started at 1:30 a.m. (ET) from the same frozen hero checkpoint at step 77,846 as the pair before it and stopped at 2:19 a.m. All three arms read the same batches: at step 77,850 their losses agree to within 0.00003. At step 77,927, the last step all three reached, the deployed build recorded 1.28414, the dynamic upgrade 1.28472 and the static variant 1.28478 — a spread of 0.0006 on a loss near 1.28. The speed half now reads the same way. Each arm logged its first step at a different point — 421, 377 and 375 seconds in, which is queueing, loading and compilation — and the 81 steps after that took 1,337, 1,336 and 1,334 seconds. So the 2.5% gap this page reported for the earlier pair is start-up time; the training itself differs by 0.2%. Separately, the night's failures are failures. Twelve jobs ran between 5:52 p.m. (ET) on September 8 and 2:19 a.m. on September 9. Ten of them trained; the other two logged nothing past the checkpoint they restored. Every training job but the two profiling ones was written to stop at step 78,846, a thousand steps past that checkpoint. The profiling pair, set to stop at 77,864, reached its stop and ended in W&B's finished state. The other eight ended in the crashed state between steps 77,863 and 78,004, none of them a sixth of the way through its planned thousand steps. Nothing on the record explains why. mcwitt has posted no result to #8959 or to [#8870](https://github.com/marin-community/marin/issues/8870), so the merge he gated on these tests is still held. The hero trained through the whole window on its own eleven racks.
throughputablation
mcwitt 8870-quack-main-r1
throughputablation
mcwitt 8870-quack-main-upgrade-r1
mcwitt ran the tests that the kernel-library upgrade waits on, and on loss the two builds came out level. W&B records eight jobs between 8:27 p.m. (ET) on September 8 and 12:49 a.m. (ET) on September 9. Seven are training jobs of ten to thirty-three minutes; the eighth is a three-second record holding no steps. Every training job carries the hero's own parameter count and stops somewhere between steps 77,863 and 77,928 — the range the hero itself passed through on September 8 — which matches the plan mcwitt set in [#8959](https://github.com/marin-community/marin/pull/8959): matched pairs restored from one frozen hero checkpoint, so the test trains the production model rather than a small stand-in. Each job is labelled as the build now deployed or as QuACK 0.6.4, the version whose scheduler fix that pull request is for. The pair that compares cleanly is the one run with profiling on, which records where a step's time goes: both arms stopped on step 77,863, and their train losses differ by 0.00002. That answers the loss half of the gate mcwitt set on the merge. The speed half is not answered here — profiling costs time of its own, and no matched timing has been written down anywhere. Four of the seven training jobs end in W&B's crashed state and two in finished; nothing on the record says which of those endings are faults and which are how a fixed-length benchmark stops. The last job, the upgrade arm of a fresh pair started at 12:49 a.m. (ET), was still running when this page read W&B. mcwitt has posted no result to #8959 or to [#8870](https://github.com/marin-community/marin/issues/8870), so the gate is not closed. The hero kept training through the whole window.
throughputablation
Larry #hero-run-2026
evals
ClassicLarry #8827
evals
Sep 8
Larry #hero-run-2026
evals
ihodes #8970
planning
hero-wd-gate-router-p02-step58k
evals
dlwh #hero-run-2026
Larry #hero-run-2026
Larry #hero-run-2026
Larry #hero-run-2026
Larry #hero-run-2026
evals
mcwitt #8959 open
incidentthroughput
mcwitt #code-review
evals
Kaiyue-Wen #hero-run-2026
evals
dlwh #hero-run-2026
evals
Kaiyue-Wen #hero-run-2026
evals
Larry #hero-run-2026
evals
Larry #hero-run-2026
evalsscaling-law
Kaiyue-Wen #hero-run-2026
evals
dlwh #hero-run-2026
evals
hero-wd-gate-router-p02-step58k
evals
Larry #hero-run-2026
evals
Sep 7
hero-wd-gate-router-p02-step58k
evals
markhart0034 #hero-run-2026
evalstraining-dynamics
hero-wd-gate-router-p02-step58k
incident
loom-oa-dev #8934
Names the cause of the morning's stop and reports the recovery blocked at the cluster operator. At 10:01:27Z task 16, on node s1wvxs64 in rack dh1-394-14, logged Xid 149 NETIR_LINK_DOWN and Xid 154 Drain and Reset — the node's link into the GPU fabric went down and the driver reset the device. Fifteen retained disruption snapshots from rack 394 mark both fault conditions and a production reboot, and carry a reservation taint and a separate evict taint. Node-local agents came back on seventeen rack-394 nodes between 10:20:07Z and 10:28:49Z, but host reboot completion is unverified. Iris deleted the failed task and requeued the 176-task coscheduled group; at 10:55Z every task was still attempt 2 BUILDING with no start timestamp, and Kueue — which admits the job — could fit only nine of the run's eleven slices, each of which needs sixteen GB200 nodes in one NVL72 rack, because seventeen nodes still held the reservation. The latest confirmed complete resume checkpoint is step 72183, saved at 09:10:32Z; attempt 1 had reached step 72344, so a restart replays about 160 steps. Attempt 2 has not logged a checkpoint load or an optimizer step, so restore and resumption are both unverified. The asks to CoreWeave are to identify the Xid-149 source, confirm which seventeen nodes are reserved and whether each completed its reboot and health checks, finish the provider-side recovery, and advise whether to quarantine rack 394 from this run. Cluster permissions denied the investigation direct node reads, so boot-ID changes and exact taint membership are unknown; it made no changes to Iris or Kubernetes.
incident
hero-wd-gate-router-p02-step58k
incident
Larry Dial hero-wd-twitter-pertoken-72k
evals
Sep 5
mcwitt #8870
incident
mcwitt #8870
incident
mcwitt #8934
incident
Helw150 #8824
evals
loom-oa-dev #8506
incident
hero-wd-gate-router-p02-step58k
incident
Helw150 #8824
evals
mcwitt #8870
incident
mcwitt #8870
incident
Sep 4
mcwitt #8870
incident
Mooler0410 #8818
training-dynamics
ClassicLarry #8818
training-dynamics
mcwitt #8925
incident
mcwitt #8912
incident
mcwitt #8911 open
evalsincident
mcwitt #8870
incident
mcwitt #8870
incident
rjpower #8896 open
incidentthroughput
Helw150 #8824
evals
mcwitt #8506
incidentcheckpointing
hero-wd-gate-router-p02-step58k
evals
Sep 3
mcwitt #8890 merged
checkpointing
mcwitt #8884 merged
throughput
Mooler0410 #8818
training-dynamics
hero-12d8b6f0-dee637
incident
mcwitt #8506
incident
hero-12d8b6f0-dee637
incident
mcwitt #8506
incident
mcwitt #8870
incident
mcwitt #8861
incidentevals
mcwitt #8506
incident
Sep 2
ClassicLarry #8868
checkpointing
mcwitt #8868 open
checkpointing
hero-ragged_a2a-ep-step54k
incidentthroughput
hero-12d8b6f0-dee637
incident
ClassicLarry #8833
optimizer
mcwitt #8862 open
evals
mcwitt #8861
evals
Larry #hero-run-2026
training-dynamics
Kaiyue-Wen #hero-run-2026
optimizer
Larry #hero-run-2026
training-dynamics
Kaiyue-Wen #hero-run-2026
training-dynamics
mcwitt #8859 open
evals
Larry #code-review
checkpointing
ClassicLarry #8854 open
checkpointing
ClassicLarry #8818
optimizer
Larry #hero-run-2026
training-dynamics
Larry #hero-run-2026
training-dynamics
Percy Liang #hero-run-2026
training-dynamics
Larry #hero-run-2026
optimizer
Kaiyue-Wen #hero-run-2026
optimizer
ClassicLarry #8833
optimizer
ClassicLarry #8818
training-dynamics
ClassicLarry #8818
ablation
Larry #hero-run-2026
ClassicLarry #8833
optimizer
ClassicLarry #8818
ablation
loom-oa-dev #8506
incident
Sep 1
hero-12d8b6f0-dee637
incident
ClassicLarry #8833 open
optimizer
Larry #hero-run-2026
ClassicLarry #8818
ClassicLarry #8818
ablation
Helw150 #8824
percyliang #8824
evals
ClassicLarry #8818
ClassicLarry #8827
ClassicLarry #8818
training-dynamics
ClassicLarry #8818
ClassicLarry #8818
ablation
Helw150 #8824
evals
Aug 31
ClassicLarry #8818
ablation
Helw150 #8435
data
ClassicLarry #8818
ClassicLarry #8818
training-dynamics
Aug 28
mcwitt #8754
long-context
Aug 27
loom-oa-dev #8506
incident
Aug 25
mcwitt #8684 open
throughput
Aug 24
rjpower #8506
ablation
ravwojdyla #8506
incident
rjpower #8506
Aug 23
rjpower #8506
incident
Aug 22
rjpower #8506
Aug 21
rjpower #8506
muchanem #8506
incident
Aug 20
rjpower #8506
incident
ravwojdyla-agent #8480 merged
Aug 19
ClassicLarry #8435
spec