4-hour window · 56 points · 58.3% coverage
- gemma4:31b is the strongest model at 155.75 token/s average throughput, peaking at 189.38 token/s, while nemotron-3-ultra is the weakest at 9.56 token/s average and just 4.81 token/s in its latest sample.
- glm-5.3 shows the most operationally significant volatility: it fell from 195.89 token/s at 00:20 to roughly 60-69 token/s across 00:40-01:40 before rebounding to 171.36 token/s, a 50.9% coefficient of variation. deepseek-v4-flash similarly swung between 209.95 and 53.9 token/s.
- Every model recorded only 7 of 12 expected samples, yielding 58.3% coverage and 56 valid points overall, so the 22:03-23:40 UTC window is absent and averages may not represent the full four hours.