← Performance dashboard

Hourly performance insights

Summaries of rolling four-hour performance data

1202 retained summaries
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model, averaging 153.12 token/s with a p95 of 181.52 token/s, while nemotron-3-ultra is the weakest, averaging 18.10 token/s with a p95 of only 65.41 token/s.
  • The most operationally significant volatility is nemotron-3-ultra's coefficient of variation of 111.9%, swinging between 1.38 and 83.15 token/s, including a near-hour stretch below 2 token/s from roughly 01:05 to 01:30; glm-5.3 also showed a sharp single-interval dip to 32.04 token/s at 02:00.
  • No missing-data limitation applies: all eight models report 48 of 48 expected samples at 100.0% coverage, matching the dataset's 384 valid points.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model with an average throughput of 155.11 token/s (p95 181.42 token/s), while nemotron-3-ultra is the weakest at 19.02 token/s average, roughly eight times slower.
  • The most operationally significant volatility is nemotron-3-ultra's coefficient of variation of 106.7%, with throughput collapsing to around 1.4 token/s for roughly 30 minutes (01:05–01:30) before spiking to 79.64 token/s at 01:45; minimax-m3 also shows a -21.6% trend, ending at 34.86 token/s.
  • No missing-data limitation: all eight models report 48 of 48 expected samples, 100.0% coverage, and 384 valid points across the four-hour window.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model with an average throughput of 144.70 token/s (p95 181.42 token/s), while nemotron-3-ultra is the weakest at 22.13 token/s average, peaking at only 83.36 token/s.
  • The most operationally significant volatility is nemotron-3-ultra: it swings between 1.25 and 83.36 token/s with a 104.0% coefficient of variation and a -43.7% trend, including a sustained stretch near 1.4 token/s from 01:05 to 01:30 UTC.
  • No missing-data limitation exists: all eight models report 48 of 48 expected samples, 100.0% coverage, and 384 valid points across the four-hour window.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model by average throughput at 137.96 token/s (p95 179.95 token/s), while nemotron-3-ultra is the weakest at 22.89 token/s average, never exceeding 83.36 token/s.
  • The most operationally significant volatility comes from nemotron-3-ultra: 87.3% coefficient of variation, a drop to 1.25 token/s at 23:05, and a period close at 2.8 token/s with a -22.1% trend; glm-5.3 also shows sharp isolated dips to 19.9 token/s at 22:15 and 20.86 token/s at 22:45 despite its high average.
  • No missing-data limitation applies in this window: all eight models report 48 of 48 expected samples, 100.0% coverage each, and 384 valid points overall.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model at 137.3 token/s average throughput (peaking at 175.89 token/s), while nemotron-3-ultra is the weakest at 23.23 token/s average, ranging from 1.25 to 83.36 token/s.
  • The most operationally significant volatility is nemotron-3-ultra, with a coefficient of variation of 85.7% and a +55.9% trend; it collapsed to 1.25–1.92 token/s between 22:50 and 23:05 UTC. glm-5.3 also showed repeated single-interval drops to roughly 20 token/s at 21:15, 22:15, 22:30, and 22:45 UTC.
  • No missing-data limitation applies: all eight models delivered 48 of 48 expected samples, 100.0% coverage, and 384 valid points, so throughput dips reflect observed values rather than collection gaps.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model with an average throughput of 134.59 token/s (p95 172.74 token/s), while nemotron-3-ultra is the weakest at 17.48 token/s average, far below the next-lowest glm-5.2 at 25.04 token/s.
  • The most operationally significant volatility is glm-5.3's repeated collapses from roughly 150 token/s to near 20 token/s at 21:15, 22:15, 22:30, and 22:45, alongside nemotron-3-ultra's extreme swings (cv 107.0%) ending at 1.85 token/s, its minimum for the period.
  • No missing-data limitation exists: all eight models report 48 of 48 expected samples, 384 valid points total, and 100.0% coverage across the four-hour window.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model with average throughput of 138.84 token/s (p95 171.43 token/s); nemotron-3-ultra is the weakest at 12.98 token/s average, with a maximum of just 75.06 token/s.
  • The most operationally significant volatility is nemotron-3-ultra, whose coefficient of variation is 90.7% and throughput ranges from 4.14 to 75.06 token/s, including a 75.06 token/s spike at 21:50; gemma4:31b also swings between 24.87 and 150.42 token/s (39.7% CV).
  • There is no missing-data limitation: all eight models report 48 of 48 expected samples, 100.0% coverage each, and 384 valid points overall, so no gaps affect these figures.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model at 136.44 token/s average throughput, peaking at 182.28 token/s; nemotron-3-ultra is the weakest at 8.04 token/s average, never exceeding 24.22 token/s.
  • gemma4:31b shows the most operationally significant volatility, ranging from 7.4 to 138.1 token/s with a 59.1% coefficient of variation and a 97.3% upward trend over the window, while deepseek-v4-flash spiked to 179.03 token/s at 19:20.
  • No missing-data limitation exists: all eight models report 48 of 48 expected samples, 384 valid points, and 100.0% coverage across the four-hour window.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model at 127.74 token/s average throughput, ahead of deepseek-v4-flash at 114.2 token/s; nemotron-3-ultra is the weakest at 7.42 token/s average.
  • gemma4:31b shows the most operationally significant volatility: throughput ranged from 7.4 to 124.82 token/s with a 72.9% coefficient of variation, and its +208.1% trend reflects a climb from sub-10 token/s lows near 17:10 to peaks above 108 token/s after 19:15.
  • No missing-data limitation applies: all eight models reported 48 of 48 expected samples, 384 valid points total, and 100.0% coverage across the four-hour window.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model at 119.89 token/s average throughput, ahead of deepseek-v4-flash at 104.96 token/s; nemotron-3-ultra is the weakest at 6.16 token/s average, with a p95 of only 13.17 token/s.
  • The most operationally significant movement is glm-5.3-flash's decline, trending -36.3% from roughly 89-118 token/s early in the window to roughly 50-60 token/s later; gemma4:31b also shows high volatility (cv 66.5%), dipping to 7.4 token/s near 17:10 before peaking at 124.82 token/s at 18:45.
  • No missing-data limitation applies: every model reports 48 of 48 expected samples, 384 valid points overall, and 100.0% coverage across the four-hour window.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 posted the strongest average throughput at 118.64 token/s, peaking at 167.33 token/s, while nemotron-3-ultra was the weakest at 4.68 token/s average and never exceeded 42.93 token/s.
  • The most operationally significant movement is gemma4:31b's 51% decline, from a 98.0 token/s peak at 15:45 down to readings around 7.4 token/s near 17:10; nemotron-3-ultra also showed extreme volatility, spiking to 42.93 token/s at 16:55 against its 4.68 token/s average.
  • No missing-data limitation applies: all eight models report 48 of 48 expected samples at 100.0% coverage, so the 384 valid points fully represent the four-hour window.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model by average throughput at 114.59 token/s, with a maximum of 167.33 token/s; nemotron-3-ultra is the weakest at 4.32 token/s average, a p95 of only 10.26 token/s, and a floor of 0.44 token/s.
  • The most operationally significant volatility is nemotron-3-ultra's coefficient of variation of 143.1%, including a drop to 0.44 token/s at 16:25 and a spike to 42.93 token/s at 16:55; gemma4:31b also swung between 7.61 and 98.0 token/s (cv 53.2%).
  • No missing-data limitation applies: all eight models report 48 of 48 expected samples with 100.0% coverage, matching the 384 valid points, so the four-hour window is fully observed.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model at 110.11 token/s average throughput, peaking at 167.33 token/s; nemotron-3-ultra is the weakest at 2.95 token/s average, never exceeding 12.23 token/s.
  • gemma4:31b shows the most significant trend, climbing 95.8% from an 8.1 token/s floor around 13:05-13:10 to a 98.0 token/s peak at 15:45; glm-5.3 is the most volatile high-throughput model, swinging between 10.98 and 167.33 token/s with a 33.6% coefficient of variation.
  • No missing-data limitation applies: all eight models report 48 of 48 expected samples, 384 valid points, and 100.0% coverage across the four-hour window.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model with an average throughput of 109.06 token/s (p95 149.91 token/s), while nemotron-3-ultra is the weakest at 3.83 token/s average, peaking at only 12.23 token/s.
  • gemma4:31b shows the highest volatility, with a coefficient of variation of 59.9% and throughput ranging from 7.62 to 63.54 token/s; glm-5.3 also swings sharply, dropping to 10.98 token/s at 12:40 after peaking at 166.18 token/s at 11:25.
  • No missing-data limitation exists: all eight models have 48 of 48 expected samples, 384 valid points, and 100.0% coverage, though the four-hour window limits longer-term assessment.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model at 113.92 token/s average (p95 172.74 token/s), while nemotron-3-ultra is weakest at 5.15 token/s average with a maximum of only 12.79 token/s.
  • The steepest operational movement is nemotron-3-ultra's decline (trend -46.2%), from a 12.79 token/s peak at 10:55 to 0.98 token/s at 12:30, closing at 3.76 token/s; glm-5.3 is also volatile, swinging between 10.98 and 190.39 token/s (cv 37.7%).
  • No missing-data limitation: all eight models report 48 of 48 samples at 100.0% coverage (384 valid points), so results are constrained only by the four-hour window ending 14:03 UTC.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model at 112.81 token/s average throughput (p95 172.74 token/s), while nemotron-3-ultra is the weakest at 6.45 token/s average, peaking at only 14.55 token/s.
  • The most significant trend is gemma4:31b's decline of 62.4%, falling from roughly 60 token/s at 09:05 to 9.06 token/s at 13:00; nemotron-3-ultra also sagged to 0.98 token/s near 12:30. glm-5.3 is volatile (cv 40.4%), swinging between 10.98 and 190.39 token/s.
  • No missing-data limitation applies: all eight models report 48 of 48 expected samples, 384 valid points, and 100.0% coverage, so averages are computed over complete four-hour windows.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model by average throughput at 113.41 token/s, peaking at 190.39 token/s; nemotron-3-ultra is the weakest at 8.70 token/s average, with a p95 of only 14.45 token/s.
  • gemma4:31b shows the sharpest deterioration, trending down 54.3% from 42.43 token/s at 08:05 to a low of 7.62 token/s at 11:25, while nemotron-3-ultra is the most volatile, with a 116.9% coefficient of variation and swings between 1.31 and 66.14 token/s.
  • No missing-data limitation applies: all eight models report 48 of 48 expected samples, 100.0% coverage each, and 384 valid points across the four-hour window.
4-hour window · 384 points · 100.0% coverage
  • Strongest average throughput was glm-5.3 at 109.96 token/s, followed by deepseek-v4-pro at 97.26 token/s; weakest was nemotron-3-ultra at 12.29 token/s, with glm-5.2 also low at 20.17 token/s.
  • The most operationally significant volatility is nemotron-3-ultra, which ranged from 66.14 token/s down to 1.31 token/s with a coefficient of variation of 117.6% and a -47.9% trend; glm-5.3 also showed sharp dips to 6.01 token/s between sustained stretches above 130 token/s.
  • No missing data: all eight models delivered 48 of 48 expected samples, 384 valid points total, at 100.0% coverage; the limitation is the four-hour window, which may not represent other periods.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 posted the highest average throughput at 105.09 token/s (peak 151.70 token/s), while nemotron-3-ultra was the weakest at 16.31 token/s average, ending the window at just 6.27 token/s.
  • glm-5.3 was also the most volatile strong performer, swinging between 6.01 and 151.70 token/s (CV 42.0%), and nemotron-3-ultra showed the sharpest deterioration, trending down 51.2% with sustained readings near 2 token/s between 07:50 and 09:00 UTC.
  • No missing-data limitation applies: all eight models recorded 48 of 48 expected samples, totaling 384 valid points at 100.0% coverage across the four-hour window.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model with an average throughput of 121.05 token/s and a peak of 174.59 token/s; nemotron-3-ultra is the weakest at 18.04 token/s average, with a minimum of 1.31 token/s.
  • glm-5.3 shows the most operationally significant volatility, swinging between 174.59 and 6.01 token/s, including drops to 38.25 and 28.70 token/s around 07:15–07:25 and to 27.69 and 6.83 token/s at 08:30–08:35; nemotron-3-ultra is also unstable, with a 113.4% coefficient of variation.
  • No missing-data limitation applies: all eight models report 48 of 48 samples each, 384 valid points total, and 100.0% coverage across the four-hour window.
4-hour window · 376 points · 97.9% coverage
  • glm-5.3 is the strongest model by average throughput at 134.51 token/s (p95 173.75 token/s), while glm-5.2 is the weakest at 21.98 token/s average, peaking at only 45.72 token/s.
  • The most operationally significant movement is gemma4:31b's decline: its average of 64.99 token/s masks a -54.0% trend, falling from roughly 110-144 token/s early in the window to sustained readings near 18-40 token/s after 06:45, with 52.0% coefficient of variation. nemotron-3-ultra is even more volatile at 90.2% CV, ranging from 1.67 to 102.21 token/s.
  • Every model recorded 47 of 48 expected samples (97.9% coverage), leaving one five-minute gap per model and 376 valid points overall, so short-lived drops may be underrepresented.
4-hour window · 383 points · 99.7% coverage
  • glm-5.3 is the strongest model with an average throughput of 138.16 token/s (p95 174.37 token/s), while nemotron-3-ultra is the weakest at 19.82 token/s average, far below the next-lowest glm-5.2 at 36.0 token/s.
  • glm-5.2 shows the most operationally significant decline, trending down 54.0% from an early peak of 177.85 token/s at 02:40 to 18.3 token/s at 06:00, with 91.8% coefficient of variation; nemotron-3-ultra is similarly volatile, ranging from 0.51 to 91.48 token/s.
  • nemotron-3-ultra is missing one of 48 expected samples (47 collected, 97.9% coverage), leaving 383 valid points overall, so its low average may be understated.
4-hour window · 383 points · 99.7% coverage
  • glm-5.3 is the strongest model with an average throughput of 125.68 token/s (p95 171.31 token/s), while nemotron-3-ultra is the weakest at 17.48 token/s average, with a minimum of 0.51 token/s.
  • The most operationally significant movement is glm-5.2's decline: after peaking at 177.85 token/s at 02:40, it fell to sustained 10-30 token/s readings from 03:00 onward, a -66.5% trend with 86.8% coefficient of variation. Conversely, nemotron-3-ultra recovered from near 1 token/s lows around 02:00-03:00 to a 91.48 token/s maximum at 04:55.
  • nemotron-3-ultra is missing one observation (47 of 48 samples, 97.9% coverage, no 03:05 point), leaving 383 valid points overall (99.7% coverage), so its recovery timing is partially unverified.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model at 135.88 token/s average throughput (p95 183.58 token/s), while nemotron-3-ultra is the weakest at 9.68 token/s average, never exceeding 82.5 token/s across the window.
  • The most operationally significant volatility is nemotron-3-ultra's collapse: after peaking at 82.5 token/s at 23:10 UTC, it stayed below roughly 1.5 token/s for most of 00:00–03:00, with a 156.3% coefficient of variation and a -71.0% trend; glm-5.2 also swung between 7.78 and 183.05 token/s.
  • No missing-data limitation applies: all eight models report 48 of 48 samples at 100.0% coverage (384 valid points), so the only constraint is the four-hour observation window itself.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model by average throughput at 141.89 token/s (p95 183.09 token/s); nemotron-3-ultra is the weakest at 15.86 token/s average, with a maximum of only 82.5 token/s.
  • nemotron-3-ultra shows the most operationally significant volatility: after peaking at 82.5 token/s at 23:10, it fell below 1 token/s from 00:10 through 00:50 (minimum 0.47 token/s), a -84.1% trend with 133.5% coefficient of variation; glm-5.2 also swung between 7.78 and 183.05 token/s.
  • No missing-data limitation applies: all eight models report 48 of 48 expected samples, 384 valid points overall, and 100.0% coverage, so the four-hour window is fully observed.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model at 153.44 token/s average throughput, peaking at 190.9 token/s; nemotron-3-ultra is the weakest at 14.58 token/s average, with a low of 0.47 token/s.
  • nemotron-3-ultra shows the most operationally significant volatility, with a coefficient of variation of 148.8% and swings between 0.47 and 82.5 token/s, including near-zero output from roughly 00:10 to 00:50 UTC; gemma4:31b also declined 22.9% across the window, ending at 38.97 token/s.
  • No missing-data limitation applies: all eight models report 48 of 48 expected samples, 384 valid points, and 100.0% coverage, though the dataset spans only this four-hour period.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model by average throughput at 144.94 token/s (p95 184.93 token/s, maximum 190.9 token/s), while nemotron-3-ultra is the weakest at 16.47 token/s average, never exceeding 82.5 token/s.
  • The most operationally significant volatility comes from nemotron-3-ultra: coefficient of variation 129.4%, swinging between 1.2 and 82.5 token/s, with roughly 1.2-4.8 token/s readings dominating before 22:20. deepseek-v4-flash shows the clearest upward trend at +16.5%, peaking at 145.51 token/s.
  • No missing-data limitation applies: all eight models report 48 of 48 samples at 100.0% coverage, totaling 384 valid points, though the dataset spans only this four-hour window.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model at 146.70 token/s average throughput (48 samples, range 32.61 to 186.40 token/s), while nemotron-3-ultra is the weakest at 11.56 token/s average, never exceeding 80.63 token/s and dipping to 1.20 token/s.
  • The most operationally significant pattern is glm-5.2's decline: a -29.0% trend, with throughput falling from roughly 148-160 token/s early in the window to a 44.47 token/s latest reading; nemotron-3-ultra is the most volatile, with a 153.0% coefficient of variation and readings spanning 1.20 to 80.63 token/s.
  • No missing-data limitation applies: all eight models report 48 of 48 expected samples, 100.0% coverage, and 384 valid points across the four-hour window.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model at 144.35 token/s average throughput, ahead of gemma4:31b at 126.02 token/s; nemotron-3-ultra is the weakest at 6.75 token/s average, with a maximum of only 31.72 token/s.
  • The most operationally significant volatility is nemotron-3-ultra: after spiking to 31.72 token/s at 20:35, it fell back below 2.2 token/s for the rest of the window, showing a 95.8% coefficient of variation and a -29.8% trend; deepseek-v4-flash posted the strongest upward trend at +18.7%.
  • No missing-data limitation applies: all eight models report 48 of 48 samples, 384 valid points overall, and 100% coverage, so the four-hour window is fully represented.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model at 140.87 token/s average throughput (p95 177.21 token/s), while nemotron-3-ultra is the weakest at 7.99 token/s average, never exceeding 31.72 token/s.
  • The most operationally significant pattern is nemotron-3-ultra's extreme volatility, with a coefficient of variation of 84.5% and swings from 1.04 to 31.72 token/s; glm-5.2 shows the steepest trend, rising 34.8% to a 213.21 token/s peak.
  • No missing-data limitation applies: all eight models delivered 48 of 48 expected samples, 384 valid points, and 100.0% coverage across the four-hour window.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model at 143.59 token/s average throughput (48 samples, peak 183.55 token/s), while nemotron-3-ultra is the weakest at 4.15 token/s average (48 samples, peak 27.04 token/s).
  • The most operationally significant volatility is nemotron-3-ultra: coefficient of variation 123.0%, trend of 448.9%, a low of 0.4 token/s at 17:00, and a spike to 27.04 token/s at 17:15; glm-5.2 also swung from 167.12 token/s at 17:20 down to 36.14 token/s at 17:35.
  • No missing-data limitation exists: all eight models have 48 of 48 expected samples with 100.0% coverage, totaling 384 valid points across the four-hour window.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model at 143.14 token/s average throughput (p95 172.52 token/s), while nemotron-3-ultra is the weakest at 2.74 token/s, far below the next-lowest model, glm-5.3-flash at 82.17 token/s.
  • The most operationally significant volatility is nemotron-3-ultra: it stayed near 1 token/s for most of the window (minimum 0.4 token/s at 17:00), spiked to 27.04 token/s at 17:15, then fell back to 1.2 token/s by 18:00, with a 151.4% coefficient of variation; minimax-m3 also swung between 31.63 and 144.38 token/s.
  • No missing-data limitation applies: all eight models report 48 of 48 expected samples, 384 valid points, and 100.0% coverage across the four-hour window.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model at 144.73 token/s average throughput (peak 186.92 token/s, p95 176.49 token/s), while nemotron-3-ultra is weakest at 1.57 token/s average, never exceeding 2.75 token/s.
  • minimax-m3 is the most volatile (cv 40.9%), ranging 31.63–133.39 token/s with a long stretch near 40–50 token/s; glm-5.2 shows the steepest trend at +13.7%, rising from a 27.55 token/s low toward a 176.38 token/s peak.
  • No missing-data limitation exists: all eight models report 48 of 48 samples with 100.0% coverage, and the 384 valid points match the expected total.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model at 144.15 token/s average throughput (peak 186.92 token/s), while nemotron-3-ultra is the weakest at 2.01 token/s average, peaking at only 4.07 token/s.
  • The most significant volatility is minimax-m3, with a 43.3% coefficient of variation and swings between 31.63 and 133.39 token/s; deepseek-v4-pro shows the steepest decline among mainstream models at -13.7% trend, and nemotron-3-ultra fell 37.1% to a 0.92 token/s low near 15:30.
  • No missing-data limitation exists: all eight models report 48 of 48 expected samples, 100.0% coverage, and 384 valid points across the four-hour window.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model at 151.83 token/s average throughput (p95 184.87 token/s), while nemotron-3-ultra is the weakest at 4.56 token/s average, far below the next-lowest model minimax-m3 at 63.60 token/s.
  • The most operationally significant change is nemotron-3-ultra's collapse: it peaked at 34.17 token/s at 11:10 UTC, then fell to roughly 1.3–2.4 token/s for the remainder of the window, a 74.2% decline, with a 135.6% coefficient of variation.
  • No missing-data limitation applies: all eight models report 48 of 48 expected samples, 384 valid points, and 100.0% coverage across the four-hour window, so the averages rest on complete five-minute observations.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model at 153.34 token/s average throughput (p95 184.87 token/s, maximum 187.43 token/s), while nemotron-3-ultra is the weakest at 9.45 token/s average, with a maximum of only 41.0 token/s.
  • The most operationally significant movement is nemotron-3-ultra's sustained collapse: after running roughly 15-35 token/s early in the window, it fell below 3 token/s from around 12:45 UTC onward, ending at 1.49 token/s, a -84.9% trend; glm-5.2 also swung widely between 22.39 and 165.25 token/s (cv 35.7%).
  • No missing-data limitation applies: every model reports 48 of 48 expected samples at 100.0% coverage, and the 384 valid points match the full expected count.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model by average throughput at 156.44 token/s (p95 of 184.4 token/s), while nemotron-3-ultra is the weakest at 9.91 token/s average, with a maximum of only 41.0 token/s.
  • The most operationally significant volatility is nemotron-3-ultra: after swinging between 1.66 and 41.0 token/s, it settled near 2 to 5 token/s from 11:25 onward, showing a 101.6% coefficient of variation and a -42.5% trend, far below all other models.
  • No missing-data limitation applies: all eight models report 48 of 48 expected samples, 384 valid points total, and 100.0% coverage, so the four-hour window is fully populated.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model by average throughput at 156.83 token/s (peak 187.43 token/s, p95 184.4 token/s), while nemotron-3-ultra is the weakest at 10.98 token/s average, never exceeding 41.0 token/s.
  • The most operationally significant volatility comes from nemotron-3-ultra, with a 98.2% coefficient of variation and swings between 1.5 and 41.0 token/s; glm-5.2 also shows abrupt single-interval drops, such as 175.4 token/s at 08:20 falling to 37.07 token/s at 08:30.
  • No missing-data limitation applies: all eight models report 48 of 48 expected samples, 384 valid points total, and 100.0% coverage, though the dataset spans only this four-hour window.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model with average throughput of 156.98 token/s (maximum 187.71 token/s), while nemotron-3-ultra is the weakest at 13.82 token/s average, far below every other model.
  • The most operationally significant volatility is nemotron-3-ultra (cv 108.8%): after peaking at 80.28 token/s at 07:25, it fell to a 1.49 token/s minimum and stayed below roughly 5 token/s for most of the 07:50–09:55 window, with a -16.4% trend.
  • No missing-data limitation applies: all eight models report 48 of 48 expected samples at 100.0% coverage, totaling 384 valid points across the four-hour window.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model with an average throughput of 155.76 token/s (p95 182.96 token/s), while nemotron-3-ultra is the weakest at 11.94 token/s average, never exceeding 80.28 token/s in any sample.
  • The most operationally significant volatility is nemotron-3-ultra: coefficient of variation 137.8% and a -69.9% trend, with sustained stretches near 1.5 to 3 token/s (minimum 1.49 token/s at 08:00) after early spikes of 57.2 and 80.28 token/s.
  • No missing-data limitation applies: all eight models report 48 of 48 expected samples, 384 valid points, and 100.0% coverage across the four-hour window.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model at 159.28 token/s average throughput (p95 184.30 token/s), while nemotron-3-ultra is the weakest at 15.95 token/s average, peaking at only 80.28 token/s.
  • nemotron-3-ultra shows the most operationally significant volatility, with a 118.2% coefficient of variation and sustained stretches near 2-3 token/s, including 1.49 token/s at 08:00; deepseek-v4-pro shows the largest upward trend at +13.8%.
  • No missing-data limitation exists: every model reports 48 of 48 expected samples with 100.0% coverage (384 valid points), though the window covers only four hours.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model at 164.47 token/s average throughput, with a 187.71 token/s maximum and the lowest relative volatility (16.4% CV); nemotron-3-ultra is the weakest at 24.54 token/s average, peaking at only 101.62 token/s and ending at 1.49 token/s.
  • The most operationally significant volatility is nemotron-3-ultra, which swings between 1.49 and 101.62 token/s with a 105.8% CV and a -40.2% trend, including sustained stretches near 2-4 token/s from roughly 05:35 to 06:55; glm-5.2 also shows a sharp single-interval drop to 14.74 token/s at 05:35.
  • No missing-data limitation applies: all eight models report 48 of 48 expected samples, 384 valid points, and 100.0% coverage for the 04:03-08:03 UTC window.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model at 165.69 token/s average throughput, while nemotron-3-ultra is the weakest at 26.04 token/s average, roughly a sixfold gap; glm-5.3 also shows the tightest spread at 13.5% CV.
  • nemotron-3-ultra shows the most operationally significant volatility: 107.0% CV, a -52.2% trend, and repeated collapses to 1.63 token/s against a 101.62 token/s peak; glm-5.2 is next most erratic with 44.85 stddev and dips to 14.74 token/s.
  • No missing-data limitation applies: all eight models report 48 of 48 expected samples, 384 valid points, and 100.0% coverage across the four-hour window.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model by average throughput at 171.05 token/s (p95 184.61 token/s, minimum 146.47 token/s), while nemotron-3-ultra is the weakest at 34.42 token/s average, peaking at only 105.72 token/s.
  • The most operationally significant volatility is nemotron-3-ultra, with a coefficient of variation of 85.3% and repeated collapses to near-zero output, including 1.65 token/s at 06:00 and readings between roughly 2 and 4 token/s from 03:50 to 04:15 and again from 05:35 onward. glm-5.2 also swings widely, from 14.74 to 205.32 token/s.
  • No missing-data limitation exists in this window: all eight models report 48 of 48 expected samples, with 384 valid points and 100.0% coverage, so the averages fully represent the four-hour period.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model at 167.36 token/s average throughput, holding a tight 116.24–187.06 token/s range; nemotron-3-ultra is weakest at 38.02 token/s average, never exceeding 105.72 token/s.
  • The most operationally significant volatility is nemotron-3-ultra's 77.9% coefficient of variation, swinging from a 2.12 token/s low at 04:05 to 101.62 token/s at 04:20; deepseek-v4-flash also dropped to 10.18 token/s at 04:50.
  • No missing-data limitation applies: all eight models recorded 48 of 48 samples with 100.0% coverage, totaling 384 valid points across the four-hour window.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model at 162.32 token/s average throughput; nemotron-3-ultra is the weakest at 31.82 token/s average, with a 1.74–105.72 token/s range.
  • nemotron-3-ultra shows the most operationally significant volatility: an 88.7% coefficient of variation and six consecutive samples below 4 token/s from 00:40 to 01:05, plus a latest reading of 2.73 token/s, versus isolated dips elsewhere such as glm-5.2's 25.68 token/s at 02:25 and deepseek-v4-pro's 39.5 token/s at 03:00.
  • No missing-data limitation applies: all eight models report 48 of 48 expected samples, total valid points are 384, and coverage is 100.0%, so the only constraint is the four-hour observation window itself.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 posted the highest average throughput at 161.74 token/s (p95 184.12 token/s), while nemotron-3-ultra was weakest at 37.40 token/s average, well below the next-lowest model minimax-m3 at 60.49 token/s.
  • The most operationally significant volatility is nemotron-3-ultra's coefficient of variation of 74.3%: after running near 78 token/s around 00:00, it collapsed to a sustained trough of 1.74–3.47 token/s from 00:40 through 01:05, an outage-like degradation. glm-5.2 also showed a -17.1% trend with a dip to 25.68 token/s at 02:25.
  • No missing-data limitation applies: all eight models recorded 48 of 48 expected samples, 384 valid points total, and 100.0% coverage, so the four-hour window is complete.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model at 161.40 token/s average throughput (p95 184.12 token/s); nemotron-3-ultra is the weakest at 36.79 token/s average, never exceeding 101.92 token/s.
  • The most operationally significant volatility is nemotron-3-ultra: 83.3% coefficient of variation and a -46.9% trend, with throughput collapsing from a 101.92 token/s peak to below 4 token/s for most of 00:30–01:15 UTC; gemma4:31b also swings widely (49.53–176.08 token/s, cv 27.3%).
  • No missing-data limitation exists: all eight models report 48 of 48 samples at 100.0% coverage, totaling 384 valid points, so results are constrained only by the four-hour observation window.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model at 158.43 token/s average throughput, edging glm-5.2 at 152.55 token/s, while nemotron-3-ultra is weakest at 29.34 token/s average, roughly a fifth of the leader despite peaking at 101.92 token/s.
  • nemotron-3-ultra shows the most operationally significant volatility: a 105.6% coefficient of variation, swings between 1.74 and 101.92 token/s, and a late-window collapse from about 78 token/s at 00:00 to 1.74 token/s at 00:40; glm-5.2 also dipped sharply to 21.12 token/s at 21:55.
  • No missing-data limitation applies: all eight models report 48 of 48 expected samples, 384 valid points, and 100.0% coverage, though the window spans only four hours.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model by average throughput at 159.81 token/s (p95 of 182.91 token/s), while nemotron-3-ultra is the weakest at 26.08 token/s average, far below the next-lowest model, minimax-m3, at 68.89 token/s.
  • The most operationally significant volatility is nemotron-3-ultra's step change around 22:25 UTC: it ran between 2.07 and 6.78 token/s for roughly the first two and a half hours, then jumped to peaks of 101.92 token/s, producing a 118.3% coefficient of variation.
  • No missing-data limitation applies: all eight models recorded 48 of 48 expected samples, totaling 384 valid points at 100.0% coverage across the four-hour window.