← Performance dashboard

Hourly performance insights

Summaries of rolling four-hour performance data

1201 retained summaries
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model at 152.93 token/s average throughput (p95 180.97 token/s), while nemotron-3-ultra is the weakest at 5.45 token/s average, peaking at only 9.57 token/s.
  • glm-5.2 shows the highest volatility (CV 34.8%), dropping to 7.02 token/s at 03:35 and 10.8 token/s at 04:20; glm-5.3 also dipped sharply to 36.92 token/s at 04:20. deepseek-v4-pro was the steadiest performer (CV 16.1%, stddev 16.85 token/s).
  • No missing-data limitation applies: all eight models recorded 48 of 48 expected samples, 100.0% coverage, and 384 valid points across the four-hour window.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 delivered the highest average throughput at 152.71 token/s (peak 185.15 token/s), while nemotron-3-ultra was weakest at 5.4 token/s average, never exceeding 9.57 token/s.
  • glm-5.2 showed the sharpest volatility, with a 40.0% coefficient of variation and swings from 7.02 to 83.45 token/s; glm-5.3 and glm-5.3-flash both trended down about 11% (trend -11.3% and -11.8%), while deepseek-v4-flash rose 12.5%.
  • No missing-data limitation applies: all eight models recorded 48 of 48 expected samples, 384 valid points total, at 100.0% coverage across the four-hour window.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model, averaging 158.15 token/s (p95 183.67 token/s), while nemotron-3-ultra is the weakest at 5.97 token/s average, roughly 26 times slower.
  • glm-5.3 is also the steadiest performer (CV 16.7%), while nemotron-3-ultra shows the sharpest decline at -25.3% trend, dropping from a 15.64 token/s peak to a 2.0 token/s low; glm-5.2 is the most volatile at 41.8% CV with dips to 7.02 token/s.
  • No missing-data limitation exists: all eight models report 48 of 48 expected samples at 100.0% coverage, so the four-hour window is fully represented with no gaps to caveat these figures.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model at 161.83 token/s average throughput (p95 183.91 token/s), while nemotron-3-ultra is the weakest at 7.8 token/s average, with a maximum of only 62.89 token/s.
  • The most operationally significant volatility comes from nemotron-3-ultra, which shows a 140.5% coefficient of variation, a 0.86 token/s minimum, and a -54.8% trend; glm-5.3-flash also declined 23.7%, dropping from early readings near 120 token/s to a 45.23 token/s low.
  • No missing-data limitation applies: all eight models report 48 of 48 expected samples with 100.0% coverage, and the 384 valid points fully cover the four-hour window.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model by average throughput at 158.41 token/s (p95 183.91 token/s), while nemotron-3-ultra is the weakest at 17.93 token/s, roughly a ninth of the leader's average.
  • The most operationally significant change is nemotron-3-ultra's collapse after 23:20 UTC, falling from roughly 48 token/s to readings below 2 token/s and a floor of 0.86 token/s, with a -76.4% trend and 120.9% coefficient of variation.
  • No missing-data limitation exists in this window: all eight models have 48 of 48 expected samples, 384 valid points total, and 100.0% coverage, though the dataset spans only four hours.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model at 154.96 token/s average throughput, ahead of gemma4:31b at 117.81 token/s; nemotron-3-ultra is the weakest at 26.24 token/s average, ending the window at 4.35 token/s.
  • The most operationally significant volatility is nemotron-3-ultra's collapse after 23:20 UTC, falling from roughly 48 token/s to sustained sub-2 token/s readings (minimum 0.86 token/s), producing a 96.6% coefficient of variation and a -74.3% trend; glm-5.3-flash trended up 17.0% to a 100.01 token/s average.
  • No missing-data limitation applies: all eight models logged 48 of 48 expected samples, 384 valid points total, and 100.0% coverage, though the dataset covers only this four-hour window.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model with an average throughput of 145.72 token/s (p95 180.61 token/s); nemotron-3-ultra is the weakest at 27.13 token/s average, ending the window at 1.56 token/s.
  • The most operationally significant volatility is nemotron-3-ultra's late-window collapse: from 30.88 token/s at 23:20 UTC to 0.86 token/s at 23:40, then holding near 1 token/s through 00:00, driving its 91.2% coefficient of variation.
  • No missing-data limitation applies: every model reports 48 of 48 expected samples, 384 valid points overall, and 100.0% coverage across the four-hour window.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model at 138.9 token/s average throughput, peaking at 185.93 token/s; nemotron-3-ultra is the weakest at 26.0 token/s average, with a low of 4.05 token/s.
  • The most operationally significant volatility is nemotron-3-ultra, whose throughput swings between 4.05 and 96.85 token/s (cv 89.9%), including a spike to 96.85 token/s at 21:25 UTC; glm-5.3 also shows a sharp single-interval dip to 30.03 token/s at 21:50 UTC.
  • No missing-data limitation applies: all eight models report 48 of 48 samples, 384 valid points, and 100.0% coverage across the four-hour window.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model with an average throughput of 126.15 token/s (p95 171.55 token/s), while nemotron-3-ultra is the weakest at 16.15 token/s, roughly one-eighth of glm-5.3's rate.
  • nemotron-3-ultra shows the sharpest volatility: after mostly 3-13 token/s, it spiked to 96.85 token/s at 21:25 and 82.53 token/s at 21:55, producing a 116.5% coefficient of variation; glm-5.2 also swung between 2.1 and 81.74 token/s (cv 64.9%).
  • No samples are missing: all eight models report 48 of 48 expected samples, 384 valid points total, and 100.0% coverage; the four-hour window is the only scope limitation.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model at 118.82 token/s average throughput (peak 173.74 token/s), while nemotron-3-ultra is the weakest at 8.03 token/s average (minimum 0.48 token/s).
  • glm-5.2 shows the most operationally significant volatility, with a coefficient of variation of 68.0% and a sustained trough from 19:20 to 19:50 where throughput fell to between 2.1 and 10.43 token/s before recovering to 81.74 token/s at 20:25; glm-5.3 also trended upward 20.1% to 159.95 token/s at 21:00.
  • No missing-data limitation applies: all eight models recorded 48 of 48 expected samples, valid_point_count is 384, and coverage is 100.0%, so the four-hour window is complete.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model, averaging 121.36 token/s (p95 155.07 token/s), while nemotron-3-ultra is the weakest, averaging 5.64 token/s with a maximum of just 12.98 token/s.
  • The most operationally significant change is glm-5.2's 49.3 percent decline, with throughput falling to between 2.1 and 10.43 token/s from 19:20 to 19:50 UTC; nemotron-3-ultra trended up 67.2 percent and glm-5.3-flash up 24.1 percent over the window.
  • No missing-data limitation exists: every model reports 48 of 48 expected samples at 100.0 percent coverage, and the 384 valid points match the expected total for the four-hour period.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model with an average throughput of 117.75 token/s (p95 157.74 token/s, maximum 167.56 token/s), while nemotron-3-ultra is the weakest at 3.94 token/s average, never exceeding 9.26 token/s.
  • The most operationally significant movement is glm-5.2's decline of 28.5% to a 42.42 token/s average, with high relative volatility (cv 48.4%) and repeated drops near 9-10 token/s; nemotron-3-ultra rose 176.9% but stays far below every other model.
  • There is no missing-data limitation: all eight models recorded 48 of 48 expected samples, 384 valid points total, with 100.0% coverage across the four-hour window.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model at 114.56 token/s average throughput, peaking at 167.56 token/s, while nemotron-3-ultra is the weakest at 3.01 token/s average, never exceeding 9.26 token/s.
  • The most operationally significant volatility is nemotron-3-ultra: throughput fell to 0.48 token/s at 17:05 UTC, then climbed to 9.26 token/s by 17:25, with a 79.4% coefficient of variation and a 133.4% trend across the window.
  • No missing-data limitation exists: all eight models report 48 of 48 expected samples, 384 valid points, and 100.0% coverage, so averages rest on complete five-minute observations across the four-hour period.
4-hour window · 376 points · 97.9% coverage
  • glm-5.3 is the strongest model with an average throughput of 114.03 token/s (p95 of 153.59 token/s), while nemotron-3-ultra is the weakest at an average of 1.82 token/s, never exceeding 5.13 token/s.
  • The most operationally significant volatility is glm-5.3's drop to 11.94 token/s at 14:25 followed by recovery to 130.48 token/s at 14:35; glm-5.3-flash also shows the steepest sustained decline, trending -28.2% from 121.94 token/s at 13:25 to a 39.05 token/s low at 16:20.
  • Every model records 47 of 48 expected samples (97.9% coverage), so one five-minute interval is missing per model, leaving a small gap in the four-hour window.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model by average throughput at 112.26 token/s (p95 148.23 token/s), while nemotron-3-ultra is the weakest at 1.85 token/s average, never exceeding 3.46 token/s.
  • glm-5.3 shows the most operationally significant volatility: coefficient of variation 32.3%, with throughput collapsing from 45.92 token/s at 14:15 to 11.94 token/s at 14:25, then recovering to 130.48 token/s at 14:30; its four-hour trend is -10.1%.
  • No missing-data limitation applies: all eight models report 48 of 48 expected samples, 384 valid points total, and 100.0% coverage across the 12:03 to 16:03 UTC window.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model by average throughput at 114.21 token/s (peaking at 171.47 token/s), while nemotron-3-ultra is the weakest at 2.08 token/s average, roughly 55 times slower than the leader.
  • glm-5.3 also shows the most operationally significant volatility: coefficient of variation 33.0%, with a plunge from 124.47 token/s at 14:10 to 11.94 token/s at 14:25 before recovering to 130.48 token/s at 14:35; nemotron-3-ultra trended down 43.3% overall, ending at 1.13 token/s.
  • No missing-data limitation applies: all eight models report 48 of 48 expected samples, 384 valid points, and 100.0% coverage across the four-hour window.
4-hour window · 376 points · 97.9% coverage
  • glm-5.3 is the strongest model with an average throughput of 125.29 token/s (p95 164.82 token/s), while nemotron-3-ultra is the weakest at 2.53 token/s average, roughly fifty times slower.
  • The most operationally significant movement is nemotron-3-ultra's steady decline, trending -36.2% from 3.12 token/s at 10:05 to 1.14 token/s at 13:55, with a low of 0.79 token/s at 13:50; glm-5.3 also shows sharp dips, falling to 37.83 token/s at 13:40.
  • Each model recorded 47 of 48 expected samples (97.9% coverage), giving 376 valid points out of 384, so one five-minute interval per model is missing and short-window figures may be slightly affected.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model at 122.8 token/s average throughput, ahead of gemma4:31b at 113.35 token/s; nemotron-3-ultra is the weakest at 3.0 token/s average, roughly 40x slower than the leader.
  • glm-5.2 shows the largest trend, up 41.6% over the window, climbing from 20.06 token/s at 09:05 to a peak of 81.98 token/s at 11:40; glm-5.3 is the most volatile high-throughput model, ranging 18.59 to 171.47 token/s (CV 27.9%).
  • No missing-data limitation: all eight models have 48 of 48 samples, 384 valid points, and 100.0% coverage for the period.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model at 122.47 token/s average throughput, ahead of gemma4:31b at 114.0 token/s; nemotron-3-ultra is weakest at 3.2 token/s average, roughly 38 times slower than glm-5.3.
  • glm-5.3 showed the sharpest volatility, dropping from 132.8 token/s at 08:50 to 35.53 token/s at 08:55 and to 18.59 token/s at 09:20, while minimax-m3 spiked to 104.62 token/s at 10:30 against a 41.34 token/s average; glm-5.2 trended upward 30.6%.
  • No missing-data limitation applies: all eight models report 48 of 48 samples with 100.0% coverage, and the dataset's 384 valid points match the expected count, so conclusions are not constrained by gaps.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model by average throughput at 126.84 token/s (peak 176.60 token/s), while nemotron-3-ultra is the weakest at 3.37 token/s average (peak 5.28 token/s).
  • The most operationally significant volatility is glm-5.3's collapse from 132.80 token/s at 08:50 to 35.53 token/s at 08:55 and a low of 18.59 token/s at 09:20, far below its 126.84 token/s average; minimax-m3 also spiked to 104.62 token/s at 10:30 against a 39.84 token/s average.
  • No missing-data limitation applies: all 384 expected observations are present, 48 per model at 100.0% coverage, so the only constraint is that results reflect this single four-hour window.
4-hour window · 384 points · 100.0% coverage
  • Strongest average throughput was glm-5.3 at 123.03 token/s; weakest was nemotron-3-ultra at 4.21 token/s, with minimax-m3 also low at 34.97 token/s and gemma4:31b second at 110.76 token/s.
  • glm-5.3 showed the sharpest operational swings, peaking at 176.6 token/s at 07:40 before dropping to 18.59 token/s at 09:20, while nemotron-3-ultra had the highest relative volatility (CV 74.9%, range 1.8 to 22.94 token/s) and deepseek-v4-pro was steadiest (CV 17.2%).
  • No missing-data limitation applies: all eight models report 48 of 48 expected samples, 384 valid points overall, and 100.0% coverage, so observed dips reflect measured values rather than collection gaps.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model with an average throughput of 129.90 token/s (p95 171.97 token/s, maximum 176.60 token/s), while nemotron-3-ultra is the weakest at 4.69 token/s average, roughly 28 times slower.
  • nemotron-3-ultra shows the most operationally significant volatility: a coefficient of variation of 69.5%, a -43.1% trend, and throughput ranging from 1.80 to 22.94 token/s, with most readings between 2 and 6 token/s; glm-5.3-flash trended up 15.3% to an 83.72 token/s average.
  • No missing-data limitation applies: all eight models report 48 of 48 samples, 384 valid points total, and 100.0% coverage across the four-hour window.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model at 137.31 token/s average throughput (peak 176.6 token/s), while nemotron-3-ultra is the weakest at 5.41 token/s average (minimum 2.0 token/s), roughly 25 times slower.
  • minimax-m3 shows the steepest decline, trending down 36.3% to a 43.48 token/s average, and nemotron-3-ultra is the most volatile with a 60.7% coefficient of variation; glm-5.3-flash is the only clear gainer, up 13.2%.
  • No missing-data limitation exists: all eight models report 48 of 48 expected samples, 100.0% coverage, and 384 valid points across the four-hour window.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model with an average throughput of 134.46 token/s (peak 176.53 token/s), while nemotron-3-ultra is the weakest at 5.52 token/s average, roughly 24 times lower than the leader.
  • The most significant volatility comes from nemotron-3-ultra (CV 58.7%) and minimax-m3 (CV 43.8%), with minimax-m3 also showing the steepest decline at -18.9% trend, falling from a 107.49 token/s peak to a 35.8 token/s latest reading; glm-5.3 dropped sharply from 176.53 to 39.72 token/s at 07:00, and deepseek-v4-pro was steadiest (CV 12.4%).
  • No missing-data limitation: all eight models have 48 of 48 expected samples with 100.0% coverage, totaling 384 valid points across the four-hour window.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model with an average throughput of 140.4 token/s (p95 176.31 token/s), while nemotron-3-ultra is the weakest at an average of 5.15 token/s, never exceeding 8.8 token/s across the window.
  • The most operationally significant volatility appears in glm-5.3, which repeatedly dropped from sustained 150-185 token/s levels to brief lows of 38.65, 61.08, 45.59, and 53.01 token/s; minimax-m3 also swung between 16.05 and 107.49 token/s with 40.6% coefficient of variation.
  • No missing-data limitation applies: all eight models report 48 of 48 samples with 100.0% coverage, and the full dataset contains 384 valid points, so conclusions rest on complete five-minute observations.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model with an average throughput of 139.96 token/s (peaking at 185.94 token/s), while nemotron-3-ultra is the weakest at 5.05 token/s average, roughly 28 times slower than the leader.
  • glm-5.2 shows the most operationally significant volatility, with a 38.0% coefficient of variation and swings between 86.67 token/s and 10.18 token/s; minimax-m3 is similarly erratic at 41.7% CV, whereas deepseek-v4-pro runs steadiest at 13.1% CV around its 103.53 token/s average.
  • No missing-data limitation applies: all eight models report 48 of 48 expected samples, 384 valid points total, and 100.0% coverage, though the dataset spans only this four-hour window.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model at 136.38 token/s average throughput, ahead of gemma4:31b at 114.85 token/s and deepseek-v4-pro at 101.58 token/s; nemotron-3-ultra is weakest at 4.41 token/s average, with a maximum of only 8.06 token/s.
  • glm-5.3-flash shows the steepest trend, rising 28.7% to average 76.55 token/s, while glm-5.2 declined 17.2%; nemotron-3-ultra is the most volatile at 41.0% coefficient of variation, and glm-5.3 dipped to 36.03 token/s at 01:45 UTC.
  • No missing-data limitation applies: every model reports 48 of 48 expected samples at 100.0% coverage, totaling 384 valid points, so the four-hour window is fully covered.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model, averaging 138.18 token/s (peaking at 185.94 token/s), while nemotron-3-ultra is the weakest, averaging 5.22 token/s with a maximum of only 45.93 token/s.
  • The most operationally significant volatility is glm-5.2's collapse to 4.01 token/s at 00:00 before recovering to 86.12 token/s at 00:20; nemotron-3-ultra also shows extreme instability with a coefficient of variation of 120.1%.
  • No missing-data limitation exists: all eight models report 48 of 48 expected samples, with 384 valid points and 100.0% coverage across the four-hour window.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model with an average throughput of 138.78 token/s, peaking at 185.94 token/s; nemotron-3-ultra is the weakest at 7.49 token/s average, never exceeding 45.93 token/s.
  • The most operationally significant volatility is nemotron-3-ultra's coefficient of variation of 108.8% alongside a -61.0% trend, while glm-5.2 collapsed from 34.59 token/s at 23:40 to 4.01 token/s at 00:00 before recovering to 81.18 token/s at 00:15.
  • No missing-data limitation applies: all eight models report 48 of 48 expected samples, 384 valid points, and 100.0% coverage across the four-hour window.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model at 137.38 token/s average throughput (p95 178.65 token/s), while nemotron-3-ultra is the weakest at 7.9 token/s average, with a maximum of only 45.93 token/s.
  • The most operationally significant volatility is glm-5.2's collapse between 23:45 and 00:10 UTC, dropping from a 46.37 token/s average to a low of 4.01 token/s before recovering to 86.12 token/s; nemotron-3-ultra also shows high instability with a 102.4% coefficient of variation and a -47.9% trend.
  • No missing-data limitation exists: all eight models report 48 of 48 expected samples, 384 valid points total, and 100.0% coverage across the four-hour window.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model at 133.5 token/s average throughput, ahead of deepseek-v4-pro at 106.82 token/s; nemotron-3-ultra is the weakest at 8.06 token/s average, with a p95 of only 22.43 token/s.
  • The most significant volatility is nemotron-3-ultra's 98.9% coefficient of variation, ranging 1.58 to 45.93 token/s, while glm-5.2 shows a late-window collapse from 34.59 token/s at 23:40 to 4.01 token/s at 00:00; glm-5.3 trended up 18.0%.
  • No missing-data limitation exists: all eight models report 48 of 48 samples with 100.0% coverage, totaling 384 valid points, though the window covers only four hours.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model at 127.17 token/s average throughput, with a p95 of 178.65 token/s and a peak of 180.73 token/s; nemotron-3-ultra is the weakest at 7.13 token/s average, peaking at only 22.85 token/s.
  • The most operationally significant movement is nemotron-3-ultra's late surge: after holding mostly between 2 and 9 token/s, it jumped to 22.85 token/s at 22:35 and stayed near 19-22 token/s through 23:00, driving its 80.0% coefficient of variation; glm-5.3-flash declined 12.3% overall.
  • No missing-data limitation applies: all eight models report 48 of 48 samples at 100.0% coverage, and the 384 valid points match the expected total.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model at 121.74 token/s average throughput (48 of 48 samples, peak 180.73 token/s), while nemotron-3-ultra is the weakest at 4.13 token/s average, roughly 30x slower.
  • glm-5.3 shows the most operationally significant volatility: repeated collapses from ~140-160 token/s to 33.17-47.57 token/s, with stddev 38.25 token/s and cv 31.4%; deepseek-v4-flash also declined 22.0% to a 40.88 token/s low near 21:40.
  • No missing-data limitation exists: all eight models report 48 of 48 expected samples with 100.0% coverage (384 valid points), so the only constraint is the four-hour window itself.
4-hour window · 384 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model by average throughput at 122.99 token/s, peaking at 174.04 token/s, while nemotron-3-ultra is the weakest at 3.05 token/s average and never exceeded 7.36 token/s.
  • glm-5.3 shows the most operationally significant volatility: otherwise sustained throughput near 130-160 token/s is interrupted by sudden drops to 6.74 token/s at 17:20, 36.67 token/s at 19:50, and roughly 33-34 token/s at 20:50-20:55, with a stddev of 42.04 token/s.
  • No missing-data limitation applies: all eight models delivered 48 of 48 expected samples at 100.0% coverage, totaling 384 valid points across the four-hour window.
4-hour window · 368 points · 95.8% coverage
  • Strongest average throughput was deepseek-v4-flash at 128.19 token/s (p95 156.19 token/s), while nemotron-3-ultra was weakest at 2.61 token/s, never exceeding 4.67 token/s across the window.
  • The most significant volatility came from glm-5.2, which spiked to 169.51 token/s at 18:25 after sitting near 40 token/s, with a 51.2% coefficient of variation; glm-5.3 also collapsed to 6.74 token/s at 17:20, and minimax-m3 trended down 30.6%, settling near 45-49 token/s after peaking at 127.84 token/s.
  • Each model recorded 46 of 48 expected samples (95.8% coverage), with observations missing around 16:50-16:55 UTC, so short-lived throughput dips or spikes in those gaps are not captured in the averages.
4-hour window · 368 points · 95.8% coverage
  • deepseek-v4-flash is the strongest model at 120.52 token/s average throughput, peaking at 174.04 token/s, while nemotron-3-ultra is the weakest at 2.11 token/s average, roughly 57 times slower.
  • The most operationally significant shift is gemma4:31b, which climbed from a 14.49 token/s low at 16:20 to a 140.99 token/s peak at 17:10, ending the window up 102.7%; glm-5.2 shows the highest volatility with a 63.7% coefficient of variation.
  • Each model recorded 46 of 48 expected samples (95.8% coverage), with the 16:50 and 16:55 UTC intervals absent across all eight models, so short-lived drops during that gap are unobserved.
4-hour window · 352 points · 91.7% coverage
  • Strongest average throughput was deepseek-v4-flash at 116.61 token/s (p95 150.9 token/s, max 174.04 token/s), while nemotron-3-ultra was weakest at 2.02 token/s average, never exceeding 3.06 token/s across the window.
  • The most operationally significant volatility was gemma4:31b, which swung from a 7.55 token/s minimum to a 140.99 token/s maximum with a coefficient of variation of 80.9% and a +291.6% trend, shifting from roughly 8-30 token/s before 16:30 to sustained 85-140 token/s afterward; glm-5.3 also dipped to 6.74 token/s at 17:20.
  • Each of the eight models recorded 44 of 48 expected samples (91.7% coverage), leaving 352 valid points of 384 expected, so four five-minute observations per model are missing and short-window averages may be skewed.
4-hour window · 352 points · 91.7% coverage
  • glm-5.3 is the strongest model by average throughput at 112.82 token/s (p95 147.26 token/s), while nemotron-3-ultra is the weakest at 1.97 token/s average, never exceeding 2.69 token/s across the window.
  • The most operationally significant movement is gemma4:31b's late surge: it ran near 20 token/s for most of the period, then jumped from 61.56 token/s at 16:30 to a sustained 109.81 token/s at 16:40–16:45, driving its 187.4% trend and 96.7% coefficient of variation. glm-5.2 also swung widely, from a 14.27 token/s low to a 184.81 token/s peak.
  • Every model is missing 4 of 48 expected samples (44 collected, 91.7% coverage; 352 valid points total), with gaps at 14:05, 14:10, 16:50, and 16:55 UTC, so short-lived throughput changes in those intervals cannot be assessed.
4-hour window · 344 points · 89.6% coverage
  • glm-5.3 delivered the highest average throughput at 109.76 token/s (peak 149.73 token/s), while nemotron-3-ultra was weakest at 1.91 token/s, never exceeding 2.69 token/s.
  • glm-5.2 showed the most operationally significant volatility, with a coefficient of variation of 85.2% and a range of 13.86 to 184.81 token/s, including a spike from 109.0 token/s at 15:40 to 184.81 token/s at 15:45 before dropping to 48.93 token/s at 16:00.
  • Every model recorded 43 of 48 expected samples (89.6% coverage), with the same five timestamps missing per model, leaving 344 valid points, so short-lived degradations during those gaps would go undetected.
4-hour window · 344 points · 89.6% coverage
  • glm-5.3 is the strongest model at 112.97 token/s average throughput (peak 171.99 token/s), while nemotron-3-ultra is the weakest at 2.22 token/s average, roughly fifty times lower.
  • gemma4:31b shows the sharpest decline, falling from an 81.66 token/s peak at 11:50 to 7.55 token/s at 14:50, a -65.9% trend with 69.7% coefficient of variation; glm-5.2 moved the opposite way, jumping from 14.27 token/s at 14:00 to 124.7 token/s at 14:50.
  • Each model has 43 of 48 expected samples (89.6% coverage), with shared gaps at 12:40-12:50 and 14:05-14:10 UTC, so averages and trends may not reflect those unobserved intervals.
4-hour window · 352 points · 91.7% coverage
  • glm-5.3 is the strongest model by average throughput at 119.51 token/s (peak 171.99 token/s), while nemotron-3-ultra is the weakest at 2.11 token/s average, roughly 57 times slower.
  • gemma4:31b shows the most operationally significant volatility: throughput trended down 53.3%, falling from an 81.66 token/s peak at 11:50 to an 8.22 token/s low at 13:10, with a 52.4% coefficient of variation; glm-5.3 also swung sharply, dipping to 20.28 token/s at 13:10.
  • Each model is missing 4 of 48 expected samples (44 collected, 91.7% coverage), including a shared gap between the 12:35 and 12:55 observations, so throughput during that interval is unobserved for all models.
4-hour window · 360 points · 93.8% coverage
  • glm-5.3 delivered the highest average throughput at 131.45 token/s (peaking at 171.99 token/s), while nemotron-3-ultra was weakest at 2.07 token/s average, never exceeding 3.9 token/s.
  • The most operationally significant volatility is gemma4:31b: coefficient of variation 45.5%, a swing from a 118.92 token/s peak to a 12.04 token/s latest reading, and a -27.7% trend, versus deepseek-v4-flash's steadier 103.43 token/s average.
  • Every model recorded 45 of 48 expected samples (93.8% coverage), with a shared gap between 12:35 and 12:55 UTC, so throughput during that 20-minute window is unmeasured.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model at 133.61 token/s average throughput, peaking at 171.99 token/s, while nemotron-3-ultra is the weakest at 2.13 token/s average, never exceeding 3.9 token/s.
  • The most operationally significant volatility is deepseek-v4-flash, which swung between a 146.07 token/s maximum and a 17.99 token/s minimum (26.0% CV); gemma4:31b shows the steepest decline, trending -29.3% from 96.14 token/s at 08:05 to 34.57 token/s at 12:00.
  • No missing-data limitation applies: all eight models recorded 48 of 48 expected samples, with 384 valid points and 100.0% coverage across the four-hour window.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model at 132.02 token/s average throughput, ahead of deepseek-v4-flash at 110.93 token/s and deepseek-v4-pro at 104.11 token/s; nemotron-3-ultra is weakest at 3.0 token/s average, far below glm-5.2's 22.83 token/s.
  • The most significant movement is a broad decline: gemma4:31b fell 23.2% to 37.87 token/s at 10:00, glm-5.2 fell 22.7%, and nemotron-3-ultra fell 48.1% to 1.35 token/s, while glm-5.3 held steady at +3.1%; deepseek-v4-flash showed sharp volatility, dropping to 17.99 token/s at 09:40.
  • No missing-data limitation applies: all eight models delivered 48 of 48 expected samples with 100.0% coverage, so the four-hour window is fully represented.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model at 132.26 token/s average throughput, peaking at 177.27 token/s; nemotron-3-ultra is the weakest at 4.03 token/s average, never exceeding 10.5 token/s.
  • Every model shows a negative trend over the window, with glm-5.2 down 28.3% and nemotron-3-ultra down 38.4%; glm-5.3 shows sharp isolated dips to 37.68, 41.49, and 47.03 token/s, and glm-5.2 fell to 6.6 token/s near 08:45.
  • No missing-data limitation applies: all eight models recorded 48 of 48 expected samples, with 384 valid points and 100.0% coverage, so the four-hour window is fully represented.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model by average throughput at 136.64 token/s (peaking at 179.99 token/s), while nemotron-3-ultra is the weakest at 6.03 token/s average and never exceeds 15.01 token/s.
  • The most operationally significant movement is nemotron-3-ultra's sustained decline of 51.2%, falling from 14.81 token/s at 04:05 to 1.59 token/s at 08:00, with the highest volatility (CV 60.5%); glm-5.3 also shows abrupt isolated dips, dropping to 37.68 token/s at 05:30.
  • No missing-data limitation applies: all eight models report 48 of 48 expected samples, 384 valid points total, and 100.0% coverage across the four-hour window.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model with an average throughput of 148.71 token/s (p95 177.34 token/s), while nemotron-3-ultra is the weakest at 11.70 token/s average, exceeding 83.15 token/s only once at 03:35.
  • The most operationally significant pattern is nemotron-3-ultra's sustained decline: from 29.88 token/s at 03:15 to below 8 token/s after 05:00, a 72.9% downward trend with 110.4% coefficient of variation, indicating progressive degradation rather than random fluctuation.
  • No missing-data limitation exists: all eight models report 48 of 48 expected samples with 100.0% coverage, totaling 384 valid points across the four-hour window.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model at 148.29 token/s average throughput, ahead of deepseek-v4-flash at 125.52 token/s; nemotron-3-ultra is the weakest at 14.18 token/s average, roughly a tenth of the leader.
  • The most operationally significant volatility is nemotron-3-ultra: coefficient of variation 104.7% with a 60% downward trend, sliding from an 83.15 token/s peak at 03:35 to single digits after 05:00 and 5.09 token/s at 06:00; glm-5.3 also showed an isolated dip to 37.68 token/s at 05:30.
  • No missing-data limitation applies: all eight models report 48 of 48 expected samples, 100% coverage, and 384 valid points across the four-hour window.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model with an average throughput of 152.12 token/s (p95 179.59 token/s), while nemotron-3-ultra is the weakest at 18.21 token/s average, roughly an eighth of the leader.
  • The most operationally significant volatility is nemotron-3-ultra's coefficient of variation of 111.2%, with throughput swinging between 1.38 and 83.15 token/s, including six consecutive readings between 1.38 and 1.49 token/s from 01:05 to 01:30 UTC; glm-5.3 also briefly dropped to 32.04 token/s at 02:00 UTC.
  • No missing-data limitation applies: all eight models report 48 of 48 expected samples, coverage is 100.0%, and the 384 valid points match the full four-hour window.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model, averaging 153.12 token/s with a p95 of 181.52 token/s, while nemotron-3-ultra is the weakest, averaging 18.10 token/s with a p95 of only 65.41 token/s.
  • The most operationally significant volatility is nemotron-3-ultra's coefficient of variation of 111.9%, swinging between 1.38 and 83.15 token/s, including a near-hour stretch below 2 token/s from roughly 01:05 to 01:30; glm-5.3 also showed a sharp single-interval dip to 32.04 token/s at 02:00.
  • No missing-data limitation applies: all eight models report 48 of 48 expected samples at 100.0% coverage, matching the dataset's 384 valid points.