← Performance dashboard

Hourly performance insights

Summaries of rolling four-hour performance data

1191 retained summaries
4-hour window · 84 points · 87.5% coverage
  • gemma4:31b is the strongest model with an average throughput of 123.88 token/s (peaking at 173.39 token/s), while nemotron-3-ultra is the weakest at 42.92 token/s average, ranging from 5.24 to 89.59 token/s.
  • The most significant volatility comes from glm-5.2 and nemotron-3-ultra: glm-5.2 fell from 108.13 token/s at 08:20 to 5.89 token/s at 12:00 (a -23.3% trend), and nemotron-3-ultra shows a 55.8% coefficient of variation with a -37.2% trend, ending at 19.72 token/s.
  • deepseek-v4-flash reported zero samples out of 12 expected, contributing to overall coverage of 87.5% (84 valid points), so its throughput cannot be assessed for this window.
4-hour window · 84 points · 87.5% coverage
  • Highest average throughput was gemma4:31b at 128.19 token/s (peak 173.39 token/s); lowest was nemotron-3-ultra at 47.10 token/s, which dipped to 8.55 token/s at 09:20.
  • glm-5.3-flash was the most volatile model, with a 44.7% coefficient of variation and swings between 50.40 and 172.07 token/s, including a drop from 172.07 to 57.67 token/s between 10:20 and 10:40; glm-5.3 trended down 35.6%, from 161.37 to 135.10 token/s.
  • deepseek-v4-flash produced no samples (0 of 12 expected, 0% coverage), lowering overall coverage to 87.5% with 84 valid points, leaving its throughput unknown for this window.
4-hour window · 84 points · 87.5% coverage
  • glm-5.3 is the strongest model by average throughput at 124.02 token/s, slightly ahead of gemma4:31b at 120.47 token/s, while nemotron-3-ultra is the weakest at 51.74 token/s, never exceeding 89.59 token/s in any observation.
  • glm-5.3-flash shows the most operationally significant volatility, with a coefficient of variation of 54.8%, swings between 13.34 and 185.16 token/s, and a 36.8% downward trend, including a drop from 184.90 token/s at 05:20 to 13.34 token/s at 05:40.
  • deepseek-v4-flash reported 0 of 12 expected samples (0% coverage), reducing overall coverage to 87.5% with 84 valid points, so its throughput cannot be assessed for this window.
4-hour window · 84 points · 87.5% coverage
  • glm-5.3 is the strongest model with an average throughput of 134.57 token/s (peaking at 186.42 token/s), while nemotron-3-ultra is the weakest at 53.02 token/s average, never exceeding 101.64 token/s across the window.
  • glm-5.3-flash shows the most operationally significant volatility, swinging between 13.34 and 187.66 token/s with a 55.9% coefficient of variation and a -29.4% trend, including a sharp drop to 13.34 token/s at 05:40; minimax-m3, by contrast, climbed 43.4% to a 176.75 token/s peak at 07:20.
  • deepseek-v4-flash reported no samples (0 of 12 expected, 0% coverage), so its throughput is unknown; overall dataset coverage is 87.5% (84 valid points), limiting conclusions about fleet-wide performance.
4-hour window · 84 points · 87.5% coverage
  • glm-5.3 is the strongest model with average throughput of 127.27 token/s (peak 186.42 token/s), while nemotron-3-ultra is the weakest at 50.47 token/s average, never exceeding 101.64 token/s.
  • glm-5.3-flash shows the most volatility, swinging between 13.34 and 187.66 token/s with a 53.3% coefficient of variation, including a drop to 13.34 token/s at 05:40 UTC; glm-5.3 also oscillates in alternating high-low cycles.
  • deepseek-v4-flash has no samples (0 of 12 expected), reducing overall coverage to 87.5% (84 valid points), so its throughput cannot be assessed for this window.
4-hour window · 84 points · 87.5% coverage
  • glm-5.3 delivered the highest average throughput at 119.53 token/s (p95 175.08 token/s), while nemotron-3-ultra was weakest at 43.96 token/s average with a minimum of 1.91 token/s.
  • glm-5.3-flash showed the most extreme volatility, ranging from 13.34 to 187.66 token/s with a coefficient of variation of 54.5%, including a collapse to 13.34 token/s at 05:40 after two readings near 185 token/s; nemotron-3-ultra was similarly unstable (cv 62.7%).
  • deepseek-v4-flash reported zero of 12 expected samples (0% coverage), reducing overall coverage to 87.5% (84 valid points), so fleet-wide comparisons exclude that model entirely.
4-hour window · 84 points · 87.5% coverage
  • glm-5.3 is the strongest model by average throughput at 121.08 token/s, while nemotron-3-ultra is the weakest at 39.06 token/s; the next-lowest averages are glm-5.2 at 59.71 token/s and minimax-m3 at 86.79 token/s.
  • The most operationally significant volatility is nemotron-3-ultra, which swings from a 1.91 token/s low to a 101.64 token/s high with a 72.6% coefficient of variation and a 77.3% trend over the window; gemma4:31b also shows a sustained decline, trending -22.4% and ending at 56.93 token/s versus its 192.34 token/s peak.
  • deepseek-v4-flash reported 0 of 12 expected samples (0% coverage), so its throughput is entirely unknown for this window, and overall dataset coverage is 87.5% with 84 valid points.
4-hour window · 84 points · 87.5% coverage
  • glm-5.3 posted the highest average throughput at 114.62 token/s, narrowly ahead of gemma4:31b at 112.21 token/s, while nemotron-3-ultra was weakest at 30.29 token/s average with a minimum of 1.91 token/s.
  • glm-5.3 was also the most volatile, swinging between 66.52 and 183.72 token/s (cv 42.3%), and glm-5.2 declined 34.1% across the window, ending at 24.53 token/s; nemotron-3-ultra dropped to 1.91 token/s at 02:40.
  • deepseek-v4-flash reported 0 of 12 expected samples (0.0% coverage), so its throughput is unknown; overall coverage was 87.5% (84 of 96 points), limiting window-wide comparisons.
4-hour window · 84 points · 87.5% coverage
  • glm-5.3 had the highest average throughput at 115.21 token/s, ahead of gemma4:31b (111.87 token/s) and glm-5.3-flash (110.96 token/s); nemotron-3-ultra was weakest at 31.65 token/s average, dipping to 1.91 token/s.
  • glm-5.3-flash fell hardest, trending -30.4% from a 201.63 token/s peak at 00:20 to 71.44 token/s at 03:00, while gemma4:31b rose 42.4% despite 41.7% coefficient of variation; deepseek-v4-pro was steadiest at 109.31 token/s average with 8.8% cv.
  • deepseek-v4-flash reported 0 of 12 expected samples, leaving 84 of 96 expected points (87.5% coverage), so its throughput is unknown for the entire window.
4-hour window · 84 points · 87.5% coverage
  • gemma4:31b posted the strongest average throughput at 116.59 token/s, peaking at 192.34 token/s and narrowly ahead of glm-5.3-flash at 116.00 token/s; nemotron-3-ultra was weakest at 33.27 token/s average, dipping to 4.57 token/s at 00:00 UTC.
  • glm-5.3 and glm-5.3-flash showed the most operationally significant volatility, with coefficients of variation of 44.7% and 43.7%, swinging between roughly 65 and 184 token/s and 62 and 202 token/s respectively across twenty-minute intervals.
  • deepseek-v4-flash reported zero of its 12 expected samples (0% coverage), leaving 84 valid points out of 96 expected (87.5% overall coverage), so its throughput cannot be assessed for this window.
4-hour window · 84 points · 87.5% coverage
  • glm-5.3-flash is the strongest model at 114.38 token/s average throughput, peaking at 201.63 token/s, while nemotron-3-ultra is the weakest at 33.73 token/s average, dipping as low as 4.57 token/s.
  • glm-5.3 shows the most operationally significant volatility, swinging between roughly 65 and 184 token/s with a 46.5% coefficient of variation; gemma4:31b declined 20.2% across the window, ending at 41.79 token/s versus a 104.01 token/s average.
  • deepseek-v4-flash returned no samples (0 of 12 expected, 0% coverage), leaving overall coverage at 87.5% (84 of 96 valid points), so its throughput cannot be assessed for this period.
4-hour window · 84 points · 87.5% coverage
  • glm-5.3 delivered the highest average throughput at 112.03 token/s (peaking at 181.05 token/s), while nemotron-3-ultra was weakest at 30.94 token/s average, dipping as low as 3.22 token/s.
  • The most operationally significant movement is gemma4:31b's upward trend of 94.1%, climbing from 29.98 token/s at 19:20 to a 161.38 token/s peak at 22:40, though it fell back to 66.53 token/s by 23:00; glm-5.3 also swung sharply between 64.28 and 181.05 token/s (stddev 46.84, CV 41.8%).
  • deepseek-v4-flash reported zero samples (0 of 12 expected, 0.0% coverage), so its throughput is unmeasured this window; overall dataset coverage is 87.5% with 84 valid points, limiting any conclusions about that model.
4-hour window · 84 points · 87.5% coverage
  • glm-5.3 is the strongest model at 113.05 token/s average throughput, while nemotron-3-ultra is the weakest at 27.33 token/s average, roughly a quarter of the leader's pace.
  • glm-5.2 shows the most operationally significant decline, with a -51.0% trend from 87.92 token/s at 18:20 to 70.42 token/s at 22:00, including a drop to 5.53 token/s at 20:20 and readings below 36 token/s for five consecutive intervals from 20:20 to 21:40.
  • deepseek-v4-flash reported no samples (0 of 12 expected, 0.0% coverage), so its throughput is unknown; overall coverage is 87.5% with 84 valid points.
4-hour window · 84 points · 87.5% coverage
  • glm-5.3 is the strongest model at 119.52 token/s average throughput, while nemotron-3-ultra is the weakest at 26.72 token/s average, with a low of 3.22 token/s at 20:20.
  • glm-5.2 shows the highest volatility at 50.1% coefficient of variation, dropping to 5.53 token/s at 20:20, while deepseek-v4-flash trended up 43.2% and spiked to 182.29 token/s at 20:00.
  • deepseek-v4-flash reported zero samples across all 12 expected observations, leaving its throughput unmeasured and reducing overall coverage to 87.5% with 84 valid points.
4-hour window · 84 points · 87.5% coverage
  • glm-5.3 delivered the highest average throughput at 111.39 token/s (peak 181.05 token/s at 19:40), while nemotron-3-ultra was weakest at 30.88 token/s average, peaking at only 56.94 token/s.
  • The sharpest volatility came from glm-5.3-flash, which swung from 16.24 token/s at 19:00 to 182.29 token/s at 20:00; minimax-m3 also trended down 21.1% across the window despite peaking at 145.78 token/s at 18:00.
  • deepseek-v4-flash reported zero of 12 expected samples (0% coverage), lowering overall coverage to 87.5% with 84 valid points, so its throughput is unknown for this period.
4-hour window · 84 points · 87.5% coverage
  • glm-5.3 delivered the strongest average throughput at 97.26 token/s (peak 165.61 token/s at 18:00), while nemotron-3-ultra was weakest at 29.01 token/s average, never exceeding 56.94 token/s.
  • Volatility is the dominant operational signal: deepseek-v4-pro swung between 17.11 and 115.76 token/s (cv 46.8%), and glm-5.3 showed the steepest trend at +40.6%, climbing from roughly 70 token/s early in the window to 165.61 token/s at 18:00 before settling at 115.84 token/s.
  • deepseek-v4-flash reported zero of 12 expected samples (0% coverage), so its throughput is unknown; overall coverage is 87.5% (84 of 96 expected points), limiting conclusions about fleet-wide performance.
4-hour window · 84 points · 87.5% coverage
  • minimax-m3 leads average throughput at 100.32 token/s (12 of 12 samples), just ahead of glm-5.3 at 96.05 token/s and gemma4:31b at 95.18 token/s; nemotron-3-ultra is weakest at 34.67 token/s, below glm-5.2's 48.55 token/s.
  • glm-5.3 shows the sharpest upward trend at +29.0%, closing at 165.61 token/s at 18:00, while deepseek-v4-pro fell -24.9% to a period-low 17.11 token/s; deepseek-v4-pro also has the highest volatility at 41.2% cv.
  • deepseek-v4-flash reported 0 of 12 samples (0% coverage), lowering overall coverage to 87.5% (84 valid points), so its throughput is unknown and rankings exclude it.
4-hour window · 84 points · 87.5% coverage
  • Strongest average throughput was gemma4:31b at 96.61 token/s; weakest was nemotron-3-ultra at 39.56 token/s, which also ranged widely from 5.64 to 91.22 token/s.
  • The sharpest operational movement is glm-5.3-flash, down 38.0% from a 192.62 token/s peak at 13:20 to 61.47 token/s at 17:00, while minimax-m3 improved 24.6% to 110.38 token/s; nemotron-3-ultra was the most volatile model with a 57.5% coefficient of variation.
  • deepseek-v4-flash returned 0 of 12 expected samples (0% coverage), lowering overall coverage to 87.5% with 84 valid points, so its throughput is unknown for this window.
4-hour window · 84 points · 87.5% coverage
  • gemma4:31b delivered the highest average throughput at 111.26 token/s, while nemotron-3-ultra was weakest at 43.39 token/s; glm-5.2 also averaged low at 50.81 token/s.
  • The sharpest operational movement was glm-5.3-flash, down 53.7% over the window, falling from a 192.62 token/s peak to roughly 56.83–80.49 token/s in the final hour; nemotron-3-ultra fell 41.0% and hit a 5.64 token/s low at 14:00. glm-5.3 stayed flat at +0.7%.
  • deepseek-v4-flash reported zero of 12 expected samples (0% coverage), leaving 84 of 96 expected points valid (87.5% coverage) and preventing any throughput assessment for that model.
4-hour window · 84 points · 87.5% coverage
  • gemma4:31b posted the highest average throughput at 113.88 token/s, while nemotron-3-ultra was the weakest at 50.12 token/s; deepseek-v4-flash cannot be ranked because it reported no samples.
  • The most significant movement is minimax-m3's 36.5% decline across the window, falling from a 208.7 token/s peak at 12:20 to 85.39 token/s at 15:00; glm-5.2 also swung widely between 26.12 and 132.35 token/s with a 56.5% coefficient of variation.
  • deepseek-v4-flash returned 0 of 12 expected samples (0% coverage), dropping overall dataset coverage to 87.5% with 84 valid points, so fleet-level comparisons exclude that model entirely.
4-hour window · 84 points · 87.5% coverage
  • minimax-m3 is the strongest model at 112.84 token/s average throughput, edging gemma4:31b (111.79 token/s) and glm-5.3 (110.17 token/s); glm-5.2 is weakest at 48.84 token/s average, below nemotron-3-ultra's 51.11 token/s.
  • glm-5.3-flash shows the sharpest movement, trending up 138.1% and swinging from 46.99 token/s at 12:00 to 192.62 token/s at 13:20, while glm-5.3 fell 39.3% and glm-5.2 has the highest volatility at 62.0% coefficient of variation.
  • deepseek-v4-flash returned no samples (0 of 12 expected, 0.0% coverage), so its throughput is unknown; overall coverage is 87.5% with 84 of 96 expected points, limiting fleet-wide conclusions.
4-hour window · 84 points · 87.5% coverage
  • Strongest average throughput was gemma4:31b at 121.0 token/s (peaking at 177.4 token/s), while glm-5.2 was weakest at 48.89 token/s average, never exceeding 78.38 token/s.
  • Volatility is the dominant operational signal: glm-5.3-flash swung between 46.99 and 187.11 token/s (49.0% coefficient of variation), and minimax-m3 ranged from 29.23 to 208.7 token/s (44.0% CV), making capacity planning on these models unreliable.
  • deepseek-v4-flash reported zero of 12 expected samples (0.0% coverage), so its throughput is unknown; overall dataset coverage is 87.5% with 84 valid points, limiting conclusions about fleet-wide behavior.
4-hour window · 84 points · 87.5% coverage
  • glm-5.3 is the strongest model by average throughput at 134.6 token/s (p95 183.96 token/s), while glm-5.2 is the weakest at 48.91 token/s average, never exceeding 78.38 token/s across the window.
  • The most operationally significant volatility is glm-5.3-flash, which fell 44.5% over the window, from 186.42 token/s at 08:20 to 46.99 token/s at 12:00, with the highest variability of any model (cv 51.7%, stddev 44.41 token/s). minimax-m3 also swung between 29.23 and 210.64 token/s.
  • deepseek-v4-flash reported zero of 12 expected samples (0% coverage), reducing overall dataset coverage to 87.5% (84 of 96 expected points), so its throughput cannot be assessed for this period.
4-hour window · 84 points · 87.5% coverage
  • glm-5.3 is the strongest model with average throughput of 140.71 token/s (p95 183.96 token/s); nemotron-3-ultra is the weakest at 49.21 token/s average, with a minimum of 3.6 token/s.
  • The most operationally significant movement is nemotron-3-ultra's 110.4% upward trend, rising from 3.6 token/s at 08:00 to a 93.09 token/s peak at 09:20; glm-5.3-flash shows the highest volatility at 54.1% coefficient of variation.
  • deepseek-v4-flash returned zero samples against 12 expected intervals (0% coverage), lowering overall coverage to 87.5% (84 of 96 valid points), so its throughput cannot be evaluated this period.
4-hour window · 87 points · 90.6% coverage
  • glm-5.3 is the strongest model by average throughput at 117.29 token/s, while nemotron-3-ultra is the weakest at 52.75 token/s; glm-5.3-flash and gemma4:31b follow at 109.80 and 103.04 token/s respectively.
  • The most operationally significant movement is nemotron-3-ultra's recovery from 3.6 token/s at 08:00 to 90.25 token/s at 10:00, an 86.0% trend with the dataset's highest coefficient of variation at 55.4%; minimax-m3 also dropped sharply to 29.23 token/s in its latest sample.
  • deepseek-v4-flash has only 3 of 12 expected samples (25.0% coverage), so its 98.14 token/s average is unreliable; overall coverage is 90.6% with 87 valid points.
4-hour window · 90 points · 93.8% coverage
  • Strongest average throughput was glm-5.3 at 115.06 token/s, edging minimax-m3 at 108.52 token/s, while nemotron-3-ultra was weakest at 38.64 token/s, below glm-5.2 (53.74) and deepseek-v4-pro (53.97 token/s).
  • The most significant movement was glm-5.3's climb of 95.2 percent, rising from a 23.09 token/s low at 06:00 to 183.08 token/s at 08:20 and holding 143.17 token/s at 09:00; nemotron-3-ultra fell 30.4 percent with the highest volatility (cv 62.5 percent) and a 3.6 token/s floor.
  • deepseek-v4-flash reported only 6 of 12 expected samples (50.0 percent coverage), with no observations after 07:00, so its 75.33 token/s average is unreliable; overall coverage was 93.8 percent (90 valid points).
4-hour window · 93 points · 96.9% coverage
  • glm-5.3 posted the highest average throughput at 116.43 token/s, while nemotron-3-ultra was weakest at 44.78 token/s; gemma4:31b and minimax-m3 followed at 102.93 and 99.03 token/s respectively.
  • glm-5.3 showed the sharpest operational swing: near 171-176 token/s through 05:20, a drop to 23.09 token/s at 06:00, then recovery to 176.17 token/s by 08:00; nemotron-3-ultra ranged 3.6 to 95.15 token/s with 72.1% CV.
  • deepseek-v4-flash has only 9 of 12 samples (75% coverage) with no observations after 07:00, so its 65.3 token/s average may not represent the full window; dataset-wide valid points total 93 of 96.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model on average throughput at 116.26 token/s, ahead of glm-5.3 at 109.32 token/s; nemotron-3-ultra is the weakest at 52.70 token/s, below deepseek-v4-flash at 63.88 token/s.
  • The most significant event is a broad throughput collapse after 05:40 UTC: glm-5.3 fell from 171.75 token/s at 05:20 to 23.09 token/s at 06:00, deepseek-v4-pro ended at 6.35 token/s, and glm-5.2 showed the highest volatility with a 75.7% coefficient of variation.
  • No missing-data limitation applies: all 96 expected points are valid, coverage is 100.0%, and each of the eight models has 12 of 12 samples.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 129.46 token/s average throughput, with a peak of 163.86 token/s; nemotron-3-ultra is the weakest at 42.74 token/s average, ranging from 9.95 to 95.15 token/s.
  • glm-5.2 shows the highest volatility (75.4% CV, stddev 45.09 token/s), swinging from 202.82 token/s at 04:00 to 24.84 token/s at 06:00; most models also dipped sharply at 06:00, e.g. glm-5.3 fell to 23.09 token/s.
  • No missing-data limitation: all eight models have 12 of 12 samples, 96 valid points, and 100.0% coverage over the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 delivered the highest average throughput at 129.42 token/s (peaking at 184.08 token/s), narrowly ahead of gemma4:31b at 127.19 token/s, while nemotron-3-ultra was weakest at 42.12 token/s average despite recovering to 95.15 token/s at the final sample.
  • deepseek-v4-flash showed the most operationally significant volatility, falling from 118.77 token/s at 01:20 to 35.99 token/s at 05:00 (trend -47.4%, cv 48.6%); glm-5.2 also swung from a 202.82 token/s spike at 04:00 down to 32.18 token/s by 05:00.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, 96 valid points, and 100.0% coverage over the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model by average throughput at 128.71 token/s (range 64.11–163.86 token/s), while nemotron-3-ultra is the weakest at 39.22 token/s, dipping as low as 9.95 token/s.
  • glm-5.2 shows the most operationally significant volatility: its coefficient of variation is 74.4%, and its final observation jumps from 32.85 to 202.82 token/s, a single-sample spike far above its 62.0 token/s average.
  • No missing-data limitation applies: all eight models recorded 12 of 12 expected samples, and overall coverage is 100.0% with 96 valid points, so the four-hour window is fully represented.
4-hour window · 96 points · 100.0% coverage
  • Highest average throughput was gemma4:31b at 124.22 token/s (peak 162.63 token/s); lowest was nemotron-3-ultra at 37.68 token/s, with a floor of 9.95 token/s, roughly a third of the leader's mean.
  • The most operationally significant pattern is a synchronized dip at 02:00 UTC, where gemma4:31b fell to 64.11 token/s and deepseek-v4-flash to 41.75 token/s; glm-5.3 is the most volatile model (CV 51.9%), ranging 18.34 to 178.32 token/s, while nemotron-3-ultra trended down 52.1% to 35.33 token/s.
  • No missing-data limitation applies: all eight models delivered 12 of 12 expected samples, yielding 96 valid points and 100.0% coverage, so every percentage and average rests on complete data.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model by average throughput at 124.05 token/s (peak 154.25 token/s), while nemotron-3-ultra is the weakest at 41.48 token/s, roughly a third of the leader's average.
  • glm-5.2 shows the most operationally significant volatility, with a coefficient of variation of 59.7% and a -43.7% trend, swinging between 150.11 and 29.51 token/s; glm-5.3 is similarly unstable at 55.4% CV, dipping to 18.34 token/s at 00:20.
  • No missing-data limitation exists in this window: all 8 models report 12 of 12 expected samples, giving 96 valid points and 100.0% coverage across the four-hour period.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 120.22 token/s average throughput, while nemotron-3-ultra is the weakest at 43.48 token/s average; deepseek-v4-flash ranks second at 90.26 token/s.
  • glm-5.3 shows the sharpest volatility: it jumped from 68.34 token/s at 23:00 to 178.32 token/s at 23:40, then fell to 18.34 token/s at 00:20, and glm-5.2 has the highest coefficient of variation at 70.0%.
  • No missing-data limitation applies: all eight models have 12 of 12 expected samples, valid_point_count is 96, and coverage is 100.0%, so the four-hour window is fully covered.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 119.74 token/s average throughput, while nemotron-3-ultra is the weakest at 47.76 token/s average, with a low of 7.27 token/s at 23:00 UTC.
  • glm-5.2 shows the most operationally significant volatility: coefficient of variation 69.6%, swings from 15.29 token/s at 21:40 to 150.11 token/s at 22:40, and a 133.5% trend ending at 148.55 token/s; nemotron-3-ultra also declined 35.8% to 17.92 token/s.
  • No missing-data limitation exists: all eight models report 12 of 12 samples, 96 valid points, and 100.0% coverage over the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput over the four-hour window: gemma4:31b at 110.39 token/s; weakest: glm-5.2 at 50.76 token/s, with nemotron-3-ultra close behind at 53.16 token/s.
  • glm-5.2 shows the sharpest decline, trending -42.0% to 26.15 token/s at 22:00, while nemotron-3-ultra is the most volatile (cv 64.2%, range 11.58–113.27 token/s); glm-5.3-flash and deepseek-v4-flash improved 22.1% and 24.7%.
  • No missing-data limitation is present: all 8 models delivered 12 of 12 expected samples, giving 96 valid points and 100.0% coverage.
4-hour window · 96 points · 100.0% coverage
  • Highest average throughput was gemma4:31b at 114.01 token/s (p95 157.09 token/s); lowest was nemotron-3-ultra at 59.26 token/s, ranging from 11.58 to 113.27 token/s.
  • The clearest trend is glm-5.3's late-window acceleration: up 49.7% overall to a 97.52 token/s average, sustaining 125.67 to 142.75 token/s after 19:00 UTC, while glm-5.2 fell 10.4% to a 64.58 token/s average. Volatility remained high, with nemotron-3-ultra's coefficient of variation at 56.0% and minimax-m3 swinging between 25.96 and 153.01 token/s.
  • No missing-data limitation applies: all eight models reported 12 of 12 expected samples, totaling 96 valid points at 100.0% coverage, so no figures are affected by gaps.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput is gemma4:31b at 116.76 token/s across 12 samples (range 55.05 to 168.89 token/s); weakest is nemotron-3-ultra at 46.08 token/s (range 17.66 to 88.83 token/s).
  • The most operationally significant volatility is nemotron-3-ultra, with a 55.4% coefficient of variation and swings between lows of 17.66 to 25.88 token/s and highs of 71.82 to 88.83 token/s; deepseek-v4-flash also swung from 33.63 to 120.71 token/s with 48.4% variation and a 49.0% trend.
  • No missing-data limitation applies: all 8 models reported 12 of 12 expected samples, yielding 96 valid points and 100.0% coverage over the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • Highest average throughput was gemma4:31b at 116.44 token/s; lowest was nemotron-3-ultra at 36.77 token/s, roughly a third of the leader's rate.
  • glm-5.3 showed the sharpest decline, trending -56.9% from a 175.58 token/s peak at 14:40 to 63.62 token/s at 18:00, while nemotron-3-ultra was the most volatile, with a 69.8% coefficient of variation and a range of 6.98 to 88.83 token/s.
  • No missing-data limitation applies: all eight models delivered 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • Highest average throughput was gemma4:31b at 116.38 token/s; lowest was nemotron-3-ultra at 24.96 token/s, with glm-5.2 also near the bottom at 43.39 token/s.
  • glm-5.3 shows the sharpest deterioration, trending -45.1% from 139.25 token/s at 13:20 to 59.15 token/s at 17:00, including a low of 13.27 token/s at 16:20; minimax-m3 was the most volatile at 45.2% CV, swinging between 179.88 and 50.72 token/s.
  • No missing-data limitation applies: all 8 models have 12 of 12 samples, 96 valid points, and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model at 127.76 token/s average throughput, ahead of gemma4:31b at 112.19 token/s, while nemotron-3-ultra is the weakest at 18.81 token/s average and never exceeded 36.15 token/s.
  • deepseek-v4-flash shows the sharpest deterioration, trending down 49.6% from a 129.22 token/s peak at 12:40 to 45.26 token/s at 16:00; glm-5.2 is the most volatile, with a 60.5% coefficient of variation and swings between 13.06 and 92.01 token/s.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, and the dataset records 96 valid points with 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model with an average throughput of 128.58 token/s (p95 176.25 token/s), while nemotron-3-ultra is the weakest at 14.35 token/s average, never exceeding 30.07 token/s.
  • The most operationally significant volatility is deepseek-v4-pro dropping to 7.08 token/s at 12:00 and 8.45 token/s at 14:00 despite an 83.08 token/s maximum; glm-5.2 shows the highest relative variability at 51.1% CV, and deepseek-v4-flash declined 28.6% over the window.
  • No missing-data limitation applies: all 8 models recorded 12 of 12 samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model at 111.45 token/s average throughput, peaking at 177.07 token/s at 14:00; nemotron-3-ultra is the weakest at 16.02 token/s average, never exceeding 30.07 token/s.
  • deepseek-v4-pro shows the most operationally significant volatility, dropping to 7.08 token/s at 12:00 and 8.45 token/s at 14:00 against an 85.3 token/s maximum; minimax-m3 also swung from roughly 72 token/s early to 179.88 token/s at 13:20, a 35.0% trend.
  • No missing-data limitation applies: all eight models delivered 12 of 12 expected samples at 100.0% coverage, and the 96 valid points match full coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 delivered the highest average throughput at 95.49 token/s, narrowly ahead of gemma4:31b at 95.02 token/s, while nemotron-3-ultra was weakest at 14.29 token/s, roughly one-seventh of the leader's pace.
  • glm-5.3 showed the sharpest operational swings, ranging from 41.73 to 151.27 token/s with a 40.7% coefficient of variation, including a drop to 41.73 token/s at 12:40 followed by recovery to 134.02 token/s; deepseek-v4-pro also dipped to 7.08 token/s at 12:00.
  • No missing-data limitation applies: all eight models recorded 12 of 12 samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput was gemma4:31b at 103.49 token/s (peaking at 157.36 token/s), while nemotron-3-ultra was weakest at 14.29 token/s and never exceeded 30.49 token/s.
  • The most operationally significant event is deepseek-v4-pro collapsing from 83.08 token/s at 11:40 to 7.08 token/s at 12:00, its interval minimum; glm-5.3-flash also swung from a 195.03 token/s peak at 09:20 down to 55.53 token/s at 11:20, a -35.9% trend.
  • No missing data exists: all 8 models report 12 of 12 samples, 96 valid points, and 100.0% coverage; the limitation is that results cover only this four-hour window at 20-minute intervals.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 103.2 token/s average throughput, ahead of glm-5.3-flash at 98.13 token/s and glm-5.3 at 91.05 token/s; nemotron-3-ultra is the weakest at 14.08 token/s average, roughly one-seventh of gemma4:31b.
  • glm-5.3-flash shows the widest volatility (cv 46.6%), swinging from 52.25 to 195.03 token/s, while glm-5.3 fell from a 153.55 token/s peak at 08:40 to 59.16 token/s at 10:40 before rebounding to 141.31 token/s at 11:00; gemma4:31b also dropped sharply to 36.33 token/s at 09:20.
  • No missing-data limitation exists: all eight models have 12 of 12 samples, 96 valid points, and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 104.29 token/s average throughput, peaking at 155.77 token/s, while nemotron-3-ultra is the weakest at 15.93 token/s average and never exceeding 38.0 token/s.
  • glm-5.3-flash shows the sharpest volatility, swinging from 52.25 token/s at 07:40 to 195.03 token/s at 09:20 and back to 65.19 token/s by 10:00, with a 48.3% coefficient of variation; nemotron-3-ultra also declines 28.3% over the window.
  • No missing-data limitation applies: all eight models have 12 of 12 samples at 100.0% coverage, with 96 valid points, so the main constraint is the four-hour window itself.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 120.45 token/s average throughput, ahead of glm-5.3 at 97.45 token/s; nemotron-3-ultra is the weakest at 20.18 token/s average, with a maximum of only 38.0 token/s.
  • glm-5.3 shows the highest volatility (cv 50.5%), swinging between 16.28 and 183.18 token/s, while deepseek-v4-flash recorded the steepest upward trend at +67.4%, ending at 88.56 token/s; nemotron-3-ultra fell 42.0% to 5.95 token/s.
  • No missing-data limitation applies: all eight models delivered 12 of 12 expected samples, totaling 96 valid points at 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput was gemma4:31b at 119.61 token/s (peak 162.1 token/s); weakest was nemotron-3-ultra at 21.46 token/s (peak 38.0 token/s).
  • glm-5.3 was the most volatile model, with a coefficient of variation of 51.3% and swings from 16.28 to 183.18 token/s; deepseek-v4-flash was similarly unstable at 51.4% cv, ranging 24.81 to 128.82 token/s.
  • No missing-data limitation applies: all eight models delivered 12 of 12 expected samples, giving 96 valid points and 100.0% coverage over the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 122.49 token/s average throughput, ahead of glm-5.3 at 102.07 token/s; nemotron-3-ultra is the weakest at 27.83 token/s average, never exceeding 55.94 token/s.
  • deepseek-v4-flash shows the steepest decline at -40.1% trend, and deepseek-v4-pro dropped to 11.76 token/s at 07:00; glm-5.3 and deepseek-v4-flash share the highest volatility at 51.9% coefficient of variation.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0%.