← Performance dashboard

Hourly performance insights

Summaries of rolling four-hour performance data

1199 retained summaries
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model at 111.82 token/s average throughput, narrowly ahead of gemma4:31b at 110.42 token/s and deepseek-v4-pro at 108.16 token/s; nemotron-3-ultra is the weakest at 3.71 token/s average, never exceeding 6.96 token/s.
  • glm-5.3 shows the highest volatility (cv 45.7%), ranging from 64.79 to 189.77 token/s, and glm-5.3-flash dropped to 13.93 token/s at 04:00 before rebounding to 129.99 token/s at 04:20; deepseek-v4-flash trended up 36.4% to a 150.70 token/s peak.
  • No samples are missing: all eight models report 12 of 12 expected observations, 96 valid points, and 100.0% coverage, so the only limitation is the short four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model at 114.27 token/s average throughput, ahead of deepseek-v4-pro (107.36 token/s) and gemma4:31b (105.74 token/s); nemotron-3-ultra is the weakest at 2.67 token/s average, never exceeding 3.94 token/s.
  • The most operationally significant movement is minimax-m3's decline of 38.6% across the window, from a 120.57 token/s peak at 00:40 to 37.39 token/s at 04:00; deepseek-v4-flash also swung widely between 43.15 and 192.88 token/s.
  • No missing-data limitation applies: all eight models report 12 of 12 expected samples, totaling 96 valid points at 100.0% coverage over the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model at 116.95 token/s average throughput, peaking at 192.88 token/s, while nemotron-3-ultra is the weakest at 5.85 token/s average, with a maximum of only 21.97 token/s.
  • nemotron-3-ultra shows the most operationally significant volatility, with a 112.2% coefficient of variation and a -66.9% trend, falling from 21.97 token/s at 23:40 to 3.22 token/s at 03:00; minimax-m3 also declined 36.8% to 38.91 token/s, while deepseek-v4-pro improved 25.5% to 108.65 token/s.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model at 130.25 token/s average throughput, while nemotron-3-ultra is the weakest at 8.59 token/s average, roughly 15 times slower than the leader.
  • The most operationally significant movement is nemotron-3-ultra's collapse: it peaked at 21.97 token/s at 23:40 UTC, then fell to 1.49 token/s at 00:40 and ended at 1.9 token/s, an 82.5 percent decline over the window. glm-5.3-flash also dropped 40.9 percent, from a 167.64 token/s high to 59.61 token/s.
  • No missing-data limitation applies: all eight models have 12 of 12 expected samples, 96 valid points, and 100.0 percent coverage across the four-hour period.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model at 142.85 token/s average throughput (range 96.27 to 192.88 token/s), while nemotron-3-ultra is the weakest at 9.08 token/s average, peaking at only 21.97 token/s.
  • glm-5.3-flash shows the sharpest decline, trending down 36.6% from 168.28 token/s at 21:20 to 49.37 token/s at 01:00, including a low of 22.34 token/s at 00:40; glm-5.3 also swings between 65.45 and 203.52 token/s.
  • No missing-data limitation: all 8 models report 12 of 12 samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model by average throughput at 135.74 token/s (range 101.32–192.84 token/s), while nemotron-3-ultra is the weakest at 9.47 token/s average, peaking at only 21.97 token/s.
  • The most operationally significant volatility is glm-5.3, which swings between roughly 65 and 200 token/s across the window (coefficient of variation 45.0%), including a drop to 63.95 token/s at 20:20 and a spike to 203.52 token/s at 00:00; deepseek-v4-pro also dipped to 20.24 token/s at 23:40 before recovering to 117.54 token/s.
  • No missing-data limitation exists: all 96 expected points are valid (100.0% coverage), and every model has 12 of 12 samples.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model at 125.65 token/s average throughput (range 100.1–158.48 token/s), while nemotron-3-ultra is the weakest at 3.52 token/s average, never exceeding 7.77 token/s.
  • glm-5.3 shows the most significant shift, trending up 89.8% from roughly 63 token/s early in the window to 174.08 token/s at 22:00, though it is also the most volatile model at 47.1% coefficient of variation; gemma4:31b declined 23.2%, dipping to 35.64 token/s at 21:20.
  • No missing-data limitation applies: all eight models report 12 of 12 expected samples, totaling 96 valid points at 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model by average throughput at 118.59 token/s, ahead of deepseek-v4-flash at 128.04 token/s only if latest values are used, while nemotron-3-ultra is the weakest at 2.82 token/s average, peaking at just 5.89 token/s.
  • glm-5.3 shows the most operationally significant volatility, with a 50.2% coefficient of variation and a swing from 54.02 to 199.8 token/s, including a drop from 199.8 to 77.83 token/s between 17:20 and 17:40; minimax-m3 also dipped to 22.02 token/s at 20:20.
  • No missing-data limitation exists in this window: all 8 models report 12 of 12 expected samples, valid_point_count is 96, and coverage is 100.0%, so every twenty-minute observation from 17:20 to 21:00 UTC is present.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model at 124.27 token/s average throughput, peaking at 158.48 token/s; nemotron-3-ultra is the weakest at 2.36 token/s average, never exceeding 4.07 token/s.
  • glm-5.3 shows the most operationally significant volatility: a spike to 199.8 token/s at 17:20 UTC followed by a drop to 54.02 token/s at 18:00, with 46.5% coefficient of variation and a -28.6% trend over the window.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0%, so the four-hour window is fully represented.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model at 120.47 token/s average throughput (range 84.72–153.83 token/s), ahead of gemma4:31b at 111.06 token/s; nemotron-3-ultra is the weakest at 2.46 token/s average, never exceeding 4.86 token/s.
  • glm-5.3 shows the sharpest decline, dropping from roughly 142–146 token/s between 15:20 and 16:00 to 54.02–63.57 token/s after 18:00, a -25.4% trend; minimax-m3 is the most volatile, spiking to 147.75 token/s at 16:00 against a 37.88 token/s low with 50.3% coefficient of variation.
  • No missing-data limitation exists: all eight models report 12 of 12 expected samples, totaling 96 valid points at 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model at 114.1 token/s average throughput, edging out glm-5.3 at 106.18 token/s, while nemotron-3-ultra is the weakest at 3.32 token/s average, never exceeding 14.68 token/s in any observation.
  • The most operationally significant volatility is nemotron-3-ultra's coefficient of variation of 106.7% with a -57.6% trend, plus a single 147.75 token/s spike for minimax-m3 at 16:00 against its 54.05 token/s average, indicating unstable capacity rather than steady degradation.
  • No missing-data limitation exists: all 8 models report 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window, so every percentage and count is fully supported by the dataset.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model with average throughput of 102.8 token/s (12 samples, range 27.65–175.53 token/s), while nemotron-3-ultra is the weakest at 3.37 token/s average, never exceeding 14.68 token/s.
  • The most significant volatility is nemotron-3-ultra's coefficient of variation of 104.7%, and glm-5.2 shows the steepest decline, falling 57.2% from 128.66 to 19.46 token/s; deepseek-v4-flash improved 40.7% despite dipping to 14.58 token/s at 14:20.
  • No missing data: all 8 models report 12 of 12 samples, 96 valid points, and 100.0% coverage; the limitation is the short four-hour window with only 12 observations per model.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput belongs to glm-5.3 at 112.91 token/s (p95 169.71 token/s, maximum 175.53 token/s); weakest is nemotron-3-ultra at 3.41 token/s average, never exceeding 14.68 token/s.
  • The most operationally significant volatility is minimax-m3 jumping from 47.58 token/s at 15:40 to 147.75 token/s at 16:00, roughly triple its 47.15 token/s average (cv 66.1%), while glm-5.2 declined from a 155.03 token/s peak at 13:40 to 13.08 token/s at 16:00.
  • No missing-data limitation applies: all eight models report 12 of 12 samples, 96 valid points total, and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3-flash is the strongest model at 103.71 token/s average throughput, ahead of glm-5.3 (99.83 token/s) and gemma4:31b (90.9 token/s); nemotron-3-ultra is the weakest at 3.17 token/s average, far below every other model.
  • glm-5.3 shows the widest volatility, with a 49.2% coefficient of variation and swings between 27.65 and 175.53 token/s; deepseek-v4-flash dropped to 14.58 token/s at 14:20 before recovering to 143.5 token/s at 14:40, and deepseek-v4-pro declined 30.9% overall.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3-flash is the strongest model at 108.32 token/s average throughput, ahead of deepseek-v4-flash at 103.96 token/s; nemotron-3-ultra is the weakest at 2.68 token/s average, roughly 40x slower than the leader.
  • glm-5.3 shows the most operationally significant volatility, with a 50.3% coefficient of variation and swings from 27.65 to 175.53 token/s, including a drop to its minimum at 14:00 UTC; deepseek-v4-flash declined 19.3% over the window, ending at 66.32 token/s.
  • No missing-data limitation: all eight models have 12 of 12 samples, 96 valid points, and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3-flash is the strongest model at 115.86 token/s average throughput, ahead of deepseek-v4-flash (107.38 token/s); nemotron-3-ultra is weakest at 2.85 token/s average, peaking at only 5.67 token/s.
  • glm-5.3 shows the most operationally significant volatility, ranging from 11.63 to 164.94 token/s with a coefficient of variation of 51.2% and a +56.0% trend; gemma4:31b also swung between 62.08 and 157.65 token/s.
  • No missing-data limitation: all 8 models delivered 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model at 113.83 token/s average throughput, narrowly ahead of glm-5.3-flash at 112.87 token/s; nemotron-3-ultra is the weakest at 2.80 token/s average, never exceeding 5.67 token/s.
  • gemma4:31b shows the sharpest decline, trending down 34.8% from 118.54 token/s at 08:20 to 70.34 token/s at 12:00, while glm-5.3 is the most volatile, swinging between 11.63 and 160.44 token/s with a 45.9% coefficient of variation.
  • No missing-data limitation applies: all 8 models recorded 12 of 12 expected samples, totaling 96 valid points at 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model by average throughput at 113.25 token/s, narrowly ahead of deepseek-v4-flash at 113.34 token/s... correction: deepseek-v4-flash leads at 113.34 token/s, with gemma4:31b at 113.25 token/s; nemotron-3-ultra is weakest at 2.68 token/s average.
  • Volatility is the dominant operational signal: minimax-m3 shows the highest relative variability (55.1% coefficient of variation, ranging 14.4 to 93.2 token/s), while glm-5.3 declined 22.7% over the window, falling from a 157.48 token/s peak to 76.37 token/s at the latest sample.
  • No missing-data limitation exists: all eight models have 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model at 118.55 token/s average throughput, ahead of glm-5.3-flash (103.75 token/s) and glm-5.3 (101.23 token/s); nemotron-3-ultra is the weakest at 2.09 token/s average, never exceeding 4.07 token/s.
  • glm-5.3 shows the sharpest decline, trend -48.9%, falling from a 174.45 token/s maximum at 06:40 to 11.63 token/s at 09:40 before rebounding to 120.26 token/s; minimax-m3 is the most volatile model (cv 57.7%), ranging from 14.4 to 93.2 token/s.
  • No missing-data limitation applies: all eight models recorded 12 of 12 samples, valid_point_count is 96, and coverage is 100.0%, so every expected observation is present.
4-hour window · 88 points · 91.7% coverage
  • glm-5.3-flash is the strongest model at 133.32 token/s average throughput (peak 157.71 token/s), while nemotron-3-ultra is the weakest at 2.11 token/s average, never exceeding 3.58 token/s across the window.
  • The 08:00 UTC observation shows a synchronized dip: glm-5.2 fell to 12.31 token/s, minimax-m3 to 14.4 token/s, and glm-5.3-flash to 19.86 token/s. glm-5.3 also declined 28.7% overall, sliding from a 174.45 token/s peak at 06:40 to 61.65 token/s at 08:40.
  • Every model recorded 11 of 12 expected samples (91.7% coverage, 88 valid points), so each series is missing one observation, which limits confidence in interval-level comparisons.
4-hour window · 88 points · 91.7% coverage
  • Strongest average throughput was deepseek-v4-flash at 128.76 token/s (peaking at 157.71 token/s), narrowly ahead of glm-5.3 at 128.45 token/s; weakest was nemotron-3-ultra at 2.77 token/s, never exceeding 5.39 token/s.
  • The most operationally significant movement is nemotron-3-ultra's decline of 44.4 percent, falling to 1.06 token/s at 07:00, while glm-5.2 recovered 66.0 percent from a 10.4 token/s low at 04:40 to 111.52 token/s at 07:00; glm-5.3 also swung between 67.85 and 191.28 token/s.
  • Every model recorded 11 of 12 expected samples (91.7 percent coverage, 88 valid points), so one interval per model is missing and averages may not reflect the full four-hour window.
4-hour window · 88 points · 91.7% coverage
  • glm-5.3 is the strongest model with 121.63 token/s average throughput and a peak of 191.28 token/s; nemotron-3-ultra is the weakest at 3.83 token/s average, never exceeding 7.88 token/s.
  • glm-5.2 shows the sharpest operational swing, falling from 67.05 token/s at 04:20 to 10.4 token/s at 04:40 and 13.56 token/s at 05:00, then recovering to 107.7 token/s at 06:00 (52.5% trend); glm-5.3 is also volatile with a 40.2% coefficient of variation.
  • Each model has 11 of 12 expected samples (91.7% coverage), giving 88 valid points overall, so one 20-minute observation per model is missing across the 03:03–07:03 UTC window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model by average throughput at 123.69 token/s, peaking at 163.24 token/s, while nemotron-3-ultra is the weakest at 4.05 token/s average, never exceeding 7.88 token/s.
  • glm-5.2 shows the most operationally significant volatility (CV 49.8%), collapsing from 67.05 token/s at 04:20 to 10.4 token/s at 04:40 and 13.56 token/s at 05:00 before recovering to 107.7 token/s by 06:00; gemma4:31b also dipped to 4.06 token/s at 03:20.
  • No missing-data limitation applies: all eight models report 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model by average throughput at 115.56 token/s (peaking at 163.24 token/s), while nemotron-3-ultra is the weakest at 4.0 token/s average, never exceeding 7.88 token/s.
  • The most significant volatility is glm-5.2, whose coefficient of variation is 45.7%; it fell from 67.05 token/s at 04:20 to 10.4 token/s at 04:40 and ended at 13.56 token/s. glm-5.3-flash also swung from 17.95 token/s at 04:20 to 166.97 token/s at 05:00.
  • No missing-data limitation exists: all 8 models delivered 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model by average throughput at 108.79 token/s, narrowly ahead of deepseek-v4-flash at 112.31 token/s... correction: deepseek-v4-flash leads at 112.31 token/s, with gemma4:31b second at 108.79 token/s; nemotron-3-ultra is weakest at 3.94 token/s.
  • The most operationally significant volatility is gemma4:31b, which swings from a 4.06 token/s low at 03:20 to 153.82 token/s at peak, with a 38.7% coefficient of variation and a -24.7% trend over the window; glm-5.3 shows similar swings between 64.84 and 166.93 token/s.
  • No missing-data limitation exists: all 8 models report 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash posted the strongest average throughput at 117.11 token/s, just ahead of gemma4:31b at 116.02 token/s, while nemotron-3-ultra was weakest at 7.01 token/s, far below the next-lowest minimax-m3 at 39.89 token/s.
  • The most significant movement is nemotron-3-ultra's 70.3% decline, dropping from a 26.8 token/s peak at 23:40 to 4.05 token/s at 03:00 with 105.6% coefficient of variation; glm-5.3 was also volatile, swinging between 64.41 and 196.0 token/s with 40.9% CV.
  • No missing-data limitation exists: all eight models report 12 of 12 expected samples, 96 valid points, and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput was gemma4:31b at 131.55 token/s (peak 174.53 token/s); weakest was nemotron-3-ultra at 10.65 token/s, dipping to 1.48 token/s at 01:40.
  • Every model trended downward over the window, with nemotron-3-ultra collapsing 81.8% to 3.75 token/s by 02:00. Volatility was highest for glm-5.3 (CV 40.7%, range 64.41–196.0 token/s) and deepseek-v4-flash (CV 39.9%, range 43.3–188.6 token/s).
  • No missing-data limitation applies: all 96 expected points are valid, coverage is 100.0%, and each of the eight models has 12 of 12 samples.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 delivered the highest average throughput at 136.32 token/s, peaking at 235.48 token/s, while nemotron-3-ultra was weakest at 12.71 token/s average and ended at just 2.17 token/s.
  • glm-5.3 showed the sharpest decline, trending down 35.8% and swinging between 235.48 and 64.41 token/s; nemotron-3-ultra was the most volatile, with a 91.8% coefficient of variation across its 12 samples.
  • No missing-data limitation exists: all eight models have 12 of 12 expected samples, 96 valid points, and 100.0% coverage, so the four-hour window is fully populated.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model by average throughput at 144.41 token/s, peaking at 235.48 token/s, while nemotron-3-ultra is the weakest at 12.62 token/s average, never exceeding 43.56 token/s.
  • glm-5.3 shows the most operationally significant volatility, swinging between 64.41 and 235.48 token/s (stddev 53.68, CV 37.2%) with a 28.0% downward trend; nemotron-3-ultra is less stable still (CV 92.8%), spiking to 43.56 token/s at 23:00 after starting near 3.3 token/s.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0% across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model at 148.24 token/s average output throughput, peaking at 235.48 token/s at 22:00 UTC; nemotron-3-ultra is the weakest at 8.98 token/s average, ranging from 1.76 to 43.56 token/s.
  • The most operationally significant volatility is gemma4:31b's drop to 44.74 token/s at 21:00 UTC from roughly 130 token/s earlier, followed by recovery to 174.12 token/s at 22:20 UTC; glm-5.3 also swung between 72.31 and 235.48 token/s within the window.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, and the dataset shows 96 valid points with 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model with an average throughput of 140.64 token/s, peaking at 235.48 token/s at 22:00 UTC; nemotron-3-ultra is the weakest at an average of 5.46 token/s, never exceeding 17.52 token/s.
  • The most significant volatility is nemotron-3-ultra's coefficient of variation of 74.0% alongside a 96.6% upward trend, while glm-5.3 shows the largest absolute swing, ranging from 67.36 to 235.48 token/s with a 48.1% trend over the window.
  • No missing-data limitation applies: all eight models report 12 of 12 expected samples, and the dataset shows 96 valid points with 100.0% coverage across the four-hour period.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model at 123.43 token/s average throughput, ahead of glm-5.3-flash at 110.05 token/s; nemotron-3-ultra is the weakest at 3.37 token/s average, never exceeding 4.34 token/s.
  • glm-5.3 is also the most volatile, with a 45.1% coefficient of variation and swings between 67.36 and 244.23 token/s within the window; deepseek-v4-flash ranged from 29.48 to 153.69 token/s.
  • No missing-data limitation applies: all eight models report 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour period.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model by average throughput at 116.77 token/s (p95 214.86 token/s), while nemotron-3-ultra is the weakest at 2.94 token/s average, never exceeding 4.34 token/s.
  • glm-5.3 also shows the highest volatility, with a 47.3% coefficient of variation and swings between 67.36 and 244.23 token/s; deepseek-v4-pro declined 14.3% over the window, dipping to 49.14 token/s at 19:00 UTC.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, 96 valid points, and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model at 120.77 token/s average throughput (peak 165.34 token/s), while nemotron-3-ultra is the weakest at 2.56 token/s average, roughly 47 times slower.
  • glm-5.3 shows the highest volatility (cv 50.7%), swinging from 57.98 to 244.23 token/s, including a 244.23 token/s spike at 17:20 UTC; deepseek-v4-flash also dipped to 29.48 token/s at 17:20 before recovering to 143.08 token/s by 19:00.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0% across the four-hour window, so averages rest on complete data.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model at 112.13 token/s average throughput, slightly ahead of glm-5.3-flash at 103.81 token/s; nemotron-3-ultra is the weakest at 2.15 token/s average, never exceeding 3.23 token/s.
  • glm-5.3 shows the most operationally significant volatility, with a 60.7% coefficient of variation, a low of 19.23 token/s at 15:00, and a spike to 244.23 token/s at 17:20; deepseek-v4-pro also climbed 42.2% overall, from 37.97 token/s at 14:20 to 86.82 token/s at 18:00.
  • No missing-data limitation exists in this window: all eight models have 12 of 12 samples, 96 valid points, and 100.0% coverage.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model at 118.78 token/s average throughput (82.26–165.34 token/s range), while nemotron-3-ultra is the weakest at 2.04 token/s average, never exceeding 3.1 token/s.
  • The most significant movement is deepseek-v4-pro's upward trend of 94.3%, climbing from a 31.02 token/s low near 14:40 to a 122.52 token/s peak at 15:40 and holding near 98.59 token/s at 17:00; glm-5.3 also swung between 19.23 and 156.41 token/s.
  • No missing-data limitation exists: all eight models have 12 of 12 samples, 96 valid points, and 100.0% coverage, though the window covers only four hours.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput is glm-5.3-flash at 104.88 token/s (peaking at 134.55 token/s), narrowly ahead of deepseek-v4-flash at 104.75 token/s; weakest is nemotron-3-ultra at 2.43 token/s, with a maximum of only 3.38 token/s.
  • gemma4:31b shows the greatest volatility, with a coefficient of variation of 57.7% and a swing from 157.85 token/s at 12:20 down to 13.7 token/s at 13:00; glm-5.3 declined 29.1% overall, including a low of 19.23 token/s at 15:00.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0%.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput: deepseek-v4-flash at 104.30 token/s, narrowly ahead of glm-5.3-flash at 102.90 token/s; weakest is nemotron-3-ultra at 3.32 token/s, roughly 31x slower than the leader.
  • Most significant volatility: glm-5.3 fell from 173.56 token/s at 11:20 to 19.23 token/s at 15:00 (trend -44.0%), while gemma4:31b swung between 157.85 and 13.7 token/s (cv 58.3%).
  • No missing-data limitation: all eight models have 12 of 12 samples, 96 valid points, and 100.0% coverage over the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model at 113.46 token/s average, narrowly ahead of deepseek-v4-flash (113.44 token/s) and glm-5.3-flash (113.37 token/s); nemotron-3-ultra is the weakest at 4.29 token/s average, never exceeding 6.89 token/s.
  • Volatility dominates the window: gemma4:31b ranged from 13.7 to 176.08 token/s (CV 49.8%), and deepseek-v4-flash dropped to 34.25 token/s at 12:40 before rebounding to 130.99 token/s at 13:00; glm-5.3 posted the steepest rise, up 35.2%.
  • No missing-data limitation applies: all 96 expected points are valid (100% coverage), and each of the eight models reports 12 of 12 samples across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput belongs to deepseek-v4-flash at 118.17 token/s, with a peak of 194.85 token/s at 11:00; weakest is nemotron-3-ultra at 4.52 token/s, roughly 26 times lower.
  • gemma4:31b shows the sharpest deterioration, trending -33.7% with a 40.7% coefficient of variation and swings between 176.08 and 46.36 token/s; glm-5.2 trended +119.7% overall but dropped to 14.4 token/s in the final observation.
  • No missing-data limitation applies: all eight models report 12 of 12 expected samples, 96 valid points, and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput was deepseek-v4-flash at 117.83 token/s, peaking at 194.85 token/s; weakest was nemotron-3-ultra at 4.95 token/s, never exceeding 9.23 token/s.
  • The most operationally significant movement is gemma4:31b's late-window collapse from 176.08 token/s at 09:20 to 46.36 token/s at 11:00, a -21.0% trend; glm-5.2 also showed high relative volatility at 63.4% cv, dipping to 3.06 token/s at 09:00.
  • No missing-data limitation applies: all eight models recorded 12 of 12 expected samples, valid_point_count is 96, and coverage is 100.0% across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 120.35 token/s average throughput, ahead of glm-5.3 at 112.66 token/s and glm-5.3-flash at 112.43 token/s; nemotron-3-ultra is the weakest at 5.77 token/s average, never exceeding 14.11 token/s.
  • glm-5.2 shows the sharpest decline, trending down 56.7% to a low of 3.06 token/s at 09:00 UTC, while minimax-m3 is the most volatile with a 61.8% coefficient of variation, ranging from 21.12 to 122.93 token/s.
  • No missing-data limitation applies: all eight models delivered 12 of 12 expected samples, giving 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash leads on average throughput at 117.44 token/s, ahead of gemma4:31b (113.72 token/s) and glm-5.3-flash (112.63 token/s); nemotron-3-ultra is weakest at 14.20 token/s average, with a peak of only 41.74 token/s.
  • The most operationally significant change is nemotron-3-ultra's collapse from 41.74 token/s at 06:00 to 4.21 token/s at 09:00, a 75.4% decline, with the highest volatility at 99.6% coefficient of variation; glm-5.2 also ended at 3.06 token/s despite a 44.48 token/s average.
  • No missing-data limitation applies: all eight models report 12 of 12 samples, 96 valid points, and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model with an average throughput of 127.05 token/s (peak 246.11 token/s), while nemotron-3-ultra is the weakest at 22.82 token/s average (minimum 3.3 token/s).
  • The most operationally significant change is nemotron-3-ultra's collapse after 06:20 UTC, dropping from a 51.95 token/s peak to readings between 3.3 and 14.11 token/s and ending at 4.09 token/s, an 81.0% decline; it also shows the highest volatility at 73.7% CV.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, 96 valid points, and 100.0% coverage, though the four-hour window limits longer-term assessment.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 posted the highest average throughput at 131.0 token/s, while nemotron-3-ultra was weakest at 34.34 token/s; glm-5.3-flash averaged 115.95 token/s and gemma4:31b 112.52 token/s.
  • The most operationally significant movement is nemotron-3-ultra's sustained decline, down 50.3% over the window to a low of 3.3 token/s at 07:00; glm-5.3 also swung widely between 62.54 and 246.11 token/s with a 46.4% coefficient of variation.
  • No missing-data limitation applies: all eight models recorded 12 of 12 expected samples, yielding 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 posted the highest average throughput at 132.17 token/s, while nemotron-3-ultra was the weakest at 35.84 token/s; deepseek-v4-flash (114.04 token/s) and gemma4:31b (123.01 token/s) also ranked near the top.
  • The most significant movement was deepseek-v4-flash trending up 47.8%, climbing from a 49.83 token/s low at 03:40 to a 189.78 token/s peak at 05:40; glm-5.3 was highly volatile, ranging 62.54 to 246.11 token/s with a 45.1% coefficient of variation.
  • No missing-data limitation applies: all 8 models delivered 12 of 12 expected samples, totaling 96 valid points at 100.0% coverage for the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 142.79 token/s average throughput, peaking at 184.71 token/s; nemotron-3-ultra is weakest at 28.04 token/s average, with a low of 4.38 token/s.
  • nemotron-3-ultra shows the most extreme volatility (89.3% coefficient of variation), jumping from 4.38 to 96.70 token/s at 03:20, while glm-5.3 swings between 62.82 and 246.11 token/s (45.2% CV) and trends up 55.7%.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, 96 valid points, and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 155.66 token/s average throughput, while nemotron-3-ultra is the weakest at 20.62 token/s average, peaking at only 96.7 token/s.
  • glm-5.3 shows the most operationally significant volatility, swinging between 60.82 and 195.89 token/s with a 48.8% coefficient of variation; minimax-m3 also declined 41.9% overall, from 74.97 to 40.57 token/s.
  • No missing-data limitation exists: all eight models have 12 of 12 samples, and the dataset shows 96 valid points with 100.0% coverage.
4-hour window · 80 points · 83.3% coverage
  • Highest average throughput was gemma4:31b at 156.45 token/s; lowest was nemotron-3-ultra at 10.76 token/s, with deepseek-v4-flash second-highest at 126.73 token/s.
  • The most operationally significant movement is minimax-m3's steady decline, trending -27.2% from a 74.97 token/s peak at 00:20 to 30.67 token/s at 03:00; glm-5.3 also swung widely between 60.82 and 195.89 token/s (cv 48.5%).
  • Each model captured only 10 of 12 expected samples (83.3% coverage), with 80 valid points overall, so two twenty-minute observations per model are missing and averages may not fully represent the four-hour window.
4-hour window · 56 points · 58.3% coverage
  • gemma4:31b is the strongest model at 155.75 token/s average throughput, peaking at 189.38 token/s, while nemotron-3-ultra is the weakest at 9.56 token/s average and just 4.81 token/s in its latest sample.
  • glm-5.3 shows the most operationally significant volatility: it fell from 195.89 token/s at 00:20 to roughly 60-69 token/s across 00:40-01:40 before rebounding to 171.36 token/s, a 50.9% coefficient of variation. deepseek-v4-flash similarly swung between 209.95 and 53.9 token/s.
  • Every model recorded only 7 of 12 expected samples, yielding 58.3% coverage and 56 valid points overall, so the 22:03-23:40 UTC window is absent and averages may not represent the full four hours.