← Performance dashboard

Hourly performance insights

Summaries of rolling four-hour performance data

1196 retained summaries
4-hour window · 96 points · 100.0% coverage
  • glm-5.3-flash is the strongest model at 93.51 token/s average throughput, ahead of minimax-m3 (87.80 token/s) and gemma4:31b (87.16 token/s); nemotron-3-ultra is the weakest at 1.96 token/s average, never exceeding 3.72 token/s.
  • deepseek-v4-pro shows the most operationally significant volatility, with a 57.9% coefficient of variation, a peak of 92.28 token/s at 14:20 UTC, and a collapse to 8.89 token/s at 18:00 UTC; glm-5.3-flash also declined 39.0% across the window.
  • No missing-data limitation applies: all eight models recorded 12 of 12 samples, 96 valid points total, and 100.0% coverage over the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3-flash is the strongest model at 108.11 token/s average throughput (peak 165.75 token/s), while nemotron-3-ultra is the weakest at 1.53 token/s average, never exceeding 2.56 token/s across the window.
  • deepseek-v4-pro shows the most operationally significant volatility, with a 55.0% coefficient of variation, swings between 11.68 and 92.28 token/s, and a 32.3% downward trend ending at 14.95 token/s; deepseek-v4-flash conversely rose 67.3% to a 83.64 token/s peak.
  • No missing-data limitation exists: all 8 models have 12 of 12 expected samples, 96 valid points, and 100.0% coverage over the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3-flash is the strongest model at 111.35 token/s average throughput (peak 165.75 token/s), while nemotron-3-ultra is the weakest at 1.6 token/s average, never exceeding 2.56 token/s.
  • deepseek-v4-flash shows the sharpest decline, trending down 36.3% to a 37.42 token/s average with a low of 4.94 token/s at 14:40 UTC; deepseek-v4-pro is the most volatile at 52.0% coefficient of variation, swinging between 11.68 and 92.28 token/s.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, 96 valid points, and 100.0% coverage, so results reflect the full four-hour window.
4-hour window · 88 points · 91.7% coverage
  • glm-5.3-flash is the strongest model at 100.49 token/s average throughput, while nemotron-3-ultra is the weakest at 2.34 token/s, roughly 43 times slower over the window.
  • The most significant volatility is deepseek-v4-pro, which swung between 7.72 and 92.28 token/s with a 53.1% coefficient of variation, including repeated collapses below 15 token/s; deepseek-v4-flash also fell 59.1% to 4.94 token/s at 14:40.
  • Each of the eight models recorded 11 of 12 expected samples, leaving 88 valid points and 91.7% coverage, so one 20-minute observation per model is missing and short-lived dips may be understated.
4-hour window · 96 points · 100.0% coverage
  • Highest average throughput was minimax-m3 at 90.15 token/s, ahead of glm-5.3 (86.94 token/s) and gemma4:31b (82.93 token/s); lowest was nemotron-3-ultra at 3.09 token/s, far below glm-5.2 (54.02 token/s) and deepseek-v4-flash (54.55 token/s).
  • deepseek-v4-pro showed the widest volatility (CV 53.7%), ranging from 7.72 to 96.9 token/s with sharp dips at 10:00 and 12:00 UTC; gemma4:31b had the steepest decline at -25.9% trend, falling from 140.75 token/s at 09:40 to 32.76 token/s at 13:00.
  • No missing-data limitation applies: every model has 12 of 12 expected samples, 96 valid points overall, and 100.0% coverage.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 90.84 token/s average throughput, while nemotron-3-ultra is the weakest at 3.58 token/s average, roughly 25 times lower.
  • deepseek-v4-flash shows the sharpest trend, rising 125.8% from 36.52 token/s at 08:20 to 69.41 token/s at 12:00, while glm-5.3 fell 29.3%; deepseek-v4-pro is the most volatile, with a 56.4% coefficient of variation and swings between 7.72 and 96.9 token/s.
  • No missing-data limitation applies: all 8 models delivered 12 of 12 expected samples, and the dataset records 96 valid points with 100.0% coverage.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b delivered the highest average throughput at 100.48 token/s (peak 149.3 token/s), while nemotron-3-ultra was weakest at 3.51 token/s average, never exceeding 5.66 token/s.
  • deepseek-v4-pro showed the most operationally significant volatility, ranging from 9.59 to 96.9 token/s with a 57.2% coefficient of variation and a +41.7% trend, including drops to 9.59 and 9.92 token/s at 09:00 and 10:00.
  • No missing-data limitation applies: all eight models reported 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window ending 11:03 UTC.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model by average throughput at 99.93 token/s, edging gemma4:31b at 95.48 token/s; nemotron-3-ultra is the weakest at 3.28 token/s, roughly 30 times slower than the leader.
  • deepseek-v4-pro shows the most operationally significant volatility, swinging between 3.58 and 92.24 token/s with a coefficient of variation of 83.5%, including a drop to 9.92 token/s at 10:00; deepseek-v4-flash also declined 43.4% over the window.
  • No missing-data limitation exists: all 8 models delivered 12 of 12 expected samples, and the dataset records 96 valid points with 100.0% coverage, so the four-hour window is fully populated.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput was gemma4:31b at 106.01 token/s (p95 146.1 token/s), while nemotron-3-ultra was weakest at 3.81 token/s, never exceeding 7.08 token/s in any observation.
  • The most operationally significant volatility came from deepseek-v4-pro, which swung between 3.58 and 104.27 token/s with an 86.3% coefficient of variation; deepseek-v4-flash also fell 42.2% overall, ending at 27.03 token/s versus an early 126.98 token/s.
  • No missing-data limitation applies: all eight models recorded 12 of 12 expected samples, 96 valid points total, with 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b delivered the highest average throughput at 117.61 token/s (peaking at 161.83 token/s), while nemotron-3-ultra was weakest at 3.89 token/s average, never exceeding 7.08 token/s across all 12 observations.
  • deepseek-v4-pro showed the most extreme volatility, with a coefficient of variation of 84.9% and swings between 3.58 and 104.27 token/s, including four consecutive readings below 13 token/s from 06:00 to 07:00; deepseek-v4-flash also dropped sharply to 12.47 token/s at 06:40.
  • No missing-data limitation applies: all eight models recorded 12 of 12 expected samples, with 96 valid points and 100.0% coverage, so the four-hour window is fully represented.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b posted the strongest average throughput at 124.01 token/s, peaking at 168.45 token/s, while nemotron-3-ultra was weakest at 6.48 token/s average and never exceeded 33.18 token/s in any observation.
  • deepseek-v4-pro showed the most operationally significant volatility, swinging between 104.27 token/s at 05:40 and readings near 10 token/s, ending at 3.58 token/s with a coefficient of variation of 82.3%; gemma4:31b also trended down 25.0% across the window.
  • No missing-data limitation applies: all eight models recorded 12 of 12 expected samples, totaling 96 valid points with 100.0% coverage over the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 139.44 token/s average throughput, while nemotron-3-ultra is the weakest at 7.59 token/s average, roughly 18 times lower than the leader.
  • The most operationally significant volatility is deepseek-v4-pro, which repeatedly collapses from peaks near 104.27 token/s down to 9.98, 10.96, and 9.36 token/s; nemotron-3-ultra also shows extreme instability with a 104.6% coefficient of variation.
  • No missing-data limitation exists in this window: all eight models have 12 of 12 expected samples, 96 valid points, and 100.0% coverage, so every twenty-minute observation from 02:20 to 06:00 UTC is present.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b delivered the strongest average throughput at 131.22 token/s, peaking at 169.78 token/s, while nemotron-3-ultra was the weakest at 7.87 token/s average, never exceeding 33.18 token/s across the window.
  • deepseek-v4-pro showed the most operationally significant degradation, falling from 91.17 token/s at 01:20 to 10.96 token/s at 04:00 and 9.98 token/s at 05:00; glm-5.3-flash was the most volatile, ranging 50.22 to 177.96 token/s with a 47.7% coefficient of variation.
  • No missing-data limitation applies: all 8 models recorded 12 of 12 expected samples, yielding 96 valid points and 100.0% coverage over the four-hour period.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 114.58 token/s average throughput, peaking at 169.78 token/s; nemotron-3-ultra is weakest at 9.55 token/s average, never exceeding 33.18 token/s.
  • deepseek-v4-flash shows the sharpest decline, falling from 112.34 token/s at 00:20 to 42.37 token/s at 04:00 (trend -34.7%), while glm-5.3-flash is the most volatile with a 46.4% coefficient of variation, swinging between 50.22 and 177.96 token/s.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0%, so the four-hour window is complete.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 112.0 token/s average throughput, well ahead of glm-5.3-flash (89.17 token/s) and glm-5.3 (89.15 token/s); nemotron-3-ultra is the weakest at 11.01 token/s average, roughly one-tenth of the leader.
  • The most operationally significant movement is nemotron-3-ultra's decline of 56.9 percent, falling from a 21.93 token/s peak to 7.11 token/s at 03:00, while glm-5.3 shows the highest volatility (53.0 percent coefficient of variation, 15.63 to 162.33 token/s range).
  • No missing-data limitation applies: all eight models have 12 of 12 samples, and overall coverage is 100.0 percent across 96 valid points, so the four-hour window is fully represented.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 107.15 token/s average throughput, while nemotron-3-ultra is the weakest at 11.39 token/s, roughly one-tenth of the leader.
  • glm-5.3 shows the most operationally significant volatility, with a 57.4% coefficient of variation and swings between 15.48 and 159.95 token/s; deepseek-v4-flash posted the steepest gain, trending up 44.9% to a 94.82 token/s average.
  • No missing-data limitation exists: all eight models delivered 12 of 12 expected samples, and the dataset records 96 valid points at 100.0% coverage for the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput is gemma4:31b at 120.5 token/s (peak 178.16 token/s); weakest is nemotron-3-ultra at 11.53 token/s, roughly one-tenth of gemma4:31b's average.
  • glm-5.3 shows the most operationally significant volatility, with a 57.2% coefficient of variation and swings between 15.48 and 159.95 token/s within the window; gemma4:31b and deepseek-v4-pro each declined 25.0% over the period.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, 96 valid points total, and 100.0% coverage, so the four-hour window is fully represented.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model by average throughput at 120.89 token/s, with a peak of 178.16 token/s; nemotron-3-ultra is the weakest at 9.24 token/s average, never exceeding 21.93 token/s across the window.
  • glm-5.3 shows the most operationally significant volatility, swinging between 15.48 and 159.95 token/s with a 55.7% coefficient of variation, including abrupt drops at 22:40 and 23:20; nemotron-3-ultra also climbed from 1.01 to 20.23 token/s, a 241.7% trend.
  • No missing-data limitation exists in this dataset: all eight models recorded 12 of 12 expected samples, yielding 96 valid points at 100.0% coverage, though the four-hour window limits longer-term assessment.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model with an average throughput of 127.01 token/s (peaking at 178.16 token/s), while nemotron-3-ultra is the weakest at 4.94 token/s average, roughly 26 times slower than the leader.
  • The most operationally significant volatility comes from nemotron-3-ultra, with a coefficient of variation of 70.4% and a swing from 1.01 to 13.87 token/s; glm-5.3 also showed a sharp single-interval drop to 15.48 token/s at 22:40 before recovering to 102.15 token/s.
  • No missing-data limitation exists in this window: all eight models have 12 of 12 expected samples, with 96 valid points and 100.0% coverage, though the four-hour span limits longer-term conclusions.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput was gemma4:31b at 110.43 token/s (peaking at 178.16 token/s); weakest was nemotron-3-ultra at 3.32 token/s, never exceeding 7.66 token/s across the window.
  • The most operationally significant volatility came from gemma4:31b, which swung between 32.83 and 178.16 token/s (cv 41.5%), while deepseek-v4-pro dipped to 19.98 token/s at 20:00 before recovering to 99.99 token/s at 21:40; glm-5.2 also declined 19.3% overall, falling to 45.08 token/s at 21:00.
  • No missing-data limitation applies: all eight models recorded 12 of 12 expected samples, 96 valid points total, with 100.0% coverage.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3-flash is the strongest model at 94.67 token/s average throughput (peak 172.12 token/s), while nemotron-3-ultra is the weakest at 2.2 token/s average, peaking at only 3.7 token/s.
  • The most significant volatility is gemma4:31b, which swung from a 32.83 token/s low at 18:40 to a 163.16 token/s high at 19:40 (stddev 39.47 token/s, +79.4% trend); deepseek-v4-pro also dropped sharply from 89.39 to 19.98 token/s at 20:00.
  • No missing-data limitation exists: all 8 models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0%.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3-flash is the strongest model at 96.04 token/s average throughput, ahead of gemma4:31b at 92.79 token/s; nemotron-3-ultra is weakest at 2.34 token/s average, peaking at only 3.7 token/s.
  • deepseek-v4-pro shows the most operationally significant volatility, ranging from 3.88 to 88.64 token/s; it held roughly 76-88 token/s from 17:40 to 19:40 before dropping to 19.98 token/s at 20:00, and glm-5.3 spiked to 229.16 token/s at 17:00.
  • No missing-data limitation applies: all eight models recorded 12 of 12 samples, valid_point_count is 96, and coverage is 100.0% across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model by average throughput at 93.92 token/s (peaking at 229.16 token/s), while nemotron-3-ultra is the weakest at 1.9 token/s average, never exceeding 3.63 token/s.
  • The most operationally significant volatility is deepseek-v4-pro, which swings from a low of 3.88 token/s to a high of 88.64 token/s (coefficient of variation 56.0%), including readings of 6.34 and 10.76 token/s mid-window before recovering to 82.85 token/s.
  • No missing-data limitation exists: all 8 models delivered 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model with an average throughput of 99.19 token/s, ahead of gemma4:31b at 81.6 token/s, while nemotron-3-ultra is the weakest at 1.28 token/s, far below deepseek-v4-flash at 35.57 token/s.
  • The most significant volatility is deepseek-v4-pro, which swung between 7.81 and 78.25 token/s with a 52.0% coefficient of variation; gemma4:31b shows the steepest decline, trending down 37.3% from a 135.8 token/s peak at 12:40 UTC.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0% across the four-hour window.
4-hour window · 88 points · 91.7% coverage
  • Strongest average throughput is gemma4:31b at 94.99 token/s (peak 135.8 token/s); weakest is nemotron-3-ultra at 1.75 token/s, never exceeding 2.88 token/s across the window.
  • minimax-m3 shows the widest swings, ranging from 33.81 to 166.13 token/s with 50.1% coefficient of variation; deepseek-v4-pro is similarly volatile at 54.8%, dipping to 8.07 token/s at 12:00 UTC.
  • Every model recorded 11 of 12 expected samples, giving 88 valid points and 91.7% coverage overall, so one 20-minute observation per model is missing and the reported averages may shift with complete data.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 106.52 token/s average throughput, while nemotron-3-ultra is the weakest at 2.09 token/s average, roughly 50x lower.
  • The most operationally significant volatility is deepseek-v4-flash, whose throughput fell 55.2% over the window, from a 132.37 token/s peak at 10:40 to 36.57 token/s at 14:00, with a 73.1% coefficient of variation; deepseek-v4-pro also swung between 8.07 and 103.07 token/s.
  • No missing-data limitation applies: all 8 models delivered 12 of 12 expected samples, giving 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 103.82 token/s average throughput, ahead of glm-5.3 at 95.11 token/s; nemotron-3-ultra is the weakest at 2.33 token/s average, never exceeding 3.42 token/s.
  • deepseek-v4-flash shows the most operationally significant volatility, with a 79.4% coefficient of variation and a -62.9% trend, falling from a 132.37 token/s peak at 10:40 to 19.38 token/s at 13:00; deepseek-v4-pro also declined -27.7% to 17.82 token/s.
  • No missing-data limitation applies: all eight models report 12 of 12 samples, totaling 96 valid points with 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • Strongest average output throughput was gemma4:31b at 104.64 token/s across 12 samples; weakest was nemotron-3-ultra at 2.55 token/s, roughly 41 times lower.
  • The most operationally significant volatility came from deepseek-v4-pro, with a coefficient of variation of 70.6% and swings between 6.45 and 103.07 token/s; minimax-v4-pro aside, minimax-m3 showed the steepest trend at +105.1%, ending at 166.13 token/s.
  • No missing-data limitation applies: all eight models recorded 12 of 12 expected samples, giving 96 valid points and 100.0% coverage over the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model by average throughput at 105.36 token/s (p95 151.62 token/s), while nemotron-3-ultra is the weakest at 2.41 token/s, never exceeding 3.42 token/s.
  • glm-5.3-flash shows the sharpest trend, +100.6% across the window, rising from a 7.69 token/s low at 08:40 to a 157.03 token/s peak at 10:40; minimax-m3 is the most volatile (CV 58.9%, swings between 12.51 and 141.85 token/s), and deepseek-v4-pro also swings from 6.45 to 103.07 token/s.
  • No missing-data limitation exists: all eight models report 12 of 12 samples each, 96 valid points overall, and 100.0% coverage for the four-hour window.
4-hour window · 88 points · 91.7% coverage
  • gemma4:31b is the strongest model at 97.88 token/s average throughput, with glm-5.3 close behind at 96.93 token/s; nemotron-3-ultra is the weakest at 2.36 token/s average, never exceeding 3.23 token/s.
  • deepseek-v4-pro shows the most operationally significant volatility, ranging from 6.45 to 93.39 token/s with a 66.8% coefficient of variation and a -36.0% trend; minimax-m3 is similarly erratic at 66.6% CV, ranging 18.96 to 141.85 token/s.
  • Each model is missing one of twelve expected samples (11 collected, 91.7% coverage), leaving 88 valid points overall, so the absent observations could affect the reported averages and trends.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model at 105.52 token/s average throughput, peaking at 167.38 token/s; nemotron-3-ultra is the weakest at 3.14 token/s average, never exceeding 5.87 token/s.
  • minimax-m3 shows the sharpest volatility, with a coefficient of variation of 70.9%: it jumped from 29.72 token/s at 07:00 to 136.61 token/s at 07:20 and 141.85 token/s at 07:40, then dropped to 18.96 token/s at 08:20. deepseek-v4-pro also ended the window at its minimum of 6.45 token/s.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, 96 valid points, and 100% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model at 108.62 token/s average throughput, ahead of glm-5.3-flash at 95.61 token/s; nemotron-3-ultra is the weakest at 3.95 token/s average, roughly 27x slower than glm-5.3.
  • Volatility is the dominant operational signal: gemma4:31b swung between 181.21 and 25.1 token/s (cv 51.8%), and most models declined late in the window, with deepseek-v4-pro falling to 11.74 token/s at 07:00 (trend -37.4%) and glm-5.2 to 17.82 token/s.
  • No missing-data limitation applies: all eight models delivered 12 of 12 expected samples, giving 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput was glm-5.3 at 125.36 token/s, ahead of glm-5.3-flash at 99.18 token/s; weakest was nemotron-3-ultra at 4.5 token/s, far below the next-lowest, minimax-m3 at 46.35 token/s.
  • The most operationally significant volatility was minimax-m3, which spiked to 158.25 token/s at 05:00 against a 46.35 token/s average, with a coefficient of variation of 73.9%; gemma4:31b also swung between 37.97 and 181.21 token/s.
  • No missing-data limitation applies: all eight models report 12 of 12 expected samples, 96 valid points total, and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model at 124.37 token/s average throughput, peaking at 177.82 token/s; nemotron-3-ultra is the weakest at 3.78 token/s average, never exceeding 5.59 token/s.
  • minimax-m3 shows the most operationally significant volatility, with a 71.9% coefficient of variation and a -34.5% trend, falling from 106.72 token/s at 01:20 to 25.28 token/s at 04:40 before spiking to 158.25 token/s at 05:00.
  • No missing-data limitation applies: all eight models have 12 of 12 expected samples, and the dataset reports 96 valid points with 100.0% coverage, so the findings reflect complete observations across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model at 112.58 token/s average throughput, edging gemma4:31b at 108.94 token/s; nemotron-3-ultra is the weakest at 3.38 token/s average, roughly 33x slower than the leader.
  • minimax-m3 shows the sharpest volatility: it peaked at 156.65 token/s at 01:40 then fell to 29.21 token/s at 03:00, a -63.2% trend over the window, with a 65.3% coefficient of variation, the highest in the dataset.
  • No missing-data limitation applies: all 8 models delivered 12 of 12 expected samples, giving 96 valid points and 100.0% coverage for the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model at 117.65 token/s average throughput, edging gemma4:31b at 112.85 token/s, while nemotron-3-ultra is the weakest at 4.01 token/s average, roughly 29 times slower.
  • glm-5.3 shows the most significant upward trend, climbing 61.5% to a window-high 177.82 token/s at 03:00 UTC; minimax-m3 is the most volatile, with a 62.3% coefficient of variation, swinging from 156.65 token/s at 01:40 to 44.76 token/s at 02:20.
  • No missing-data limitation applies: all 8 models delivered 12 of 12 expected samples, yielding 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model at 126.76 token/s average (peak 166.44 token/s), while nemotron-3-ultra is the weakest at 4.5 token/s average, never exceeding 12.84 token/s.
  • The most operationally significant movement is minimax-m3's 86.5% upward trend, climbing from a 15.78 token/s low at 00:00 to 156.65 token/s at 01:40; nemotron-3-ultra also shows the highest volatility at 72.6% coefficient of variation.
  • No missing-data limitation applies: all eight models delivered 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model at 139.74 token/s average throughput (91.11 to 166.44 token/s range), while nemotron-3-ultra is the weakest at 5.27 token/s average, peaking at only 12.84 token/s.
  • The most significant trend is glm-5.3-flash improving 34.0% over the window, sustaining 162.09 to 164.66 token/s from 00:00 to 00:40 before falling to 78.42 token/s at 01:00; glm-5.3 shows the highest volatility at 44.9% coefficient of variation, swinging between 52.93 and 174.58 token/s.
  • No missing-data limitation exists: all 8 models have 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour period.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model at 145.78 token/s average throughput, with gemma4:31b second at 125.69 token/s; nemotron-3-ultra is the weakest at 5.3 token/s average, peaking at only 12.84 token/s.
  • The most significant volatility is glm-5.3, whose throughput swings between 52.93 and 174.58 token/s (39.7% coefficient of variation); the final 00:00 reading shows sharp drops for minimax-m3 to 15.78 token/s and glm-5.2 to 35.09 token/s.
  • No missing-data limitation exists: all 96 expected points are valid, and every model has 12 of 12 samples at 100.0% coverage.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model at 149.99 token/s average throughput (peaking at 183.45 token/s), while nemotron-3-ultra is the weakest at 4.3 token/s average, never exceeding 9.75 token/s.
  • glm-5.3 shows the most operationally significant volatility, swinging between 52.93 and 174.03 token/s with a 42.9% coefficient of variation, alternating high and low readings roughly every 20 minutes; nemotron-3-ultra also rose 164.2% over the window but from a very low base.
  • No missing-data limitation exists: all eight models delivered 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour period.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput was deepseek-v4-flash at 147.91 token/s (peak 183.45 token/s), while nemotron-3-ultra was weakest at 3.08 token/s, roughly 48 times slower.
  • glm-5.3 showed the most operationally significant volatility, with a 40.9% coefficient of variation and swings between 61.1 and 174.03 token/s, including a jump from 75.62 to 127.92 token/s around 21:00 UTC; nemotron-3-ultra also climbed from 1.53 to 8.09 token/s.
  • No missing-data limitation exists: all eight models reported 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model at 135.46 token/s average throughput, narrowly ahead of gemma4:31b at 131.52 token/s; nemotron-3-ultra is the weakest at 1.69 token/s average, peaking at only 2.0 token/s.
  • gemma4:31b shows the most volatility, with a coefficient of variation of 32.6% and swings from 183.23 token/s at 17:20 down to 65.56 token/s at 18:00, plus a -19.6% trend over the window; deepseek-v4-flash improved 14.2% to 151.66 token/s.
  • No missing-data limitation exists: all eight models have 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour period.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 139.6 token/s average throughput, with deepseek-v4-flash second at 132.08 token/s; nemotron-3-ultra is weakest at 1.56 token/s average, roughly 90x slower than gemma4:31b.
  • glm-5.3 shows the highest volatility (41.4% CV), swinging from 29.02 to 150.55 token/s within the window; gemma4:31b fell from 183.23 token/s at 17:20 to 65.56 token/s at 18:00, a sharp drop contributing to its -17.1% trend.
  • No missing-data limitation: all eight models have 12 of 12 samples, 96 valid points, and 100.0% coverage over the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput was gemma4:31b at 148.56 token/s, peaking at 183.23 token/s; weakest was nemotron-3-ultra at 1.70 token/s, never exceeding 2.80 token/s across the window.
  • glm-5.3-flash showed the steepest decline, trending -29.1% and falling from 173.76 token/s at 14:40 to 64.25 token/s at 16:40, while glm-5.3 was the most volatile, swinging between 29.02 and 150.55 token/s with a 40.2% coefficient of variation.
  • No missing-data limitation applies: all eight models delivered 12 of 12 expected samples, yielding 96 valid points and 100.0% coverage across the four-hour period.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b posted the strongest average throughput at 132.06 token/s (range 67.59 to 179.48 token/s), while nemotron-3-ultra was the weakest at 1.88 token/s average, never exceeding 2.8 token/s across the window.
  • glm-5.3 showed the most operationally significant volatility: a 46.1% coefficient of variation, a swing from 185.55 token/s at 13:20 down to 29.02 token/s at 15:40, and a -44.3% trend, though it rebounded to 150.55 token/s at 16:00.
  • No missing-data limitation applies: all eight models recorded 12 of 12 expected samples, giving 96 valid points and 100.0% coverage, so the only constraint is the four-hour window ending 2026-09-13T16:03:01+00:00.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model at 120.74 token/s average throughput, narrowly ahead of gemma4:31b at 119.05 token/s, while nemotron-3-ultra is the weakest at 2.32 token/s average, roughly 50 times slower.
  • The most significant volatility is deepseek-v4-flash collapsing to 14.51 token/s at 14:40 before recovering to 138.93 token/s at 15:00; glm-5.3-flash shows the strongest upward trend, rising 49.8% to peak at 173.76 token/s.
  • No missing-data limitation exists: all eight models have 12 of 12 samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 delivered the highest average throughput at 143.76 token/s (peak 185.55 token/s), while nemotron-3-ultra was weakest at 2.89 token/s average, peaking at only 4.97 token/s.
  • nemotron-3-ultra shows the most operationally significant trend, declining 51.3% across the window to 1.4 token/s at 14:00 UTC; deepseek-v4-pro was the most volatile (CV 37.0%), swinging between 126.49 and 26.36 token/s.
  • No missing-data limitation exists: all eight models have 12 of 12 samples, 96 valid points, and 100.0% coverage over the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput is deepseek-v4-flash at 143.39 token/s (peak 223.43 token/s); weakest is nemotron-3-ultra at 3.67 token/s, roughly 39 times lower.
  • Most operationally significant volatility is gemma4:31b, swinging between 51.55 and 155.75 token/s (cv 35.0%), while glm-5.2 shows the clearest trend, rising 27.3% from 45.21 token/s at 09:20 to a peak of 87.31 token/s at 11:20.
  • No missing-data limitation applies: all eight models report 12 of 12 expected samples, 96 valid points, and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model at 144.6 token/s average throughput (peak 223.43 token/s), while nemotron-3-ultra is the weakest at 3.82 token/s average, peaking at only 6.04 token/s.
  • glm-5.3-flash shows the sharpest decline, trending down 42.1% from a 172.69 token/s peak at 10:00 to 64.23 token/s at 11:00; gemma4:31b is the most volatile, with a 37.5% coefficient of variation and swings between 42.58 and 162.02 token/s.
  • No missing-data limitation applies: all eight models have 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model at 140.99 token/s average throughput, ahead of deepseek-v4-flash at 136.53 token/s and gemma4:31b at 113.93 token/s; nemotron-3-ultra is the weakest at 4.51 token/s average, never exceeding 6.99 token/s.
  • The most significant movement is glm-5.3's 28.6% upward trend, closing at 183.25 token/s, plus deepseek-v4-flash's spike to 223.43 token/s at 09:40; volatility is highest for glm-5.3 (33.9% CV) and gemma4:31b (33.3% CV), which swung between 42.58 and 162.02 token/s.
  • No missing-data limitation applies: all 8 models recorded 12 of 12 expected samples, totaling 96 valid points at 100.0% coverage over the four-hour window.