← Performance dashboard

Hourly performance insights

Summaries of rolling four-hour performance data

1210 retained summaries
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model at 183.38 token/s average throughput (peak 227.33 token/s), while nemotron-3-ultra is the weakest at 21.14 token/s average, roughly one-ninth of the leader.
  • glm-5.2 shows the most operationally significant volatility, with a coefficient of variation of 73.7 percent, swinging from 19.99 token/s at 21:40 to 191.19 token/s at 22:20; gemma4:31b also dipped to 50.18 token/s around 20:40.
  • No missing-data limitation applies: all eight models have 12 of 12 samples and 100.0 percent coverage, with 96 valid points overall, so the four-hour window is fully represented.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model at 180.13 token/s average throughput (range 129.09–227.33 token/s), while nemotron-3-ultra is the weakest at 21.11 token/s average, never exceeding 41.32 token/s.
  • glm-5.3 shows the steadiest gain, trending +19.0% to a 123.86 token/s average with the lowest volatility (cv 22.4%); glm-5.2 is the most volatile (cv 57.4%), swinging between 19.99 and 125.78 token/s, and minimax-m3 dipped to 6.95 token/s at 19:20.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, with 96 valid points and 100.0% coverage, so the four-hour window is fully represented.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model at 163.48 token/s average throughput, peaking at 227.33 token/s, while nemotron-3-ultra is the weakest at 20.78 token/s average and never exceeding 38.18 token/s.
  • The most operationally significant trend is glm-5.3's 78.7% rise, from 49.21 token/s at 18:00 to 143.23 token/s at 20:00; glm-5.2 shows the greatest volatility with a 48.3% coefficient of variation, swinging between 31.0 and 125.78 token/s.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, 96 valid points, and 100.0% coverage, so figures represent the complete four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model with 138.64 token/s average throughput, peaking at 227.33 token/s at 19:40; nemotron-3-ultra is the weakest at 22.01 token/s average, never exceeding 39.76 token/s across the window.
  • The most operationally significant movement is glm-5.3's late surge, climbing from 49.21 token/s at 18:00 to 143.23 token/s at 20:00 for a 63.4% trend gain, while glm-5.2 shows the highest relative volatility (cv 56.8%) and minimax-m3 dipped to 6.95 token/s at 19:20.
  • No missing-data limitation applies: all 96 expected points are valid (100.0% coverage), and each of the eight models reports 12 of 12 samples over the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model at 121.33 token/s average throughput, peaking at 199.80 token/s, while nemotron-3-ultra is the weakest at 23.46 token/s average and never exceeding 39.76 token/s.
  • glm-5.3 shows the most operationally significant volatility, swinging from 146.03 token/s at 15:20 down to 21.87 token/s at 17:00, with a coefficient of variation of 45.0% and a -23.6% trend; deepseek-v4.1-flash also oscillates sharply between 38.85 and 199.80 token/s.
  • No missing-data limitation exists: all eight models report 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput was deepseek-v4.1-flash at 111.66 token/s, ahead of gemma4:31b at 103.83 token/s; weakest was nemotron-3-ultra at 24.99 token/s, below minimax-m3 at 43.90 token/s.
  • glm-5.2 showed the most extreme volatility, ranging from 9.97 to 239.87 token/s with a 90.7% coefficient of variation, while glm-5.3 posted the steepest decline, falling 46.4% from 145.53 token/s at 14:20 to 49.21 token/s at 18:00.
  • No missing-data limitation applies: all eight models delivered 12 of 12 expected samples, yielding 96 valid points and 100.0% coverage, though the window covers only four hours.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 delivered the highest average throughput at 114.93 token/s (peak 154.09 token/s), while nemotron-3-ultra was the weakest at 23.15 token/s average, never exceeding 60.23 token/s.
  • glm-5.2 showed the most operationally significant volatility, swinging between 9.97 and 239.87 token/s with a coefficient of variation of 102.5%; deepseek-v4.1-flash also spiked to 199.8 token/s at 15:40 before falling to 38.85 token/s at 16:40.
  • No missing-data limitation applies: all eight models report 12 of 12 samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput belongs to deepseek-v4.1-flash at 126.29 token/s (p95 190.37 token/s); weakest is nemotron-3-ultra at 23.42 token/s, peaking at only 60.23 token/s.
  • glm-5.2 shows the most operationally significant volatility, with a 95.3% coefficient of variation, swings between 9.97 and 239.87 token/s, and a -35.1% trend; glm-5.3 rose 48.5% to average 108.11 token/s.
  • No missing-data limitation: all 8 models delivered 12 of 12 expected samples, 96 valid points, and 100.0% coverage across the 12:20–16:00 UTC observations.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model with an average throughput of 135.61 token/s (maximum 182.65 token/s), while nemotron-3-ultra is the weakest at 21.70 token/s average, peaking at only 60.23 token/s.
  • glm-5.2 shows the most operationally significant volatility, swinging between 9.97 and 239.87 token/s with a coefficient of variation of 85.4%, including a drop to 9.97 token/s at 14:20 followed by a spike to 239.87 token/s at 15:00; deepseek-v4.1-flash also declined 18.2% over the window.
  • No missing-data limitation exists in this dataset: all eight models recorded 12 of 12 expected samples, yielding 96 valid points and 100.0% coverage across the four-hour period.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model with an average throughput of 149.56 token/s (peaking at 212.23 token/s), while nemotron-3-ultra is the weakest at 17.34 token/s average, never exceeding 29.54 token/s.
  • glm-5.2 shows the most operationally significant volatility, with a coefficient of variation of 76.8% and swings from 27.27 token/s at 11:00 to 224.82 token/s at 13:40; deepseek-v4-pro also dropped to 11.72 token/s at 14:00.
  • No missing-data limitation exists: all eight models recorded 12 of 12 expected samples, yielding 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model by average throughput at 154.26 token/s, while nemotron-3-ultra is the weakest at 17.55 token/s, roughly nine times lower.
  • glm-5.2 shows the most operationally significant volatility, with a coefficient of variation of 81.9% and a trend of +106.1%; it swung from 25.95 token/s at 09:20 to a 193.32 token/s peak at 13:00, including a drop to 45.18 token/s at 12:20 after reaching 177.64 token/s at 12:00.
  • No missing-data limitation exists: all eight models have 12 of 12 expected samples, and the dataset reports 96 valid points with 100.0% coverage over the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model at 152.58 token/s average throughput, while nemotron-3-ultra is the weakest at 17.79 token/s average.
  • The most operationally significant volatility is glm-5.2, whose coefficient of variation is 64.8%; it swung from a 25.95 token/s low to 177.64 token/s at 12:00 UTC, and deepseek-v4-pro dropped sharply to 16.83 token/s in the final sample versus its 93.28 token/s average.
  • No missing-data limitation exists: all eight models delivered 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model at 146.78 token/s average throughput, peaking at 212.23 token/s; nemotron-3-ultra is the weakest at 17.57 token/s average, never exceeding 28.73 token/s.
  • glm-5.3-flash shows the most operationally significant volatility, with a 45.1% coefficient of variation and swings between 53.66 and 232.91 token/s, including a collapse from 232.91 token/s at 08:40 to 53.66 token/s at 09:00; glm-5.3 trended up 42.5%.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0%, so the four-hour window is fully represented.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput is deepseek-v4.1-flash at 142.14 token/s (p95 176.22 token/s); weakest is nemotron-3-ultra at 18.32 token/s, which never exceeded 28.73 token/s.
  • glm-5.3-flash is the most volatile model, with a 46.0% coefficient of variation and swings between 53.66 and 232.91 token/s; glm-5.3 shows the steepest decline, trending -28.7% from 170.48 token/s at 06:20 to 110.28 token/s at 10:00.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, 96 valid points, and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model by average throughput at 129.35 token/s, ahead of deepseek-v4.1-flash at 148.87 token/s... correction: deepseek-v4.1-flash leads at 148.87 token/s, with gemma4:31b second at 129.35 token/s; nemotron-3-ultra is weakest at 21.35 token/s.
  • glm-5.3 shows the sharpest decline, trending down 33.3% from a 170.48 token/s peak at 06:20 to 69.86 token/s at 09:00, while glm-5.3-flash rose 19.7% and spiked to 232.91 token/s at 08:40 before dropping to 53.66 token/s at 09:00.
  • No missing-data limitation exists: all eight models have 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput was deepseek-v4.1-flash at 128.78 token/s; weakest was nemotron-3-ultra at 24.57 token/s, roughly one-fifth of the leader's rate.
  • deepseek-v4.1-flash also showed the widest swings, ranging 14.9 to 186.13 token/s with 44.2% coefficient of variation, while glm-5.3 climbed 41.0% overall, peaking at 170.48 token/s before falling back to 112.79 token/s.
  • No missing-data limitation applies: all eight models delivered 12 of 12 expected samples, giving 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model at 135.35 token/s average throughput, peaking at 209.48 token/s; nemotron-3-ultra is the weakest at 23.57 token/s average, never exceeding 46.49 token/s.
  • The most operationally significant movement is deepseek-v4.1-flash's 52.5% upward trend, including a dip to 14.9 token/s at 05:00 before recovering to 165.38 token/s by 07:00; glm-5.2 shows the highest volatility at 65.6% coefficient of variation.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0%, so the four-hour window is fully represented.
4-hour window · 96 points · 100.0% coverage
  • Highest average throughput was deepseek-v4.1-flash at 122.44 token/s, edging glm-5.3-flash (122.37 token/s) and gemma4:31b (119.24 token/s); lowest was nemotron-3-ultra at 30.60 token/s, below minimax-m3's 37.69 token/s.
  • glm-5.2 showed the sharpest volatility, with a coefficient of variation of 82.8% and swings between a 202.79 token/s spike at 02:00 and a 21.44 token/s low at 04:40; deepseek-v4-pro trended up 53.1% while glm-5.3-flash fell 24.6%.
  • No missing-data limitation applies: all eight models delivered 12 of 12 samples, with 96 valid points and 100.0% coverage, though the window spans only four hours.
4-hour window · 96 points · 100.0% coverage
  • Highest average throughput was gemma4:31b at 133.1 token/s across 12 samples (range 81.25 to 170.9 token/s), while nemotron-3-ultra was lowest at 27.93 token/s, never exceeding 54.07 token/s.
  • glm-5.2 showed the most operationally significant volatility, with a 70.5% coefficient of variation, swings between 24.83 and 202.79 token/s, and a -31.2% trend; deepseek-v4.1-flash also swung between 24.14 and 226.05 token/s.
  • No missing-data limitation applies: all 8 models reported 12 of 12 expected samples, giving 96 valid points and 100.0% coverage over the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model at 150.73 token/s average throughput, ahead of gemma4:31b at 132.85 token/s and glm-5.3-flash at 131.83 token/s; nemotron-3-ultra is the weakest at 39.67 token/s average, below minimax-m3 at 48.47 token/s.
  • glm-5.2 shows the most volatility, with a 74.2% coefficient of variation, ranging from 24.52 to 202.79 token/s, including a spike to 202.79 token/s at 02:00 UTC; deepseek-v4.1-flash also swings sharply, dropping to 24.14 token/s at 01:20 after peaking at 226.05 token/s at 00:40.
  • No missing-data limitation exists: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0% across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model by average throughput at 140.56 token/s, narrowly ahead of gemma4:31b at 140.37 token/s; nemotron-3-ultra is the weakest at 44.39 token/s, below minimax-m3 at 55.41 token/s.
  • glm-5.2 shows the sharpest swing, rising 123.0% over the window to a peak of 202.79 token/s at 02:00 after dipping to 24.52 token/s at 00:00, with the highest volatility at 72.8% coefficient of variation; minimax-m3 fell 48.6% to 11.73 token/s.
  • No missing-data limitation applies: all eight models delivered 12 of 12 expected samples, and the dataset records 96 valid points at 100.0% coverage.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model by average throughput at 143.44 token/s, ahead of deepseek-v4.1-flash at 135.52 token/s; nemotron-3-ultra is the weakest at 46.7 token/s average, below minimax-m3 at 62.55 token/s.
  • The most operationally significant volatility is glm-5.3, which swings between 19.0 and 180.67 token/s (cv 50.8%), including a drop to 19.0 token/s at 00:00; deepseek-v4.1-flash shows the steepest upward trend at +52.8%, peaking at 226.05 token/s at 00:40.
  • No missing-data limitation applies: all 8 models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0%, so the four-hour window is fully represented.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b delivered the strongest average throughput at 129.49 token/s with a peak of 173.62 token/s, while nemotron-3-ultra was weakest at 47.04 token/s average, never exceeding 94.23 token/s.
  • glm-5.3 showed the most operationally significant volatility, swinging between 19.0 and 180.67 token/s with a 45.6% coefficient of variation; deepseek-v4-pro also fell sharply from 112.6 token/s at 23:20 to 12.59 token/s at 23:40.
  • No missing-data limitation applies: all eight models recorded 12 of 12 samples, giving 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model at 129.58 token/s average output throughput, ahead of gemma4:31b at 124.19 token/s; nemotron-3-ultra is the weakest at 37.83 token/s average, never exceeding 75.69 token/s.
  • deepseek-v4.1-flash shows the most operationally significant volatility, with a 60.6% coefficient of variation and swings from 19.01 to 185.99 token/s; glm-5.3 also declined 22.0% across the window, ending at 63.23 token/s versus its 178.26 token/s peak.
  • No missing-data limitation applies: all 8 models delivered 12 of 12 expected samples, totaling 96 valid points at 100.0% coverage, though the four-hour window restricts longer-term assessment.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model with an average throughput of 123.05 token/s (peak 178.26 token/s), while nemotron-3-ultra is the weakest at 29.47 token/s average, never exceeding 50.9 token/s.
  • deepseek-v4.1-flash shows the most operationally significant volatility, swinging between 19.01 and 194.62 token/s with a coefficient of variation of 53.6%; glm-5.2 is similarly erratic at 59.2%, dipping to 12.85 token/s at 21:00 UTC.
  • No missing-data limitation applies: all 8 models delivered 12 of 12 expected samples, yielding 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model at 130.48 token/s average throughput, narrowly ahead of deepseek-v4.1-flash at 128.91 token/s; nemotron-3-ultra is the weakest at 24.05 token/s average, never exceeding 41.31 token/s.
  • deepseek-v4.1-flash shows the most operationally significant volatility, swinging from 194.62 token/s at 18:40 to 19.01 token/s at 19:40, with a 47.7% coefficient of variation and a -47.3% trend; glm-5.2 is similarly unstable at 73.3% CV, ending at 12.85 token/s.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, 96 valid points, and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 delivered the highest average throughput at 131.79 token/s, peaking at 178.26 token/s at 20:00 UTC, while nemotron-3-ultra was weakest at 24.09 token/s average and never exceeded 41.31 token/s.
  • deepseek-v4.1-flash showed the most operationally significant volatility, swinging from 12.98 to 198.28 token/s with a 61.1% coefficient of variation, including a drop from 160.86 token/s at 19:00 to 38.51 token/s at 19:20 before recovering to 185.99 token/s.
  • No missing-data limitation applies: all 8 models recorded 12 of 12 expected samples, yielding 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model by average throughput at 136.12 token/s, while nemotron-3-ultra is the weakest at 21.19 token/s, roughly a 6.4x gap between the two.
  • The most operationally significant volatility is deepseek-v4.1-flash, which swung from 209.88 token/s at 15:40 down to 12.98 token/s at 16:40, with a coefficient of variation of 46.9%; deepseek-v4-pro also climbed from a 15.9 token/s low to 115.41 token/s.
  • No missing-data limitation exists: all eight models recorded 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model by average throughput at 145.0 token/s, with a peak of 232.5 token/s; nemotron-3-ultra is the weakest at 17.82 token/s average, never exceeding 30.41 token/s.
  • The most operationally significant volatility is deepseek-v4.1-flash, which fell from 232.5 token/s at 14:40 to 12.98 token/s at 16:40 before recovering to 198.28 token/s at 17:40, a coefficient of variation of 49.2%; glm-5.2 similarly swung between 101.73 and 15.97 token/s.
  • No missing-data limitation applies: all eight models recorded 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model at 137.49 token/s average throughput, while nemotron-3-ultra is the weakest at 20.03 token/s average, roughly one-seventh of the leader's pace.
  • The most operationally significant movement is deepseek-v4.1-flash's collapse: after peaking at 243.68 token/s at 13:40 UTC, it fell to 12.98 token/s at 16:40, a -39.7% trend with 60.4% coefficient of variation, indicating severe instability in the top performer.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0% across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model at 150.03 token/s average throughput; nemotron-3-ultra is the weakest at 19.73 token/s average, roughly 7.6 times lower.
  • deepseek-v4.1-flash shows the most operationally significant volatility, ranging from 20.22 to 243.68 token/s with a 51.5% coefficient of variation and a +63.2% trend, while minimax-m3 declined 42.4% to a 42.08 token/s latest reading.
  • No missing-data limitation applies: all eight models report 12 of 12 samples, 96 valid points, and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model at 135.07 token/s average throughput, while nemotron-3-ultra is the weakest at 24.51 token/s average.
  • The most operationally significant volatility is deepseek-v4.1-flash, whose throughput swings between 20.22 and 243.68 token/s (coefficient of variation 57.9%), while gemma4:31b shows the steepest decline, falling 40.5% from a 184.32 token/s peak to roughly 68-97 token/s late in the window.
  • No missing-data limitation exists: all eight models report 12 of 12 samples, 96 valid points, and 100.0% coverage; the only limitation is the short four-hour window with twelve observations per model.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput was gemma4:31b at 135.0 token/s, peaking at 184.32 token/s at 11:40, while nemotron-3-ultra was weakest at 28.03 token/s average and never exceeded 57.64 token/s.
  • The most operationally significant volatility came from deepseek-v4.1-flash, which ranged from 243.68 token/s at 13:40 down to 20.22 token/s at 14:00 with a 56.2% coefficient of variation; deepseek-v4-pro similarly swung between 11.76 and 114.54 token/s.
  • No missing-data limitation applies: all eight models delivered 12 of 12 expected samples, yielding 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b delivered the highest average throughput at 135.78 token/s, while nemotron-3-ultra was weakest at 27.62 token/s; deepseek-v4.1-flash and glm-5.3 followed closely at 132.95 and 129.78 token/s.
  • deepseek-v4.1-flash showed the sharpest swing, peaking at 210.67 token/s at 10:00 before falling to 37.08 token/s at 12:40, with a coefficient of variation of 44.3% and a trend of -41.1%; glm-5.3 also swung between 182.87 and 65.78 token/s.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, 96 valid points, and 100.0% coverage, though the window covers only four hours.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash posted the highest average throughput at 140.28 token/s, peaking at 210.67 token/s, while nemotron-3-ultra was weakest at 25.28 token/s average and never exceeded 45.36 token/s.
  • glm-5.3 showed the sharpest volatility, ranging from 50.69 to 182.87 token/s with a 47.2% coefficient of variation, including a fall from 181.17 token/s at 10:20 to 72.23 token/s at 10:40; deepseek-v4.1-flash also dropped from 142.01 to 37.34 token/s at 12:00.
  • No missing-data limitation applies: all 8 models recorded 12 of 12 expected samples, totaling 96 valid points with 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3-flash is the strongest model at 146.88 token/s average throughput, peaking at 221.2 token/s; nemotron-3-ultra is the weakest at 25.75 token/s average, never exceeding 79.12 token/s.
  • glm-5.3 shows the most operationally significant volatility, swinging between 49.75 and 182.87 token/s with a 46.8% coefficient of variation, while deepseek-v4-pro dropped to 5.64 token/s at 10:00 UTC before recovering to 114.54 token/s by 10:40.
  • No missing-data limitation exists: all eight models recorded 12 of 12 expected samples, yielding 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3-flash is the strongest model by average throughput at 147.85 token/s, ahead of deepseek-v4.1-flash (134.79 token/s) and gemma4:31b (133.44 token/s); nemotron-3-ultra is the weakest at 24.36 token/s average, with a maximum of only 79.12 token/s.
  • The most operationally significant volatility is deepseek-v4-pro's collapse to 5.64 token/s at 10:00 after holding 85.61 to 114.36 token/s from 09:00 to 09:40; deepseek-v4.1-flash similarly fell to 24.35 token/s at 08:20 before recovering to 210.67 token/s by 10:00.
  • No missing-data limitation exists: all eight models have 12 of 12 samples, and overall coverage is 100.0% with 96 valid points.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model by average throughput at 142.88 token/s, narrowly ahead of glm-5.3-flash at 137.72 token/s; nemotron-3-ultra is the weakest at 24.25 token/s average, with a maximum of only 79.12 token/s.
  • deepseek-v4.1-flash showed the sharpest deterioration, trending down 41.8% with a fall from 196.28 token/s at 06:00 to 24.35 token/s at 08:20, while glm-5.3-flash improved 57.7% from 5.74 token/s at 05:20 to 221.2 token/s at 09:00; nemotron-3-ultra was the most volatile at 81.5% CV.
  • No missing-data limitation applies: all 96 expected points are valid, coverage is 100.0%, and every model has 12 of 12 samples across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model by average throughput at 152.3 token/s (p95 175.61 token/s), while nemotron-3-ultra is the weakest at 28.44 token/s average, peaking at only 79.12 token/s.
  • The most operationally significant movement is glm-5.3-flash, which climbed from 5.15 token/s at 04:20 to 233.42 token/s at 06:20, a 294.8% trend with 77.4% coefficient of variation; deepseek-v4.1-flash also fell to 24.65 token/s at 08:00, its window minimum.
  • No missing-data limitation applies: all eight models report 12 of 12 expected samples, 96 valid points, and 100.0% coverage, so the four-hour window is complete.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model at 154.31 token/s average, narrowly ahead of gemma4:31b at 151.3 token/s; nemotron-3-ultra is the weakest at 34.64 token/s average, ending at just 3.86 token/s.
  • glm-5.3-flash shows the most operationally significant volatility: it ran near 5-12 token/s for the first seven intervals, then jumped to 233.42 token/s at 06:20, with a coefficient of variation of 120.2%. glm-5.3 also fell from 184.0 token/s at 03:40 to 64.84 token/s at 05:40.
  • No missing-data limitation exists: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0% across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 154.77 token/s average throughput, narrowly ahead of deepseek-v4.1-flash at 154.75 token/s; nemotron-3-ultra is the weakest at 36.89 token/s average.
  • glm-5.3-flash shows the most operationally significant volatility: after peaking at 182.04 token/s at 03:00, it fell below 12 token/s for seven consecutive observations (03:20 through 05:20), with a 127.9% coefficient of variation, before recovering to 130.12 token/s.
  • No missing-data limitation applies: all 8 models delivered 12 of 12 expected samples, yielding 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model at 165.19 token/s average throughput (peak 254.14 token/s), while nemotron-3-ultra is the weakest at 42.42 token/s average, peaking at only 76.96 token/s.
  • The most operationally significant volatility is glm-5.3-flash, which fell from a 191.44 token/s maximum to roughly 5-11 token/s in most observations after 02:20 UTC, with a 106.5% coefficient of variation and a -94.4% trend.
  • No missing-data limitation applies: all eight models report 12 of 12 expected samples, 96 valid points, and 100.0% coverage for the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model at 164.14 token/s average throughput, peaking at 254.14 token/s; minimax-m3 is the weakest at 49.33 token/s average, with a low of 10.31 token/s.
  • glm-5.3-flash shows the most severe volatility, with a 78.8% coefficient of variation and repeated collapses to roughly 7-12 token/s at 01:00, 02:20, 03:20, and 03:40 UTC; deepseek-v4.1-flash also swung from 254.14 token/s at 02:20 down to 21.39 token/s at 03:40 before recovering to 201.4 token/s.
  • No missing-data limitation applies: all eight models recorded 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model at 175.68 token/s average throughput, peaking at 233.3 token/s, while nemotron-3-ultra is the weakest at 40.98 token/s average, never exceeding 77.48 token/s.
  • The most operationally significant event is a synchronized throughput dip at 01:00 UTC, when glm-5.3-flash fell to 6.91 token/s, minimax-m3 to 10.31 token/s, and gemma4:31b to 36.03 token/s; minimax-m3 also shows the highest volatility with a 47.0% coefficient of variation.
  • No missing-data limitation applies: all eight models have 12 of 12 expected samples, valid_point_count is 96, and coverage is 100.0%, so the four-hour window is fully represented.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model at 173.88 token/s average throughput, peaking at 233.3 token/s; nemotron-3-ultra is weakest at 43.08 token/s average, never exceeding 77.48 token/s.
  • The most operationally significant volatility is the 01:00 collapse in glm-5.3-flash, which fell from 175.51 token/s at 00:20 to 6.91 token/s; deepseek-v4.1-flash also dipped to 33.76 token/s at 23:40 before recovering to 229.94 token/s.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, and the dataset reports 96 valid points with 100.0% coverage.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model on average throughput at 143.0 token/s, ahead of deepseek-v4.1-flash at 164.21 token/s only on peaks (233.3 token/s maximum); glm-5.2 is the weakest at 43.18 token/s average, below nemotron-3-ultra at 42.0 token/s only marginally.
  • minimax-m3 shows the most operationally significant volatility, swinging between 11.52 and 105.05 token/s with 54.8% coefficient of variation and a 47.4% trend; deepseek-v4.1-flash also dropped to 33.76 token/s at 23:40.
  • No missing-data limitation exists: all 8 models report 12 of 12 samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model by average throughput at 163.79 token/s, while nemotron-3-ultra is the weakest at 42.45 token/s average.
  • glm-5.3 shows the most operationally significant volatility, with a coefficient of variation of 39.8% and swings between 31.52 and 182.87 token/s, alongside an 18.9% downward trend; by contrast, deepseek-v4.1-flash improved 26.3% over the window.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0%, so the full four-hour window is represented.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model with an average throughput of 140.28 token/s (peaking at 176.61 token/s), while nemotron-3-ultra is the weakest at 37.99 token/s average, never exceeding 66.59 token/s across the window.
  • deepseek-v4.1-flash shows the most operationally significant volatility, swinging from a low of 9.07 token/s at 18:40 to a high of 215.65 token/s at 21:40, with a 43.6% coefficient of variation and a 65.6% upward trend; glm-5.3 also oscillated between 31.52 and 182.87 token/s.
  • No missing-data limitation applies: all eight models have 12 of 12 expected samples, valid_point_count is 96, and coverage is 100.0% for the full four-hour period.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput was gemma4:31b at 142.86 token/s across 12 samples (range 96.61–171.19 token/s); weakest was nemotron-3-ultra at 27.06 token/s across 12 samples (range 6.09–62.14 token/s).
  • deepseek-v4.1-flash showed the widest swings, spanning 9.07 to 231.01 token/s with a 54.4% coefficient of variation, while glm-5.3 rose 63.4% overall to a 174.8 token/s latest reading; deepseek-v4-pro also dipped to 37.21 token/s at 20:40 before recovering to 115.68 token/s.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, 96 valid points, and 100.0% coverage, though the four-hour window limits longer-term conclusions.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 141.79 token/s average throughput, peaking at 163.56 token/s; nemotron-3-ultra is the weakest at 23.86 token/s average, ranging from 6.09 to 62.14 token/s.
  • deepseek-v4.1-flash shows the most operationally significant volatility, with a 53.2% coefficient of variation, swings from 231.01 token/s at 17:20 down to 9.07 token/s at 18:40, and a -33.2% trend; glm-5.3 also oscillates between 42.51 and 182.87 token/s.
  • No missing-data limitation applies: all eight models report 12 of 12 samples, giving 96 valid points and 100.0% coverage across the four-hour window.