← Performance dashboard

Hourly performance insights

Summaries of rolling four-hour performance data

1186 retained summaries
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model at an average 184.46 token/s (peaking at 235.93 token/s), while nemotron-3-ultra is the weakest at an average 25.54 token/s, never exceeding 52.47 token/s.
  • glm-5.2 shows the most operationally significant volatility, with a 64.2% coefficient of variation and swings from 10.78 to 96.62 token/s within the four hours; nemotron-3-ultra is similarly erratic at 57.9%.
  • No missing-data limitation exists: all eight models report 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the period.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model at 186.75 token/s average throughput (peak 235.93 token/s), while nemotron-3-ultra is the weakest at 22.48 token/s average (minimum 3.15 token/s).
  • glm-5.2 shows the most extreme volatility, ranging from 10.78 to 216.7 token/s with a 94.6% coefficient of variation; gemma4:31b also swung from 146.13 down to 43.95 token/s late in the window.
  • No missing-data limitation exists: all eight models report 12 of 12 expected samples, 96 valid points, and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput was deepseek-v4.1-flash at 193.04 token/s (peaking at 235.93 token/s), while nemotron-3-ultra was weakest at 21.48 token/s, never exceeding 42.75 token/s.
  • glm-5.2 showed the most extreme volatility, with a coefficient of variation of 96.7%: it spiked to 216.7 token/s at 20:40 then fell to 10.78 token/s at 21:20, and all models trended downward except deepseek-v4-pro (+3.9%) and nemotron-3-ultra (+27.3%).
  • No missing-data limitation applies: all 8 models delivered 12 of 12 expected samples, and the dataset achieved 100.0% coverage with 96 of 96 valid points.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model at 186.07 token/s average throughput, peaking at 230.07 token/s, while nemotron-3-ultra is the weakest at 20.1 token/s average, with a maximum of only 40.72 token/s.
  • glm-5.2 shows the most operationally significant volatility, ranging from 10.78 to 216.7 token/s with a 105.7% coefficient of variation, including a spike to 216.7 token/s at 20:40 UTC followed by a drop to 10.78 token/s at 21:20 UTC; deepseek-v4.1-flash also trended upward 9.3%.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0% across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model at 184.08 token/s average throughput, peaking at 230.07 token/s, while nemotron-3-ultra is the weakest at 21.36 token/s average, dipping to 4.27 token/s.
  • glm-5.2 shows the most extreme volatility, ranging from 11.04 to 216.7 token/s with a 96.0% coefficient of variation; glm-5.3 posted the steepest climb, up 52.0% to a 175.32 token/s peak at 20:20.
  • No missing-data limitation exists in this window: all eight models delivered 12 of 12 expected samples, with 96 valid points and 100.0% coverage.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model at 181.27 token/s average throughput (peak 228.81 token/s), while nemotron-3-ultra is the weakest at 20.47 token/s average, never exceeding 40.72 token/s.
  • glm-5.2 shows the most operationally significant volatility, with a 70.2% coefficient of variation and swings from 13.18 to 134.30 token/s; glm-5.3 similarly ranged from 16.72 to 163.17 token/s, indicating unstable throughput for both.
  • No missing data: all eight models have 12 of 12 samples with 100.0% coverage and 96 valid points, so the only limitation is the short four-hour window, which may not represent longer-term behavior.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model at 164.93 token/s average throughput, peaking at 228.81 token/s; nemotron-3-ultra is the weakest at 20.23 token/s average, never exceeding 40.72 token/s.
  • glm-5.2 shows the most operationally significant volatility, with a coefficient of variation of 68.5% and swings between 13.18 and 134.3 token/s; glm-5.3 similarly ranged from 16.72 to 142.96 token/s, indicating unstable throughput.
  • No missing-data limitation exists: all eight models recorded 12 of 12 expected samples, totaling 96 valid points at 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model at 162.24 token/s average throughput (peak 228.81 token/s), while nemotron-3-ultra is the weakest at 19.68 token/s average (minimum 7.74 token/s).
  • glm-5.2 shows the most operationally significant volatility, with an 80.5% coefficient of variation, swings from 16.12 to 195.22 token/s, and a -44.6% trend; deepseek-v4.1-flash trended up 32.2%, closing at 185.94 token/s.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, 96 valid points, and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model at 166.42 token/s average throughput, peaking at 229.21 token/s; nemotron-3-ultra is weakest at 19.27 token/s average, never exceeding 33.31 token/s.
  • glm-5.2 shows the most operationally significant volatility, with a coefficient of variation of 80.0% and a swing from 16.12 to 195.22 token/s, plus a -42.2% trend; deepseek-v4.1-flash is steadier at 27.8% variation despite a 229.21-to-85.42 token/s range.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0%, so the four-hour window is fully represented.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model at 160.42 token/s average throughput (peak 229.21 token/s), while nemotron-3-ultra is the weakest at 18.01 token/s average (peak 33.31 token/s).
  • glm-5.2 shows the most volatility, with a coefficient of variation of 86.5% and swings from 7.84 to 195.22 token/s; a broad dip near 14:00 UTC hit nearly all models, including minimax-m3 at 4.59 token/s and gemma4:31b at 47.71 token/s.
  • No missing-data limitation exists: all eight models have 12 of 12 samples and 96 valid points at 100% coverage, though the four-hour window alone cannot confirm whether the 14:00 dip is recurring.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model at 169.92 token/s average throughput (peak 229.21 token/s), while nemotron-3-ultra is the weakest at 19.01 token/s average (peak 33.31 token/s).
  • The 14:00 UTC observation shows a synchronized dip: gemma4:31b fell to 47.71 token/s, minimax-m3 to 4.59 token/s, and deepseek-v4-pro to 23.05 token/s; glm-5.2 is the most volatile model with a 79.8% coefficient of variation, swinging between 7.84 and 195.97 token/s.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, and the dataset reports 96 valid points with 100.0% coverage.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model at 182.63 token/s average throughput (peak 230.12 token/s), while nemotron-3-ultra is the weakest at 18.48 token/s average (peak 33.31 token/s).
  • The most significant trend is minimax-m3's decline of 55.1% over the window, ending at 4.59 token/s; glm-5.2 shows the highest volatility, swinging between 4.76 and 195.97 token/s with an 85.0% coefficient of variation.
  • No missing-data limitation: all 8 models report 12 of 12 expected samples, 96 valid points, and 100.0% coverage, so the four-hour window is complete.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model at 169.36 token/s average throughput (peak 230.12 token/s), while nemotron-3-ultra is the weakest at 17.71 token/s average (peak 25.64 token/s), a roughly tenfold gap between the two.
  • glm-5.2 shows the most operationally significant volatility, with a coefficient of variation of 83.5% and swings from 4.76 token/s at 11:00 to 211.18 token/s at 10:00, including a late dip to 7.84 token/s at 12:40 before recovering to 116.02 token/s.
  • No missing-data limitation exists in this window: all 8 models report 12 of 12 expected samples, giving 96 valid points and 100.0% coverage across the four-hour period.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model at 174.26 token/s average throughput, peaking at 230.12 token/s; nemotron-3-ultra is the weakest at 19.79 token/s average, never exceeding 26.99 token/s.
  • glm-5.2 shows the most operationally significant volatility, with a 78.4% coefficient of variation and swings from 4.76 token/s at 11:00 to 211.18 token/s at 10:00; minimax-m3 also ranged between 12.2 and 74.93 token/s.
  • No missing-data limitation applies: all eight models report 12 of 12 samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model at 172.39 token/s average throughput, peaking at 237.96 token/s; nemotron-3-ultra is the weakest at 22.34 token/s average, never exceeding 32.73 token/s.
  • glm-5.2 shows the most volatility (77.7% CV), spiking to 211.18 token/s at 10:00 UTC before collapsing to 4.76 token/s at 11:00 UTC; deepseek-v4.1-flash also dipped to 63.63 token/s at 09:40 UTC against its 172.39 token/s average.
  • No missing-data limitation exists: all eight models report 12 of 12 samples, 96 valid points, and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model at 167.71 token/s average throughput, peaking at 237.96 token/s, while nemotron-3-ultra is the weakest at 23.73 token/s average with a maximum of only 34.38 token/s.
  • glm-5.2 shows the most volatility, with a 61.5% coefficient of variation and swings from 15.88 to 211.18 token/s, including a spike to its maximum in the final 10:00 UTC sample; gemma4:31b declined 19.8% over the window.
  • No missing-data limitation exists: all eight models have 12 of 12 samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model at 189.86 token/s average throughput (peak 237.96 token/s), while nemotron-3-ultra is the weakest at 25.15 token/s average (peak 34.38 token/s).
  • Volatility is the main operational concern: glm-5.2 shows 56.5% coefficient of variation with swings from 12.61 to 140.36 token/s, and minimax-m3 dropped to 5.15 token/s at 07:00; glm-5.3 is the only clear gainer, up 14.2% to 145.77 token/s at 09:00.
  • No missing-data limitation: all 8 models have 12 of 12 samples, 96 valid points, and 100.0% coverage over the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model with an average throughput of 184.76 token/s (peaking at 237.96 token/s), while nemotron-3-ultra is the weakest at 23.7 token/s average, never exceeding 34.38 token/s across the window.
  • The most significant volatility comes from glm-5.3, whose coefficient of variation is 42.3%, swinging from 171.53 token/s at 04:40 down to 6.45 token/s at 06:00; glm-5.2 shows the steepest trend, rising 95.8% from a 12.61 token/s low to a 140.36 token/s peak.
  • No missing-data limitation applies: all eight models delivered 12 of 12 expected samples, with 96 valid points and 100.0% coverage over the four-hour period.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model at 197.65 token/s average throughput (peaking at 231.48 token/s), while nemotron-3-ultra is the weakest at 24.70 token/s average, never exceeding 36.63 token/s.
  • glm-5.2 shows the most volatility, with a 73.3% coefficient of variation and swings from 12.61 to 218.07 token/s; minimax-m3 also fell sharply to 5.15 token/s at 07:00, its lowest reading.
  • No missing-data limitation exists: all eight models have 12 of 12 expected samples, valid_point_count is 96, and coverage is 100.0% across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model at 189.69 token/s average throughput (peaking at 229.56 token/s), while nemotron-3-ultra is the weakest at 24.74 token/s average, roughly eight times slower.
  • glm-5.2 shows the most volatility, with a coefficient of variation of 89.5% and swings from 218.07 token/s at 04:00 down to 12.61 token/s at 05:20; glm-5.3 also dropped sharply to 6.45 token/s in the final 06:00 observation.
  • No missing-data limitation exists: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0% across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model at 177.60 token/s average throughput (range 122.23 to 229.51 token/s), while nemotron-3-ultra is the weakest at 26.83 token/s average (range 10.18 to 36.63 token/s).
  • glm-5.2 shows the most operationally significant volatility, with a 78.9% coefficient of variation and swings from 20.74 token/s at 03:00 to 218.07 token/s at 04:00; deepseek-v4.1-flash is the steadiest at 18.4%.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, and the dataset records 96 valid points with 100.0% coverage, though it covers only this four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model at 186.86 token/s average throughput, peaking at 229.51 token/s, while nemotron-3-ultra is the weakest at 29.10 token/s average, never exceeding 36.63 token/s.
  • glm-5.2 shows the most operationally significant volatility, with a 73.9% coefficient of variation and swings from 20.74 to 218.07 token/s; it also declined 26.3% over the window, and glm-5.3 fell 35.4%.
  • No missing-data limitation exists: all eight models recorded 12 of 12 expected samples, yielding 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model at 175.25 token/s average throughput (peak 211.09 token/s), while nemotron-3-ultra is the weakest at 31.75 token/s average, peaking at only 58.87 token/s.
  • glm-5.3 shows the sharpest decline, trending -37.8% from roughly 180 token/s early in the window to a 53.18 token/s low at 02:20; glm-5.2 is the most volatile, with a 75.7% coefficient of variation spanning 11.08 to 198.51 token/s.
  • No missing-data limitation exists: all eight models have 12 of 12 samples, 96 valid points, and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model at 182.38 token/s average throughput (peaking at 213.39 token/s), while nemotron-3-ultra is the weakest at 32.17 token/s average (minimum 11.42 token/s).
  • glm-5.2 shows the most volatility, with a 69.3% coefficient of variation and swings from 11.08 to 198.51 token/s, including a 41.0% upward trend; by contrast deepseek-v4.1-flash stays stable with only 13.5% variation.
  • No missing-data limitation exists: all 8 models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0% across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model at 184.68 token/s average throughput (peak 213.39 token/s), while nemotron-3-ultra is the weakest at 29.10 token/s average, roughly six times slower.
  • glm-5.3 shows the steadiest improvement, rising 33.4% to a 153.15 token/s average with only 19.5% coefficient of variation, whereas glm-5.2 is the most volatile at 86.1% coefficient of variation, swinging between 11.08 and 198.51 token/s.
  • No missing-data limitation applies: all eight models delivered 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model with an average throughput of 180.61 token/s (p95 215.23 token/s, max 217.48 token/s), while nemotron-3-ultra is the weakest at 27.23 token/s average, peaking at only 58.87 token/s.
  • glm-5.2 shows the most volatility, with a coefficient of variation of 77.7% and swings from 11.08 token/s at 00:00 to a 191.19 token/s spike at 22:20; glm-5.3-flash also swung between 34.83 and 185.87 token/s within the window.
  • No missing-data limitation applies: all eight models have 12 of 12 expected samples, valid_point_count is 96, and coverage is 100.0% across the four-hour period, so the dataset is complete.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model at 183.38 token/s average throughput (peak 227.33 token/s), while nemotron-3-ultra is the weakest at 21.14 token/s average, roughly one-ninth of the leader.
  • glm-5.2 shows the most operationally significant volatility, with a coefficient of variation of 73.7 percent, swinging from 19.99 token/s at 21:40 to 191.19 token/s at 22:20; gemma4:31b also dipped to 50.18 token/s around 20:40.
  • No missing-data limitation applies: all eight models have 12 of 12 samples and 100.0 percent coverage, with 96 valid points overall, so the four-hour window is fully represented.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model at 180.13 token/s average throughput (range 129.09–227.33 token/s), while nemotron-3-ultra is the weakest at 21.11 token/s average, never exceeding 41.32 token/s.
  • glm-5.3 shows the steadiest gain, trending +19.0% to a 123.86 token/s average with the lowest volatility (cv 22.4%); glm-5.2 is the most volatile (cv 57.4%), swinging between 19.99 and 125.78 token/s, and minimax-m3 dipped to 6.95 token/s at 19:20.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, with 96 valid points and 100.0% coverage, so the four-hour window is fully represented.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model at 163.48 token/s average throughput, peaking at 227.33 token/s, while nemotron-3-ultra is the weakest at 20.78 token/s average and never exceeding 38.18 token/s.
  • The most operationally significant trend is glm-5.3's 78.7% rise, from 49.21 token/s at 18:00 to 143.23 token/s at 20:00; glm-5.2 shows the greatest volatility with a 48.3% coefficient of variation, swinging between 31.0 and 125.78 token/s.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, 96 valid points, and 100.0% coverage, so figures represent the complete four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model with 138.64 token/s average throughput, peaking at 227.33 token/s at 19:40; nemotron-3-ultra is the weakest at 22.01 token/s average, never exceeding 39.76 token/s across the window.
  • The most operationally significant movement is glm-5.3's late surge, climbing from 49.21 token/s at 18:00 to 143.23 token/s at 20:00 for a 63.4% trend gain, while glm-5.2 shows the highest relative volatility (cv 56.8%) and minimax-m3 dipped to 6.95 token/s at 19:20.
  • No missing-data limitation applies: all 96 expected points are valid (100.0% coverage), and each of the eight models reports 12 of 12 samples over the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model at 121.33 token/s average throughput, peaking at 199.80 token/s, while nemotron-3-ultra is the weakest at 23.46 token/s average and never exceeding 39.76 token/s.
  • glm-5.3 shows the most operationally significant volatility, swinging from 146.03 token/s at 15:20 down to 21.87 token/s at 17:00, with a coefficient of variation of 45.0% and a -23.6% trend; deepseek-v4.1-flash also oscillates sharply between 38.85 and 199.80 token/s.
  • No missing-data limitation exists: all eight models report 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput was deepseek-v4.1-flash at 111.66 token/s, ahead of gemma4:31b at 103.83 token/s; weakest was nemotron-3-ultra at 24.99 token/s, below minimax-m3 at 43.90 token/s.
  • glm-5.2 showed the most extreme volatility, ranging from 9.97 to 239.87 token/s with a 90.7% coefficient of variation, while glm-5.3 posted the steepest decline, falling 46.4% from 145.53 token/s at 14:20 to 49.21 token/s at 18:00.
  • No missing-data limitation applies: all eight models delivered 12 of 12 expected samples, yielding 96 valid points and 100.0% coverage, though the window covers only four hours.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 delivered the highest average throughput at 114.93 token/s (peak 154.09 token/s), while nemotron-3-ultra was the weakest at 23.15 token/s average, never exceeding 60.23 token/s.
  • glm-5.2 showed the most operationally significant volatility, swinging between 9.97 and 239.87 token/s with a coefficient of variation of 102.5%; deepseek-v4.1-flash also spiked to 199.8 token/s at 15:40 before falling to 38.85 token/s at 16:40.
  • No missing-data limitation applies: all eight models report 12 of 12 samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput belongs to deepseek-v4.1-flash at 126.29 token/s (p95 190.37 token/s); weakest is nemotron-3-ultra at 23.42 token/s, peaking at only 60.23 token/s.
  • glm-5.2 shows the most operationally significant volatility, with a 95.3% coefficient of variation, swings between 9.97 and 239.87 token/s, and a -35.1% trend; glm-5.3 rose 48.5% to average 108.11 token/s.
  • No missing-data limitation: all 8 models delivered 12 of 12 expected samples, 96 valid points, and 100.0% coverage across the 12:20–16:00 UTC observations.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model with an average throughput of 135.61 token/s (maximum 182.65 token/s), while nemotron-3-ultra is the weakest at 21.70 token/s average, peaking at only 60.23 token/s.
  • glm-5.2 shows the most operationally significant volatility, swinging between 9.97 and 239.87 token/s with a coefficient of variation of 85.4%, including a drop to 9.97 token/s at 14:20 followed by a spike to 239.87 token/s at 15:00; deepseek-v4.1-flash also declined 18.2% over the window.
  • No missing-data limitation exists in this dataset: all eight models recorded 12 of 12 expected samples, yielding 96 valid points and 100.0% coverage across the four-hour period.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model with an average throughput of 149.56 token/s (peaking at 212.23 token/s), while nemotron-3-ultra is the weakest at 17.34 token/s average, never exceeding 29.54 token/s.
  • glm-5.2 shows the most operationally significant volatility, with a coefficient of variation of 76.8% and swings from 27.27 token/s at 11:00 to 224.82 token/s at 13:40; deepseek-v4-pro also dropped to 11.72 token/s at 14:00.
  • No missing-data limitation exists: all eight models recorded 12 of 12 expected samples, yielding 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model by average throughput at 154.26 token/s, while nemotron-3-ultra is the weakest at 17.55 token/s, roughly nine times lower.
  • glm-5.2 shows the most operationally significant volatility, with a coefficient of variation of 81.9% and a trend of +106.1%; it swung from 25.95 token/s at 09:20 to a 193.32 token/s peak at 13:00, including a drop to 45.18 token/s at 12:20 after reaching 177.64 token/s at 12:00.
  • No missing-data limitation exists: all eight models have 12 of 12 expected samples, and the dataset reports 96 valid points with 100.0% coverage over the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model at 152.58 token/s average throughput, while nemotron-3-ultra is the weakest at 17.79 token/s average.
  • The most operationally significant volatility is glm-5.2, whose coefficient of variation is 64.8%; it swung from a 25.95 token/s low to 177.64 token/s at 12:00 UTC, and deepseek-v4-pro dropped sharply to 16.83 token/s in the final sample versus its 93.28 token/s average.
  • No missing-data limitation exists: all eight models delivered 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model at 146.78 token/s average throughput, peaking at 212.23 token/s; nemotron-3-ultra is the weakest at 17.57 token/s average, never exceeding 28.73 token/s.
  • glm-5.3-flash shows the most operationally significant volatility, with a 45.1% coefficient of variation and swings between 53.66 and 232.91 token/s, including a collapse from 232.91 token/s at 08:40 to 53.66 token/s at 09:00; glm-5.3 trended up 42.5%.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0%, so the four-hour window is fully represented.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput is deepseek-v4.1-flash at 142.14 token/s (p95 176.22 token/s); weakest is nemotron-3-ultra at 18.32 token/s, which never exceeded 28.73 token/s.
  • glm-5.3-flash is the most volatile model, with a 46.0% coefficient of variation and swings between 53.66 and 232.91 token/s; glm-5.3 shows the steepest decline, trending -28.7% from 170.48 token/s at 06:20 to 110.28 token/s at 10:00.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, 96 valid points, and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model by average throughput at 129.35 token/s, ahead of deepseek-v4.1-flash at 148.87 token/s... correction: deepseek-v4.1-flash leads at 148.87 token/s, with gemma4:31b second at 129.35 token/s; nemotron-3-ultra is weakest at 21.35 token/s.
  • glm-5.3 shows the sharpest decline, trending down 33.3% from a 170.48 token/s peak at 06:20 to 69.86 token/s at 09:00, while glm-5.3-flash rose 19.7% and spiked to 232.91 token/s at 08:40 before dropping to 53.66 token/s at 09:00.
  • No missing-data limitation exists: all eight models have 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput was deepseek-v4.1-flash at 128.78 token/s; weakest was nemotron-3-ultra at 24.57 token/s, roughly one-fifth of the leader's rate.
  • deepseek-v4.1-flash also showed the widest swings, ranging 14.9 to 186.13 token/s with 44.2% coefficient of variation, while glm-5.3 climbed 41.0% overall, peaking at 170.48 token/s before falling back to 112.79 token/s.
  • No missing-data limitation applies: all eight models delivered 12 of 12 expected samples, giving 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model at 135.35 token/s average throughput, peaking at 209.48 token/s; nemotron-3-ultra is the weakest at 23.57 token/s average, never exceeding 46.49 token/s.
  • The most operationally significant movement is deepseek-v4.1-flash's 52.5% upward trend, including a dip to 14.9 token/s at 05:00 before recovering to 165.38 token/s by 07:00; glm-5.2 shows the highest volatility at 65.6% coefficient of variation.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0%, so the four-hour window is fully represented.
4-hour window · 96 points · 100.0% coverage
  • Highest average throughput was deepseek-v4.1-flash at 122.44 token/s, edging glm-5.3-flash (122.37 token/s) and gemma4:31b (119.24 token/s); lowest was nemotron-3-ultra at 30.60 token/s, below minimax-m3's 37.69 token/s.
  • glm-5.2 showed the sharpest volatility, with a coefficient of variation of 82.8% and swings between a 202.79 token/s spike at 02:00 and a 21.44 token/s low at 04:40; deepseek-v4-pro trended up 53.1% while glm-5.3-flash fell 24.6%.
  • No missing-data limitation applies: all eight models delivered 12 of 12 samples, with 96 valid points and 100.0% coverage, though the window spans only four hours.
4-hour window · 96 points · 100.0% coverage
  • Highest average throughput was gemma4:31b at 133.1 token/s across 12 samples (range 81.25 to 170.9 token/s), while nemotron-3-ultra was lowest at 27.93 token/s, never exceeding 54.07 token/s.
  • glm-5.2 showed the most operationally significant volatility, with a 70.5% coefficient of variation, swings between 24.83 and 202.79 token/s, and a -31.2% trend; deepseek-v4.1-flash also swung between 24.14 and 226.05 token/s.
  • No missing-data limitation applies: all 8 models reported 12 of 12 expected samples, giving 96 valid points and 100.0% coverage over the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model at 150.73 token/s average throughput, ahead of gemma4:31b at 132.85 token/s and glm-5.3-flash at 131.83 token/s; nemotron-3-ultra is the weakest at 39.67 token/s average, below minimax-m3 at 48.47 token/s.
  • glm-5.2 shows the most volatility, with a 74.2% coefficient of variation, ranging from 24.52 to 202.79 token/s, including a spike to 202.79 token/s at 02:00 UTC; deepseek-v4.1-flash also swings sharply, dropping to 24.14 token/s at 01:20 after peaking at 226.05 token/s at 00:40.
  • No missing-data limitation exists: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0% across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model by average throughput at 140.56 token/s, narrowly ahead of gemma4:31b at 140.37 token/s; nemotron-3-ultra is the weakest at 44.39 token/s, below minimax-m3 at 55.41 token/s.
  • glm-5.2 shows the sharpest swing, rising 123.0% over the window to a peak of 202.79 token/s at 02:00 after dipping to 24.52 token/s at 00:00, with the highest volatility at 72.8% coefficient of variation; minimax-m3 fell 48.6% to 11.73 token/s.
  • No missing-data limitation applies: all eight models delivered 12 of 12 expected samples, and the dataset records 96 valid points at 100.0% coverage.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model by average throughput at 143.44 token/s, ahead of deepseek-v4.1-flash at 135.52 token/s; nemotron-3-ultra is the weakest at 46.7 token/s average, below minimax-m3 at 62.55 token/s.
  • The most operationally significant volatility is glm-5.3, which swings between 19.0 and 180.67 token/s (cv 50.8%), including a drop to 19.0 token/s at 00:00; deepseek-v4.1-flash shows the steepest upward trend at +52.8%, peaking at 226.05 token/s at 00:40.
  • No missing-data limitation applies: all 8 models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0%, so the four-hour window is fully represented.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b delivered the strongest average throughput at 129.49 token/s with a peak of 173.62 token/s, while nemotron-3-ultra was weakest at 47.04 token/s average, never exceeding 94.23 token/s.
  • glm-5.3 showed the most operationally significant volatility, swinging between 19.0 and 180.67 token/s with a 45.6% coefficient of variation; deepseek-v4-pro also fell sharply from 112.6 token/s at 23:20 to 12.59 token/s at 23:40.
  • No missing-data limitation applies: all eight models recorded 12 of 12 samples, giving 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model at 129.58 token/s average output throughput, ahead of gemma4:31b at 124.19 token/s; nemotron-3-ultra is the weakest at 37.83 token/s average, never exceeding 75.69 token/s.
  • deepseek-v4.1-flash shows the most operationally significant volatility, with a 60.6% coefficient of variation and swings from 19.01 to 185.99 token/s; glm-5.3 also declined 22.0% across the window, ending at 63.23 token/s versus its 178.26 token/s peak.
  • No missing-data limitation applies: all 8 models delivered 12 of 12 expected samples, totaling 96 valid points at 100.0% coverage, though the four-hour window restricts longer-term assessment.