← Performance dashboard

Hourly performance insights

Summaries of rolling four-hour performance data

1198 retained summaries
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model at 127.99 token/s average throughput (peak 169.43 token/s), while nemotron-3-ultra is the weakest at 4.87 token/s average, never exceeding 10.46 token/s.
  • The most operationally significant volatility is gemma4:31b, which swung between 27.97 and 173.46 token/s (CV 42.4%); deepseek-v4-pro also dropped to 9.58 token/s at 06:00 UTC before recovering to 108.17 token/s, and glm-5.3-flash spiked to 166.43 token/s at 07:00.
  • No missing-data limitation exists: all eight models have 12 of 12 samples, 96 valid points, and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model at 118.24 token/s average throughput, ahead of gemma4:31b at 114.13 token/s; nemotron-3-ultra is the weakest at 4.99 token/s average, never exceeding 10.46 token/s.
  • The most operationally significant event is deepseek-v4-pro falling to 9.58 token/s at 06:00 from 97.77 token/s at 05:40; gemma4:31b also shows sharp volatility, ranging 27.97 to 173.46 token/s with cv 39.9%.
  • No missing-data limitation: all eight models recorded 12 of 12 samples, coverage is 100.0%, and all 96 valid points are present, so the four-hour window is fully represented.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model with an average throughput of 117.41 token/s, while nemotron-3-ultra is the weakest at 4.36 token/s, roughly 27 times slower.
  • deepseek-v4-flash shows the most volatility, swinging between 25.5 and 163.52 token/s with a coefficient of variation of 54.5% and a 54.6% upward trend; glm-5.3 is steadier, rising 24.8% to a peak of 144.19 token/s.
  • No missing-data limitation applies: all eight models have 12 of 12 expected samples, valid_point_count is 96, and coverage is 100.0% across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model at 110.21 token/s average throughput, peaking at 132.99 token/s; nemotron-3-ultra is the weakest at 3.56 token/s average, never exceeding 5.27 token/s.
  • deepseek-v4-flash shows the most operationally significant volatility, with a 52.5% coefficient of variation, swings between 25.5 and 142.91 token/s, and a 51.0% upward trend; gemma4:31b is also volatile at 43.0% CV.
  • No missing-data limitation: all eight models have 12 of 12 expected samples, 96 valid points, and 100.0% coverage over the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model by average throughput at 113.65 token/s, ahead of gemma4:31b at 105.62 token/s and deepseek-v4-pro at 96.95 token/s; nemotron-3-ultra is the weakest at 9.83 token/s.
  • The most significant change is nemotron-3-ultra's collapse: it peaked at 53.03 token/s at 23:40, then ran between 2.07 and 5.27 token/s from 00:00 onward, a -81.0% trend. deepseek-v4-flash was the most volatile, swinging between 25.5 and 158.01 token/s with a 58.2% coefficient of variation.
  • No missing-data limitation applies: all 8 models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0% across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model at 124.29 token/s average throughput, with a latest reading of 122.07 token/s; nemotron-3-ultra is the weakest at 11.70 token/s average, ending at just 2.07 token/s.
  • The most operationally significant volatility is nemotron-3-ultra, whose 122.4% coefficient of variation and -80.2% trend reflect a fall from a 53.03 token/s peak to 2.07 token/s; deepseek-v4-flash also swung between 158.01 and 25.5 token/s with a -57.5% trend.
  • No missing-data limitation applies: all 8 models recorded 12 of 12 samples, giving 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 delivered the strongest average throughput at 128.0 token/s, while nemotron-3-ultra was weakest at 11.79 token/s, roughly a tenth of the leader; deepseek-v4-flash (116.64 token/s) and gemma4:31b (105.96 token/s) followed closely behind glm-5.3.
  • The most significant volatility came from nemotron-3-ultra, which spiked to 53.03 token/s at 23:40 before collapsing back to 3.61 token/s at 00:00, with a coefficient of variation of 121.0%; deepseek-v4-flash also fell from 162.15 token/s at 21:20 to 36.91 token/s at 00:40.
  • No missing-data limitation exists: all 8 models report 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model at 130.09 token/s average throughput (p95 146.30 token/s); nemotron-3-ultra is the weakest at 11.48 token/s average, with a maximum of just 53.03 token/s.
  • nemotron-3-ultra shows the most operationally significant volatility: it stays between 2.21 and 5.24 token/s for the first two hours, spikes to 53.03 token/s at 23:40, then falls to 3.61 token/s at 00:00, a 125.7% coefficient of variation.
  • There is no missing-data limitation: every model has 12 of 12 samples, and the dataset reports 96 valid points at 100.0% coverage, though it spans only this single four-hour window.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput is deepseek-v4-flash at 125.22 token/s (peak 162.15 token/s); weakest is nemotron-3-ultra at 5.35 token/s (peak 11.99 token/s), roughly 23x lower than the leader.
  • The most operationally significant volatility is minimax-m3's final spike: 44.5 token/s at 22:40 jumping to 151.13 token/s at 23:00, nearly triple its 57.74 token/s average; gemma4:31b likewise swung from a 172.84 token/s peak down to 51.6 token/s at 23:00.
  • No missing-data limitation exists in this window: all eight models report 12 of 12 samples, 96 valid points total, and 100.0% coverage from 19:20 to 23:00 UTC.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash leads average throughput at 125.68 token/s (p95 159.75 token/s), while nemotron-3-ultra is weakest at 3.68 token/s average, peaking at only 5.24 token/s.
  • glm-5.3-flash shows the steepest decline, trending -47.0% from 59.65 token/s at 18:20 to 22.87 token/s at 22:00; gemma4:31b is the most volatile with cv 37.1%, swinging between 60.04 and 172.84 token/s.
  • No missing-data limitation: all eight models report 12 of 12 samples, 96 valid points, and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model at 116.07 token/s average throughput, narrowly ahead of glm-5.3 at 115.03 token/s; nemotron-3-ultra is the weakest at 3.34 token/s average, peaking at only 4.74 token/s.
  • minimax-m3 shows the most significant volatility, swinging between 167.86 and 38.41 token/s with 59.3% coefficient of variation and a -49.7% trend, falling from a 167.86 token/s peak at 18:00 to roughly 50 token/s after 19:00.
  • No missing-data limitation applies: all eight models have 12 of 12 expected samples, and the dataset shows 96 valid points with 100.0% coverage over the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput was deepseek-v4-flash at 124.41 token/s (p95 145.01 token/s), while nemotron-3-ultra was weakest at 2.95 token/s, roughly 42 times lower than the leader.
  • The most operationally significant volatility is minimax-m3, with a coefficient of variation of 65.1 percent, standard deviation of 47.41 token/s, and swings between 29.44 and 167.86 token/s, plus a 19.8 percent downward trend; glm-5.3 also spiked to 222.06 token/s at 17:40.
  • No missing-data limitation exists in this window: all eight models report 12 of 12 expected samples, 96 valid points, and 100.0 percent coverage, so the figures reflect complete four-hour observations.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model by average throughput at 117.29 token/s (12 of 12 samples, peak 145.22 token/s), while nemotron-3-ultra is the weakest at 2.40 token/s average, never exceeding 4.74 token/s.
  • minimax-m3 shows the most operationally significant volatility, with a coefficient of variation of 65.9% and swings between 29.44 and 167.86 token/s, including a jump from 55.05 to 167.86 token/s between 17:40 and 18:00 UTC.
  • No missing-data limitation exists in this window: all 8 models report 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour period ending 19:03 UTC.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput was deepseek-v4-flash at 103.26 token/s, ahead of glm-5.3 at 96.31 token/s and deepseek-v4-pro at 91.47 token/s; weakest was nemotron-3-ultra at 1.98 token/s, with a maximum of only 2.85 token/s.
  • The most significant volatility came from minimax-m3, which averaged 69.02 token/s but swung between 29.44 and 167.86 token/s with a coefficient of variation of 70.1%, including repeated spikes above 140 token/s at 16:00, 17:20, and 18:00; glm-5.3 also spiked to 222.06 token/s at 17:40.
  • No missing-data limitation exists in this window: all 8 models delivered 12 of 12 expected samples, giving 96 valid points and 100.0% coverage across the four-hour period.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model at 102.34 token/s average throughput, while nemotron-3-ultra is the weakest at 2.20 token/s, roughly 46 times lower.
  • minimax-m3 shows the most operationally significant volatility, with a 67.9% coefficient of variation, swinging between 162.45 token/s at 13:40 and 29.44 token/s at 17:00, and a -24.1% trend; glm-5.2 also dipped to 14.01 token/s at 15:40.
  • No missing-data limitation exists: all 8 models have 12 of 12 expected samples, 96 of 96 valid points, and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model at 98.14 token/s average throughput (peak 131.51 token/s), while nemotron-3-ultra is the weakest at 2.93 token/s average, peaking at only 6.11 token/s.
  • minimax-m3 shows the most operationally significant volatility, with a coefficient of variation of 66.5% and swings between 36.86 and 162.45 token/s, plus a -39.2% trend; its latest reading jumped to 144.42 token/s after five samples below 48 token/s.
  • No missing-data limitation exists: all 8 models delivered 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model at 106.45 token/s average throughput, peaking at 125.11 token/s, while nemotron-3-ultra is the weakest at 4.02 token/s average and never exceeded 6.85 token/s.
  • minimax-m3 shows the most operationally significant volatility, with a 66.1% coefficient of variation and swings between 27.73 and 170.83 token/s across consecutive 20-minute observations; deepseek-v4-pro also fell to 16.53 token/s at 12:40 before recovering to 111.08 token/s.
  • No missing-data limitation applies: all 8 models recorded 12 of 12 expected samples, yielding 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 108.65 token/s average throughput, with deepseek-v4-flash close behind at 105.68 token/s; nemotron-3-ultra is the weakest at 3.94 token/s average, peaking at only 6.85 token/s.
  • minimax-m3 shows the most operationally significant volatility, swinging between 27.73 and 170.83 token/s with a 66.1% coefficient of variation, while glm-5.3 declined 22.6% over the window, ending at 61.71 token/s after a 225.04 token/s peak.
  • No missing-data limitation applies: all 96 expected points are valid, coverage is 100.0%, and every model has 12 of 12 samples.
4-hour window · 96 points · 100.0% coverage
  • Highest average throughput was gemma4:31b at 108.6 token/s (peak 151.18 token/s), while nemotron-3-ultra was weakest at 3.66 token/s average, peaking at only 6.85 token/s.
  • minimax-m3 showed the most operationally significant volatility, with a 66.9% coefficient of variation and swings between 27.73 and 170.83 token/s, including alternating low and high readings from 10:20 onward; glm-5.3 also declined 16.6% to 51.32 token/s at 12:00.
  • No missing-data limitation applies: all 8 models delivered 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash posted the highest average throughput at 98.96 token/s (peak 135.22 token/s), narrowly ahead of gemma4:31b at 98.64 token/s, while nemotron-3-ultra was weakest at 3.86 token/s average with a maximum of only 6.85 token/s.
  • minimax-m3 showed the most operationally significant volatility, with a 63.8% coefficient of variation, dropping to 31.47 token/s at 08:00, spiking to 170.83 token/s at 10:40, then falling to 43.78 token/s at 11:00; glm-5.3 also spiked to 225.04 token/s at 10:00.
  • No missing-data limitation applies: all 96 expected points are valid, coverage is 100.0%, and every model has 12 of 12 samples.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput was deepseek-v4-flash at 105.59 token/s (peak 161.4 token/s), while nemotron-3-ultra was weakest at 3.76 token/s average, peaking at only 6.1 token/s.
  • The most operationally significant volatility was glm-5.3, with a 55.2% coefficient of variation and a spike to 225.04 token/s at 10:00 UTC, up 37.5% over the window; minimax-m3 also swung from a 27.25 token/s low to 112.1 token/s.
  • No missing-data limitation applies: all 8 models delivered 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour period.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model by average throughput at 108.04 token/s, ahead of deepseek-v4-flash at 101.03 token/s and deepseek-v4-pro at 97.79 token/s; nemotron-3-ultra is the weakest at 4.66 token/s, roughly 23 times slower than gemma4:31b.
  • glm-5.3 shows the most operationally significant volatility, with a coefficient of variation of 72.2% and swings between 25.19 and 245.62 token/s, including a spike to 245.62 token/s at 06:00 UTC followed by a drop to 45.52 token/s at 06:40.
  • No missing-data limitation exists in this window: all 8 models report 12 of 12 expected samples, and the dataset shows 96 valid points with 100.0% coverage.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 116.47 token/s average throughput, narrowly ahead of deepseek-v4-flash at 116.28 token/s, while nemotron-3-ultra is the weakest at 6.84 token/s average, never exceeding 17.16 token/s.
  • glm-5.3 shows the most operationally significant volatility, with a 61.7% coefficient of variation, swinging from 25.19 token/s at 05:20 to 245.62 token/s at 06:00; nemotron-3-ultra shows the steepest decline, trending down 53.3% to 2.46 token/s at 08:00.
  • No missing-data limitation applies: all eight models report 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash posted the highest average throughput at 121.88 token/s, narrowly ahead of gemma4:31b at 121.53 token/s, while nemotron-3-ultra was weakest at 9.59 token/s, peaking at only 18.68 token/s.
  • glm-5.3 showed the sharpest volatility, swinging between 25.19 and 245.62 token/s with a 61.1% coefficient of variation, and nemotron-3-ultra declined 54.9% over the window, ending at 4.81 token/s.
  • No missing-data limitation exists: all eight models delivered 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour period.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput was gemma4:31b at 123.66 token/s (p95 155.6 token/s, max 157.83 token/s); weakest was nemotron-3-ultra at 11.46 token/s, peaking at only 18.68 token/s and ending at 3.71 token/s.
  • glm-5.3 showed the most operationally significant volatility, with a 65.4% coefficient of variation, a low of 12.24 token/s at 02:20, and a spike to 245.62 token/s at 06:00; deepseek-v4-flash also dropped from 137.22 to 38.87 token/s at 05:40.
  • No missing-data limitation applies: all 8 models delivered 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b delivered the highest average throughput at 126.69 token/s (peaking at 168.96 token/s), while nemotron-3-ultra was weakest at 16.83 token/s average, never exceeding 66.98 token/s after the first observation.
  • glm-5.3 showed the most operationally significant volatility, swinging between 12.24 and 141.22 token/s with a 46.9% coefficient of variation, and nemotron-3-ultra was even less stable at 94.0%, collapsing from 66.98 token/s to single digits within 40 minutes.
  • No missing-data limitation applies: all eight models recorded 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model at 122.73 token/s average throughput, peaking at 147.10 token/s, while nemotron-3-ultra is the weakest at 24.67 token/s average, with a low of 5.19 token/s.
  • The most operationally significant volatility is nemotron-3-ultra's collapse from 66.98 token/s at 01:20 to 5.19 token/s at 03:20, a -62.0% trend with an 84.0% coefficient of variation; minimax-m3 also spiked to 106.49 token/s at 04:00.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0%, so the four-hour window is complete.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 128.29 token/s average throughput, narrowly ahead of deepseek-v4-flash at 126.76 token/s; minimax-m3 is the weakest at 37.4 token/s average, with nemotron-3-ultra slightly higher at 33.88 token/s.
  • nemotron-3-ultra shows the most operationally significant volatility, with a 78.3% coefficient of variation, a -56.8% trend, and a fall from 89.35 token/s at 23:40 to 9.21 token/s at 03:00; glm-5.3 also swung between 142.71 and 12.24 token/s.
  • No missing-data limitation applies: all 8 models have 12 of 12 samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model by average throughput at 137.6 token/s (peak 182.96 token/s), while nemotron-3-ultra is the weakest at 38.51 token/s average, below minimax-m3's 42.19 token/s.
  • The most operationally significant volatility is nemotron-3-ultra, with a coefficient of variation of 62.6%, a range of 6.08 to 89.35 token/s, and a latest reading of just 6.08 token/s; glm-5.3 also swings between 27.6 and 147.59 token/s (cv 42.0%).
  • No missing-data limitation exists: all 8 models have 12 of 12 expected samples, and the dataset reports 96 valid points with 100.0% coverage over the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model at 141.01 token/s average throughput (95.34–171.16 token/s range), while nemotron-3-ultra is the weakest at 36.63 token/s average, dipping to 4.5 token/s at 22:00 UTC.
  • glm-5.3 shows the most operationally significant volatility, with a 50.3% coefficient of variation and swings between 18.56 and 147.59 token/s, including a drop to 27.6 token/s at 23:00 after peaking at 147.59 token/s at 22:40; nemotron-3-ultra is similarly erratic at 60.0% CV.
  • No missing-data limitation exists: all eight models have 12 of 12 samples, 96 valid points, and 100.0% coverage, though the four-hour window limits longer-term conclusions.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model at 142.93 token/s average throughput (peak 171.16 token/s), while nemotron-3-ultra is the weakest at 30.15 token/s average, dipping to 4.5 token/s at 22:00 UTC.
  • glm-5.3 shows the most operationally significant volatility, with a 60.6% coefficient of variation and swings from 12.79 to 147.59 token/s; nemotron-3-ultra is similarly erratic at 70.4% CV, including a 4.5 token/s low.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, with 96 valid points and 100.0% coverage, so the four-hour window is fully populated.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput belongs to gemma4:31b at 118.48 token/s, with deepseek-v4-flash close behind at 127.08 token/s; weakest is nemotron-3-ultra at 19.72 token/s, roughly one-sixth of the leader.
  • glm-5.3 shows the most operationally significant volatility, with a 59.8% coefficient of variation and swings between 12.79 and 177.07 token/s, including a drop to 27.6 token/s in the final observation at 23:00.
  • No missing-data limitation exists in this window: all eight models report 12 of 12 samples, totaling 96 valid points at 100.0% coverage across the four-hour period.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput is deepseek-v4-flash at 126.53 token/s (peak 165.47 token/s); weakest is nemotron-3-ultra at 16.29 token/s, which never exceeded 36.36 token/s.
  • glm-5.3 shows the most operationally significant volatility: coefficient of variation 59.9%, swinging from 177.07 token/s at 19:20 down to 12.79 token/s at 21:00, ending at 18.56 token/s with a -46.7% trend.
  • No missing-data limitation exists: all eight models have 12 of 12 samples, 96 valid points, and 100.0% coverage; the only constraint is the four-hour observation window itself.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model at 123.83 token/s average throughput, while nemotron-3-ultra is the weakest at 18.20 token/s average, roughly a 6.8x gap between them.
  • glm-5.3 shows the most operationally significant volatility, with a 54.1% coefficient of variation, peaking at 177.07 token/s at 19:20 before collapsing to 12.79 token/s at 21:00; nemotron-3-ultra is similarly erratic at 93.1% CV, ranging from 2.67 to 70.41 token/s.
  • No missing-data limitation exists: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0%, though the four-hour window limits longer-term conclusions.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model at 118.88 token/s average throughput, while nemotron-3-ultra is the weakest at 14.46 token/s average, roughly eight times slower.
  • glm-5.3 shows the most operationally significant volatility, with a 57.9% coefficient of variation and swings from 17.67 to 177.07 token/s; deepseek-v4-flash also declined from 165.47 token/s at 18:20 to 62.14 token/s at 20:00.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0%, so the four-hour window is complete.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model by average throughput at 123.9 token/s, peaking at 165.47 token/s; nemotron-3-ultra is the weakest at 12.63 token/s average, with a low of 0.52 token/s.
  • The most significant trend is minimax-m3's decline of 43.1%, falling from 149.51 token/s at 15:40 to 36.76 token/s at 19:00. glm-5.3 shows the highest volatility (CV 56.9%), swinging between 17.67 and 172.09 token/s, while glm-5.3-flash is the steadiest (CV 9.1%).
  • No missing-data limitation exists: all 96 expected points are valid, coverage is 100.0%, and every model has 12 of 12 samples.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model by average throughput at 112.27 token/s, while nemotron-3-ultra is the weakest at 8.58 token/s, roughly thirteen times lower.
  • The most operationally significant volatility is nemotron-3-ultra, which sat near 1-5 token/s for most of the window before spiking to 70.41 token/s at 17:40, yielding a 219.5% coefficient of variation; glm-5.3 also swung from a 178.7 token/s peak down to 17.67 token/s.
  • No missing-data limitation exists: all eight models have 12 of 12 expected samples, valid_point_count is 96, and coverage is 100.0% across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model at 117.25 token/s average throughput, ahead of deepseek-v4-flash at 113.07 token/s; nemotron-3-ultra is the weakest at 1.98 token/s average, never exceeding 4.74 token/s.
  • glm-5.3 shows the sharpest volatility, swinging from a 200.95 token/s peak at 13:40 UTC to 17.67 token/s at 16:20 UTC, a -40.8% trend; minimax-m3 also ranged between 34.42 and 149.51 token/s.
  • No missing-data limitation: all 8 models have 12 of 12 samples, 96 valid points, and 100.0% coverage over the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model at 126.65 token/s average throughput, peaking at 200.95 token/s, while nemotron-3-ultra is the weakest at 2.03 token/s average, never exceeding 3.2 token/s.
  • The most operationally significant volatility is gemma4:31b, with a 57.0% coefficient of variation and swings between 5.83 and 134.49 token/s; minimax-m3 shows the steepest trend at 64.0%, spiking to 149.51 token/s at 15:40 UTC.
  • No missing-data limitation applies: all eight models report 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 88 points · 91.7% coverage
  • deepseek-v4-flash is the strongest model at 110.39 token/s average throughput, with a peak of 167.04 token/s, while nemotron-3-ultra is the weakest at 2.90 token/s average, never exceeding 6.76 token/s.
  • glm-5.3 shows the most operationally significant volatility, rising 107.2% across the window from 59.59 token/s to a 200.95 token/s spike at 13:40 UTC, then falling to 74.74 token/s at 14:20; deepseek-v4-pro also dropped to 40.0 token/s at 13:40.
  • Each model is missing one of its 12 expected samples, with 11 collected per model, giving 88 valid points out of 96 expected and 91.7% coverage, so the final interval is unobserved for all models.
4-hour window · 88 points · 91.7% coverage
  • deepseek-v4-flash posted the highest average throughput at 108.06 token/s (peak 167.04 token/s), while nemotron-3-ultra was the weakest at 4.13 token/s average, peaking at only 6.76 token/s.
  • The most operationally significant volatility is deepseek-v4-flash's swing from 167.04 token/s at 11:20 down to 21.16 token/s at 12:00, with a 47.4% coefficient of variation; glm-5.3 shows the steepest upward trend at +74.3%, ending at 200.95 token/s.
  • Every model recorded 11 of 12 expected samples (91.7% coverage), leaving 8 missing points across the window (88 valid of 96 expected), so short-lived dips or spikes may be unobserved.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput was deepseek-v4-flash at 118.05 token/s (p95 179.64 token/s, peak 195.05 token/s), while nemotron-3-ultra was weakest at 4.65 token/s average, never exceeding 6.76 token/s.
  • The most significant volatility came from deepseek-v4-flash (cv 46.5%), which collapsed to 16.24 token/s at 10:20 and 21.16 token/s at 12:00 before partially recovering to 70.48 token/s; gemma4:31b also fell sharply to 5.83 token/s at 12:40.
  • No missing-data limitation applies: all 8 models recorded 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model on average throughput at 121.78 token/s (p95 179.64 token/s), while nemotron-3-ultra is the weakest at 5.83 token/s average, never exceeding 9.84 token/s.
  • The sharpest declines hit glm-5.3 (trend -33.5%) and glm-5.3-flash (-31.4%); deepseek-v4-flash shows the highest volatility (cv 43.5%), swinging between 195.05 and 16.24 token/s, with its latest reading at 21.16 token/s.
  • No missing-data limitation: all 8 models have 12 of 12 samples, 96 valid points, and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model at 116.87 token/s average throughput, peaking at 195.05 token/s; nemotron-3-ultra is the weakest at 5.76 token/s average, never exceeding 9.84 token/s.
  • glm-5.3 shows the most operationally significant volatility and decline, swinging between 181.84 and 14.15 token/s (CV 53.6%) and ending at its 14.15 token/s minimum, a -42.7% trend; deepseek-v4-flash also dipped sharply to 16.24 token/s at 10:20 before recovering to 143.87 token/s.
  • No missing-data limitation applies: all 8 models delivered 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model with an average throughput of 119.95 token/s (peak 212.99 token/s at 06:40), while nemotron-3-ultra is the weakest at an average of 5.03 token/s, never exceeding 9.84 token/s across the window.
  • The most operationally significant movement is glm-5.3's decline: after sustaining 131.92 to 212.99 token/s through 07:40, it fell to roughly 62 to 68 token/s between 08:00 and 08:40, a trend of -40.9%. deepseek-v4-flash moved the opposite way, climbing from 58.72 token/s at 06:40 to 195.05 token/s at 09:40, a 55.6% trend.
  • No missing-data limitation applies: all 8 models delivered 12 of 12 expected samples, with 96 valid points and 100.0% coverage, so every twenty-minute interval from 06:20 to 10:00 UTC is represented.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model with an average throughput of 127.29 token/s (peaking at 212.99 token/s), while nemotron-3-ultra is the weakest at 4.72 token/s average, never exceeding 9.84 token/s across the window.
  • The most operationally significant volatility comes from glm-5.3, which swung from 212.99 token/s at 06:40 down to 61.93 token/s at 08:00, with a coefficient of variation of 41.8%; nemotron-3-ultra also trended upward 85.9% but from a very low base.
  • No missing-data limitation applies: all eight models recorded 12 of 12 expected samples, with 96 valid points and 100.0% coverage over the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model at 140.38 token/s average throughput, peaking at 212.99 token/s; nemotron-3-ultra is the weakest at 4.4 token/s average, never exceeding 7.54 token/s.
  • deepseek-v4-flash shows the steepest decline, trending -35.0% from a 150.7 token/s peak at 04:40 to 94.17 token/s at 08:00; glm-5.3 is also volatile, swinging between 61.93 and 212.99 token/s (CV 37.1%).
  • No missing-data limitation: all 8 models have 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model at 125.21 token/s average throughput, peaking at 212.99 token/s; nemotron-3-ultra is weakest at 3.91 token/s average, never exceeding 6.96 token/s.
  • glm-5.3 shows the most operationally significant volatility, swinging between 69.63 and 212.99 token/s with a 42.2% coefficient of variation while trending up 22.1%; deepseek-v4-flash fell 21.9% overall, dropping from 150.7 to 74.28 token/s.
  • No missing-data limitation exists: all eight models delivered 12 of 12 expected samples, yielding 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model at 122.46 token/s average throughput, peaking at 150.70 token/s, while nemotron-3-ultra is the weakest at 3.85 token/s average, never exceeding 6.96 token/s.
  • glm-5.3 shows the most operationally significant volatility, with a 46.0% coefficient of variation and swings between 66.81 and 189.77 token/s; glm-5.3-flash also dropped to 13.93 token/s at 04:00 before recovering to 126.30 token/s at 06:00.
  • No missing-data limitation applies: all 8 models report 12 of 12 samples each, 96 valid points overall, and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model at 111.82 token/s average throughput, narrowly ahead of gemma4:31b at 110.42 token/s and deepseek-v4-pro at 108.16 token/s; nemotron-3-ultra is the weakest at 3.71 token/s average, never exceeding 6.96 token/s.
  • glm-5.3 shows the highest volatility (cv 45.7%), ranging from 64.79 to 189.77 token/s, and glm-5.3-flash dropped to 13.93 token/s at 04:00 before rebounding to 129.99 token/s at 04:20; deepseek-v4-flash trended up 36.4% to a 150.70 token/s peak.
  • No samples are missing: all eight models report 12 of 12 expected observations, 96 valid points, and 100.0% coverage, so the only limitation is the short four-hour window.