← Performance dashboard

Hourly performance insights

Summaries of rolling four-hour performance data

1197 retained summaries
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model at 140.99 token/s average throughput, ahead of deepseek-v4-flash at 136.53 token/s and gemma4:31b at 113.93 token/s; nemotron-3-ultra is the weakest at 4.51 token/s average, never exceeding 6.99 token/s.
  • The most significant movement is glm-5.3's 28.6% upward trend, closing at 183.25 token/s, plus deepseek-v4-flash's spike to 223.43 token/s at 09:40; volatility is highest for glm-5.3 (33.9% CV) and gemma4:31b (33.3% CV), which swung between 42.58 and 162.02 token/s.
  • No missing-data limitation applies: all 8 models recorded 12 of 12 expected samples, totaling 96 valid points at 100.0% coverage over the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput is deepseek-v4-flash at 141.82 token/s (peaking at 223.43 token/s), while nemotron-3-ultra is weakest at 4.46 token/s, never exceeding 6.99 token/s across the window.
  • glm-5.3-flash shows the largest upward trend at 54.9%, climbing from roughly 102 token/s early in the window to a latest reading of 172.69 token/s; glm-5.3 is the most volatile model, with a 38.9% coefficient of variation and swings between 64.2 and 186.83 token/s.
  • No missing-data limitation exists: all eight models have 12 of 12 samples, yielding 96 valid points and 100.0% coverage for the four-hour period.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model by average throughput at 126.69 token/s, narrowly ahead of deepseek-v4-flash at 131.09 token/s... correction: deepseek-v4-flash leads at 131.09 token/s, with gemma4:31b second at 126.69 token/s; nemotron-3-ultra is weakest at 4.5 token/s.
  • glm-5.3 shows the sharpest volatility, with a coefficient of variation of 43.7% and swings between 64.2 and 186.83 token/s, including a drop from 186.83 to 73.98 token/s between 07:40 and 08:00; deepseek-v4-pro posts the largest gain, trending up 36.6% from a 35.11 token/s low to 105.41 token/s.
  • No missing-data limitation exists: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0% across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput was gemma4:31b at 136.76 token/s (p95 183.66 token/s); weakest was nemotron-3-ultra at 5.13 token/s, which never exceeded 7.36 token/s.
  • glm-5.3 showed the most operationally significant volatility: coefficient of variation 46.4% with a +61.1% trend, swinging from roughly 69 token/s early in the window to a 186.83 token/s maximum at 07:40 before falling to 73.98 token/s at 08:00.
  • No missing-data limitation applies: all 8 models delivered 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model by average throughput at 143.84 token/s, peaking at 184.99 token/s, while nemotron-3-ultra is the weakest at 5.08 token/s average and never exceeding 7.36 token/s.
  • glm-5.3 shows the most operationally significant volatility, with a 47.3% coefficient of variation and swings between 64.2 and 183.69 token/s, including a sustained dip to roughly 68–72 token/s from 04:00 to 05:40 UTC; deepseek-v4-pro also declined 35.9% overall, falling to 35.11 token/s at 06:20.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0%, so the four-hour window is complete.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput was gemma4:31b at 145.82 token/s (peak 184.99 token/s); weakest was nemotron-3-ultra at 5.36 token/s, never exceeding 7.36 token/s.
  • glm-5.3 showed the sharpest operational swing: after holding 174.24-183.69 token/s from 03:00 to 03:40, it fell to roughly 68-72 token/s for six consecutive samples before rebounding to 174.35 token/s at 06:00, a -37.4% trend with 45.5% coefficient of variation.
  • No missing-data limitation applies: all 8 models delivered 12 of 12 expected samples, valid_point_count is 96, and coverage is 100.0%, so the four-hour window is complete.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model at 147.15 token/s average throughput, narrowly ahead of gemma4:31b at 145.47 token/s; nemotron-3-ultra is the weakest at 5.75 token/s average, never exceeding 9.72 token/s.
  • glm-5.3 shows the most operationally significant volatility, with a coefficient of variation of 47.6%: throughput climbed from 65.03 token/s at 01:20 to 183.69 token/s at 03:20, then dropped to 71.47 token/s at 04:00. deepseek-v4-pro posted the largest trend change at +32.4%.
  • No missing-data limitation applies: all eight models have 12 of 12 expected samples, and the dataset reports 96 valid points with 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 144.77 token/s average throughput, with deepseek-v4-flash close behind at 142.96 token/s; nemotron-3-ultra is the weakest at 5.79 token/s average, peaking at only 9.72 token/s.
  • glm-5.3 shows the most operationally significant volatility, swinging between 191.82 and 65.03 token/s with a 44.6% coefficient of variation, including a drop from 179.48 to 71.47 token/s in the final 20-minute interval; deepseek-v4-pro also ranged from 33.26 to 122.53 token/s.
  • No missing-data limitation exists: all 8 models have 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput belongs to deepseek-v4-flash at 147.48 token/s (peaking at 200.56 token/s), narrowly ahead of gemma4:31b at 141.25 token/s; weakest is nemotron-3-ultra, averaging just 10.70 token/s with a maximum of only 37.32 token/s.
  • The most operationally significant volatility is nemotron-3-ultra, ranging 2.21 to 37.32 token/s with 86.4% CV and a 61.4% decline; deepseek-v4-pro is also erratic (8.86 to 109.27 token/s, 57.5% CV), and glm-5.2 dropped 24.6% to a 74.10 token/s average.
  • No missing-data limitation applies: all 8 models report 12 of 12 samples, totaling 96 valid points at 100.0% coverage for the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 141.62 token/s average throughput, peaking at 177.46 token/s, while nemotron-3-ultra is the weakest at 9.93 token/s average, never exceeding 37.32 token/s.
  • deepseek-v4-pro shows the most operationally significant volatility, dropping from a 107.66 token/s peak to an 8.86 token/s low with a 54.1% coefficient of variation and a -41.1% trend; glm-5.3-flash rose 54.9% to a 172.45 token/s peak.
  • No missing-data limitation exists: all 8 models have 12 of 12 samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput is deepseek-v4-flash at 136.76 token/s (p95 188.96, maximum 190.88 token/s); weakest is nemotron-3-ultra at 9.66 token/s, never exceeding 37.32 token/s. glm-5.3 is second strongest at 125.42 token/s.
  • The most operationally significant movement is deepseek-v4-pro's decline: roughly 100 token/s through 22:20, then a drop to 8.86 token/s at 23:20 and a 33.24 token/s close, a -53.0% trend with 47.4% CV. deepseek-v4-flash also dipped to 15.01 token/s at 21:40.
  • No missing-data limitation exists: all 8 models report 12 of 12 expected samples, 96 valid points total, and 100.0% coverage over the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model at 135.68 token/s average throughput (peak 190.88 token/s), while nemotron-3-ultra is the weakest at 5.65 token/s average, roughly 24x slower.
  • The most operationally significant movement is deepseek-v4-pro's late decline: it held 100-117 token/s from 19:20 through 22:20, then fell to 37.19 and 32.3 token/s in the final two intervals, a -28.2% trend. deepseek-v4-flash also swung sharply, ranging from 15.01 to 190.88 token/s.
  • No missing-data limitation applies: all eight models delivered 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput was deepseek-v4-flash at 130.51 token/s, while nemotron-3-ultra was weakest at 3.5 token/s, roughly 37 times slower.
  • The most significant trend is glm-5.3 climbing 67.7% to a 149.59 token/s peak, recovering from a 19.92 token/s low at 19:40; deepseek-v4-flash also swung sharply between 15.01 and 187.39 token/s.
  • No missing-data limitation exists: all eight models recorded 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model at 138.76 token/s average throughput, peaking at 168.96 token/s; nemotron-3-ultra is the weakest at 3.44 token/s average, never exceeding 5.8 token/s.
  • deepseek-v4-pro shows the largest sustained gain, rising 43.1% to a 110.26 token/s latest reading after hovering near 61-85 token/s early; glm-5.3 is the most volatile, swinging from 149.91 token/s at 17:40 down to 19.92 token/s at 19:40.
  • No missing-data limitation exists: all eight models report 12 of 12 samples, 96 valid points, and 100.0% coverage, though the window spans only four hours.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model at 135.62 token/s average throughput (peak 162.57 token/s), while nemotron-3-ultra is the weakest at 3.06 token/s average, roughly 44 times slower.
  • The most significant movement is in the glm-5.3 family: glm-5.3 fell 37.1% over the window, dropping to 19.92 token/s at 19:40, and glm-5.3-flash fell 35.5% from a 174.92 token/s peak; glm-5.3 also shows the highest volatility at 37.7% coefficient of variation.
  • No missing-data limitation exists: all eight models have 12 of 12 samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model at 131.15 token/s average throughput, ahead of glm-5.3 at 118.37 token/s, while nemotron-3-ultra is the weakest at 2.92 token/s average, roughly 45 times slower than the leader.
  • The most operationally significant movement is glm-5.3's decline from a 153.57 token/s peak at 16:40 to 61.74 token/s at 19:00, a -13.5% trend; deepseek-v4-flash moved the opposite way, up 20.3% to a 162.57 token/s peak at 17:00.
  • No missing-data limitation exists in this window: all eight models have 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour period.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model at 124.41 token/s average throughput, narrowly ahead of deepseek-v4-flash at 124.27 token/s; nemotron-3-ultra is the weakest at 2.87 token/s average, roughly 43 times slower than glm-5.3.
  • deepseek-v4-flash shows the widest volatility, ranging from 26.57 to 191.31 token/s with a 35.0% coefficient of variation, and gemma4:31b is similarly unstable at 37.9%; glm-5.3 is the steadiest performer at 19.3%.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0% across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model at 122.04 token/s average throughput, ahead of deepseek-v4-flash at 115.21 token/s; nemotron-3-ultra is the weakest at 2.76 token/s average, far below every other model.
  • The sharpest operational swing is glm-5.3-flash, trending down 40.9% from a 164.14 token/s peak to 78.28 token/s at the latest sample; deepseek-v4-flash also ranged widely between 26.57 and 191.31 token/s.
  • No missing-data limitation applies: all eight models report 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model at 123.22 token/s average throughput, while nemotron-3-ultra is the weakest at 3.46 token/s average, roughly 35 times slower.
  • gemma4:31b shows the sharpest decline, trending down 30.8% to 23.27 token/s at 16:00; deepseek-v4-flash swung from a 191.31 token/s peak at 14:40 to 26.57 token/s at 15:00, and minimax-m3 spiked to 105.28 token/s at 15:40.
  • No missing-data limitation applies: all eight models recorded 12 of 12 samples, the dataset holds 96 valid points, and coverage is 100.0% across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model at 124.96 token/s average throughput, narrowly ahead of glm-5.3 at 122.75 token/s; nemotron-3-ultra is the weakest at 5.09 token/s average, roughly 25 times slower than the leader.
  • minimax-m3 shows the most operationally significant volatility, with a coefficient of variation of 71.8 percent, a one-off spike to 168.33 token/s at 10:40 UTC, and a 35.9 percent downward trend ending at 43.52 token/s.
  • No missing-data limitation applies: all eight models report 12 of 12 samples, valid_point_count is 96, and coverage is 100.0 percent for the full four-hour window.
4-hour window · 96 points · 100.0% coverage
  • Strongest model by average throughput is gemma4:31b at 122.23 token/s, narrowly ahead of glm-5.3 at 115.73 token/s; weakest is nemotron-3-ultra at 5.78 token/s, with a maximum of only 8.37 token/s.
  • Most operationally significant volatility is minimax-m3: 69.3% coefficient of variation, a one-point spike to 168.33 token/s at 10:40 against a 51.57 token/s average, and a -37.2% trend; deepseek-v4-flash also swung between 10.65 and 157.94 token/s.
  • No missing-data limitation applies: all eight models report 12 of 12 samples, 96 valid points, and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model at 129.86 token/s average throughput, peaking at 213.47 token/s, while nemotron-3-ultra is the weakest at 6.69 token/s average, never exceeding 17.52 token/s.
  • The most operationally significant volatility is minimax-m3, which spiked to 168.33 token/s at 10:40 UTC against a 51.91 token/s average (68.7% coefficient of variation); deepseek-v4-flash also dipped to 10.65 token/s at 09:40 and glm-5.3 to 25.21 token/s at 10:00.
  • No missing-data limitation applies: all eight models recorded 12 of 12 expected samples, totaling 96 valid points with 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash leads average throughput at 137.78 token/s (peak 213.47 token/s), while nemotron-3-ultra is weakest at 8.36 token/s average, roughly 16x lower.
  • minimax-m3 shows the sharpest volatility: it fell from 189.19 token/s at 07:20 to a 33.49 token/s low at 11:00, with 72.8% coefficient of variation and a -40.8% trend; deepseek-v4-flash also dipped to 10.65 token/s at 09:40 before recovering.
  • No missing-data limitation: all eight models report 12 of 12 samples, 96 valid points, and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model by average throughput at 128.54 token/s, with a peak of 213.47 token/s; nemotron-3-ultra is the weakest at 8.66 token/s average, never exceeding 17.52 token/s.
  • The most operationally significant volatility is minimax-m3, which spiked from 36.89 token/s at 07:00 to 189.19 token/s at 07:20 and 181.68 token/s at 07:40, then collapsed to roughly 40-46 token/s from 08:20 onward, producing a -60.6% trend and 76.2% coefficient of variation.
  • No missing-data limitation applies: all eight models reported 12 of 12 expected samples, giving 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model at 144.58 token/s average throughput (peak 213.47 token/s), while nemotron-3-ultra is the weakest at 9.48 token/s average (peak only 17.52 token/s).
  • minimax-m3 shows the most operationally significant volatility, with a coefficient of variation of 74.3%: it jumped from 36.89 token/s at 07:00 to 189.19 token/s at 07:20, then fell back to 45.42 token/s by 08:20.
  • No missing-data limitation exists: all 8 models have 12 of 12 expected samples, with 96 valid points and 100.0% coverage over the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model at 144.66 token/s average throughput (peak 207.34 token/s), while nemotron-3-ultra is the weakest at 9.64 token/s average (minimum 2.0 token/s), roughly 15x lower.
  • The most significant volatility is minimax-m3, which held near 45-53 token/s for eight observations then jumped to 189.19 token/s at 07:20, driving a 72.0% coefficient of variation and a 126.8% trend; glm-5.3 also dipped to 25.72 token/s at 06:40.
  • No missing-data limitation applies: all eight models report 12 of 12 samples, 96 valid points, and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • Highest average throughput was deepseek-v4-flash at 138.7 token/s (peak 207.34 token/s), while nemotron-3-ultra was weakest at 9.21 token/s average, peaking at only 17.47 token/s.
  • gemma4:31b showed the widest volatility, coefficient of variation 46.3%, swinging between 37.48 and 175.35 token/s; glm-5.3 dipped to 25.72 token/s at 06:40 before recovering to 150.44 token/s at 07:00, and deepseek-v4-pro declined 16.2% overall, hitting 27.04 token/s at 06:00.
  • No missing-data limitation exists in this window: all eight models delivered 12 of 12 samples, with 96 valid points and 100.0% coverage.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model at 141.96 token/s average throughput (range 110.03–207.34 token/s), while nemotron-3-ultra is the weakest at 11.85 token/s average, peaking at only 21.8 token/s.
  • gemma4:31b shows the most volatility, with a 42.7% coefficient of variation and swings between 49.27 and 190.24 token/s; nemotron-3-ultra also declined 34.7% over the window, falling to 5.32 token/s at 06:00.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0%, so every twenty-minute observation in the four-hour window is present.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model at 144.63 token/s average throughput (peak 207.34 token/s), while nemotron-3-ultra is the weakest at 13.10 token/s average, peaking at only 22.22 token/s.
  • gemma4:31b shows the highest volatility, with a coefficient of variation of 38.9% and swings between 49.27 and 190.24 token/s, including a drop to 49.27 token/s at 03:40; deepseek-v4-pro is the only model trending upward, at +36.5%.
  • No missing-data limitation applies: all eight models recorded 12 of 12 samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput belongs to deepseek-v4-flash at 134.85 token/s (peak 185.27 token/s), while nemotron-3-ultra is weakest at 12.05 token/s, roughly one-eleventh of the leader's average.
  • gemma4:31b shows the sharpest operational volatility, with a 39.0% coefficient of variation and swings between 49.27 and 190.24 token/s within the window; glm-5.3-flash declined 11.1% overall, ending at 85.38 token/s.
  • No missing-data limitation applies: all eight models report 12 of 12 expected samples, 96 valid points, and 100.0% coverage, so the four-hour window is complete, though it captures only a single short period.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model by average throughput at 133.21 token/s, while nemotron-3-ultra is the weakest at 8.65 token/s, roughly fifteen times slower over the four-hour window.
  • The most significant movement is glm-5.3-flash, which trended up 76.6%, climbing from 58.61 token/s at 22:20 to a peak of 172.7 token/s at 01:00 before settling at 105.25 token/s; nemotron-3-ultra also shows high volatility with a 55.6% coefficient of variation.
  • No missing-data limitation applies: all eight models report 12 of 12 expected samples, yielding 96 valid points and 100.0% coverage, so every twenty-minute interval in the window is represented.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash posted the highest average throughput at 125.81 token/s, peaking at 157.39 token/s, while nemotron-3-ultra was weakest at 6.08 token/s average and never exceeded 12.34 token/s.
  • The sharpest volatility came from glm-5.3-flash, which jumped from 89.98 token/s at 00:40 to 172.7 token/s at 01:00, driving its +27.0% trend; deepseek-v4-flash trended down 18.4%, ending at 78.54 token/s.
  • No missing-data limitation applies: all eight models report 12 of 12 samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model on average throughput at 116.89 token/s, peaking at 157.39 token/s, while nemotron-3-ultra is the weakest at 5.28 token/s, roughly 22 times slower on average.
  • gemma4:31b shows the sharpest volatility, ranging from 66.21 to 190.44 token/s with a coefficient of variation of 36.2% and a -20.6% trend; in contrast, deepseek-v4-pro improved 42.4% overall, closing at 100.34 token/s.
  • No missing-data limitation applies: all eight models delivered 12 of 12 expected samples, yielding 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model at 123.17 token/s average throughput, ahead of glm-5.3 (112.16 token/s) and gemma4:31b (111.22 token/s); nemotron-3-ultra is the weakest at 3.73 token/s average, far below the next-lowest, minimax-m3, at 50.0 token/s.
  • gemma4:31b shows the highest volatility (CV 33.4%), swinging between 66.31 and 190.44 token/s, including a 190.44 token/s spike at 20:20 followed by a drop to 75.41 token/s at 20:40; deepseek-v4-flash similarly dipped to 41.3 token/s at 20:40 before recovering to 157.39 token/s at 22:20.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, 96 valid points, and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model at 118.58 token/s average throughput, peaking at 146.12 token/s; nemotron-3-ultra is the weakest at 3.03 token/s average, never exceeding 4.34 token/s.
  • The most operationally significant movement is glm-5.2's upward trend of 46.8%, climbing from 40.51 token/s at 18:20 to 103.1 token/s at 22:00, while deepseek-v4-pro shows the highest volatility (cv 42.2%), swinging between 9.18 and 103.28 token/s.
  • No missing-data limitation exists: all eight models report 12 of 12 expected samples, with 96 valid points and 100.0% coverage over the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model by average throughput at 114.11 token/s, ahead of deepseek-v4-flash at 107.66 token/s and gemma4:31b at 96.74 token/s; nemotron-3-ultra is weakest at 3.01 token/s, roughly 38 times slower than the leader.
  • Volatility is the main operational concern: deepseek-v4-pro swings from 9.18 to 103.28 token/s (47.6% coefficient of variation), and gemma4:31b ranges 29.03 to 190.44 token/s (44.7%), while deepseek-v4-flash dipped to 21.6 token/s at 19:00 before recovering.
  • No missing-data limitation exists: all eight models report 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model with an average throughput of 111.39 token/s, while nemotron-3-ultra is the weakest at 2.91 token/s; deepseek-v4-flash follows closely at 104.80 token/s.
  • deepseek-v4-pro shows the highest volatility, with a coefficient of variation of 55.9% and swings between 9.18 and 103.28 token/s; gemma4:31b also varies widely, from 29.03 to 154.43 token/s with a 41.3% coefficient of variation.
  • No missing-data limitation applies: all eight models report 12 of 12 samples, and the dataset records 96 valid points with 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model with an average throughput of 111.4 token/s (peak 152.09 token/s), while nemotron-3-ultra is the weakest at 2.7 token/s average, never exceeding 3.56 token/s.
  • deepseek-v4-flash shows the most operationally significant volatility, ranging from 21.6 to 180.81 token/s with a 42.1% coefficient of variation, and its latest reading fell to 21.6 token/s; deepseek-v4-pro is similarly erratic with a 52.6% coefficient of variation.
  • No missing-data limitation applies: all eight models report 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput was deepseek-v4-flash at 113.88 token/s; weakest was nemotron-3-ultra at 2.61 token/s, roughly 44 times lower.
  • deepseek-v4-flash also showed the most volatility, ranging from 22.77 to 180.81 token/s with a 46.3 token/s standard deviation, including a drop to 26.77 token/s at 16:20; deepseek-v4-pro fell to 9.84 token/s at 16:20.
  • No missing-data limitation applies: all eight models recorded 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model at 115.28 token/s average throughput, with deepseek-v4-flash close behind at 115.18 token/s; nemotron-3-ultra is the weakest at 2.36 token/s average, never exceeding 3.23 token/s.
  • The most operationally significant volatility is deepseek-v4-pro, whose throughput fell from 91.94 token/s at 14:40 to 9.84 token/s at 16:20, a 34.3% negative trend; deepseek-v4-flash also dipped sharply to 22.77 token/s at 14:40 and 26.77 token/s at 16:20.
  • No missing-data limitation exists: all 96 expected points are valid, coverage is 100.0%, and every model has 12 of 12 samples.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput was deepseek-v4-flash at 133.9 token/s, narrowly ahead of glm-5.3 at 122.41 token/s; weakest was nemotron-3-ultra at 2.42 token/s, far below every other model.
  • The most operationally significant volatility was at 14:40 UTC: minimax-m3 spiked to 134.78 token/s against a roughly 40 token/s baseline, while deepseek-v4-flash dropped to 22.77 token/s from 128.11 token/s at 14:20.
  • No missing-data limitation exists: all eight models reported 12 of 12 expected samples, giving 96 valid points and 100.0% coverage over the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model at 128.08 token/s average throughput, ahead of glm-5.3 at 121.15 token/s; nemotron-3-ultra is weakest at 2.66 token/s average, roughly 48x slower than the leader.
  • deepseek-v4-pro shows the most operationally significant volatility, swinging between 10.73 and 102.06 token/s with a 51.3% coefficient of variation, including drops to 11.76, 11.13, and 10.73 token/s at 12:00, 13:00, and 14:00 UTC.
  • No missing-data limitation exists: all eight models report 12 of 12 samples with 100.0% coverage, matching the 96 valid points expected for the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model at 135.47 token/s average throughput (peak 203.38 token/s), while nemotron-3-ultra is the weakest at 3.38 token/s average, peaking at only 6.07 token/s.
  • glm-5.3 shows the most significant upward trend, rising 49.2% to a 232.92 token/s peak at 13:40, while minimax-m3 declined 48.0% to 40.4 token/s; deepseek-v4-pro also showed sharp volatility, dropping to 10.73 token/s at 14:00.
  • No missing-data limitation: all eight models have 12 of 12 samples, 96 valid points, and 100.0% coverage over the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput was deepseek-v4-flash at 126.38 token/s (peak 203.38 token/s); weakest was nemotron-3-ultra at 3.81 token/s, never exceeding 9.24 token/s.
  • Most significant volatility: minimax-m3 (cv 62.6%) swung between 173.21 and 34.86 token/s and after 11:20 never exceeded 45.07 token/s; deepseek-v4-pro also fell to 11.13 token/s at 13:00, while gemma4:31b trended down 27.9% to 79.15 token/s.
  • No missing-data limitation: all 8 models delivered 12 of 12 samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model at 122.73 token/s average throughput, peaking at 203.38 token/s, while nemotron-3-ultra is the weakest at 3.65 token/s average, never exceeding 9.24 token/s.
  • minimax-m3 shows the most operationally significant volatility, with a 60.1% coefficient of variation and swings between 33.72 and 173.21 token/s; deepseek-v4-pro also dropped sharply to 11.76 token/s at 12:00 from 99.94 token/s at 11:40.
  • No missing-data limitation applies: all eight models report 12 of 12 samples, and the dataset shows 96 valid points with 100.0% coverage over the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model by average throughput at 121.52 token/s, edging deepseek-v4-flash at 117.77 token/s; nemotron-3-ultra is weakest at 3.02 token/s, roughly 40 times slower than the leader.
  • minimax-m3 shows the most operationally significant volatility, swinging between 30.65 and 133.01 token/s with a coefficient of variation of 60.4%, alternating near-flat readings around 33-36 token/s with spikes above 100 token/s at 09:40, 10:00, and 10:40 UTC.
  • No missing-data limitation exists: all eight models delivered 12 of 12 expected samples, and the dataset reports 96 valid points with 100.0% coverage over the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model by average throughput at 118.79 token/s, narrowly ahead of deepseek-v4-flash at 117.11 token/s, while nemotron-3-ultra is the weakest at 2.34 token/s, with a maximum of only 9.24 token/s across the window.
  • The most operationally significant movement is minimax-m3's late surge: it held roughly 30-39 token/s from 06:20 to 08:40, then jumped to 98.85, 133.01, and 107.97 token/s in the final three observations, a 114.9% trend. nemotron-3-ultra also shows extreme volatility, with a 102.1% coefficient of variation and a low of 0.49 token/s at 08:20.
  • No missing-data limitation applies: all eight models have 12 of 12 expected samples, and the dataset reports 96 valid points with 100.0% coverage, so the four-hour window is complete.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model at 120.16 token/s average throughput (peaking at 152.17 token/s), narrowly ahead of glm-5.3 at 118.66 token/s and gemma4:31b at 117.75 token/s; nemotron-3-ultra is the weakest at 2.79 token/s average, roughly 43x slower.
  • The most operationally significant movement is nemotron-3-ultra's sustained collapse: it fell from 5.26 token/s at 05:20 to a 0.49 token/s low at 08:20, a -66.1% trend, before recovering to 4.58 token/s at 09:00. glm-5.3 also swung between 186.03 and 76.49 token/s.
  • No missing-data limitation applies: all eight models delivered 12 of 12 expected samples, giving 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 88 points · 91.7% coverage
  • glm-5.3 is the strongest model at 131.05 token/s average throughput (peak 186.03 token/s), while nemotron-3-ultra is the weakest at 4.54 token/s average, never exceeding 10.46 token/s.
  • The most operationally significant movement is nemotron-3-ultra's steady decline from 6.21 token/s at 04:20 to 1.17 token/s at 07:40, a 63.3% downward trend. Also notable are single-interval drops: deepseek-v4-pro fell to 9.58 token/s at 06:00 and gemma4:31b to 27.97 token/s at 05:00, before both recovered.
  • Every model recorded 11 of 12 expected samples (91.7% coverage), yielding 88 valid points overall, so one observation per model is missing and short-window averages may underrepresent the full four hours.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model at 127.99 token/s average throughput (peak 169.43 token/s), while nemotron-3-ultra is the weakest at 4.87 token/s average, never exceeding 10.46 token/s.
  • The most operationally significant volatility is gemma4:31b, which swung between 27.97 and 173.46 token/s (CV 42.4%); deepseek-v4-pro also dropped to 9.58 token/s at 06:00 UTC before recovering to 108.17 token/s, and glm-5.3-flash spiked to 166.43 token/s at 07:00.
  • No missing-data limitation exists: all eight models have 12 of 12 samples, 96 valid points, and 100.0% coverage across the four-hour window.