← Performance dashboard

Hourly performance insights

Summaries of rolling four-hour performance data

1192 retained summaries
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 122.49 token/s average throughput, ahead of glm-5.3 at 102.07 token/s; nemotron-3-ultra is the weakest at 27.83 token/s average, never exceeding 55.94 token/s.
  • deepseek-v4-flash shows the steepest decline at -40.1% trend, and deepseek-v4-pro dropped to 11.76 token/s at 07:00; glm-5.3 and deepseek-v4-flash share the highest volatility at 51.9% coefficient of variation.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0%.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 123.66 token/s average throughput, while nemotron-3-ultra is the weakest at 26.74 token/s average, roughly a 4.6x gap.
  • glm-5.3 shows the most operationally significant volatility, swinging between 183.18 token/s at 05:20 and 16.28 token/s at 06:00 with a 46.6% coefficient of variation; deepseek-v4-flash is similarly erratic at 53.1%.
  • No missing-data limitation exists: all 8 models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0% across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 121.22 token/s average throughput, peaking at 147.35 token/s; nemotron-3-ultra is the weakest at 27.43 token/s average, never exceeding 55.94 token/s.
  • deepseek-v4-flash shows the highest volatility (CV 46.3%), swinging between 22.59 and 128.82 token/s, while glm-5.3 alternates sharply between roughly 58-79 and 147-171 token/s and ends at its maximum, 171.17 token/s.
  • No missing-data limitation applies: all eight models have 12 of 12 samples and 100.0% coverage, with 96 valid points overall.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model by average throughput at 113.2 token/s, narrowly ahead of gemma4:31b at 112.22 token/s, while nemotron-3-ultra is the weakest at 25.3 token/s average, peaking at only 50.73 token/s.
  • The most operationally significant volatility is glm-5.3-flash, with a coefficient of variation of 46.4% and swings from 52.72 to 192.31 token/s; glm-5.3 also swung between 55.26 and 181.11 token/s, and deepseek-v4-flash ended at 22.59 token/s after a -27.2% trend.
  • No missing-data limitation exists: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0% across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 113.08 token/s average throughput, peaking at 165.05 token/s; nemotron-3-ultra is the weakest at 25.44 token/s average, never exceeding 50.73 token/s.
  • glm-5.3-flash shows the steepest decline, dropping 38.0% from 171.63 token/s at 22:20 to 65.43 token/s at 02:00, while glm-5.3 is the most volatile with a 44.6% coefficient of variation, swinging between 181.11 and 55.26 token/s within the window.
  • No missing-data limitation applies: all eight models report 12 of 12 expected samples, and overall coverage is 100.0% across 96 valid points, so the four-hour picture is complete.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 107.85 token/s average throughput, while nemotron-3-ultra is the weakest at 21.78 token/s average, roughly one-fifth of the leader.
  • glm-5.3 shows the most operationally significant volatility, with a coefficient of variation of 43.5% and swings from 60.47 token/s at 22:00 up to 181.11 token/s at 00:40, then a drop to 55.26 token/s at 01:00; deepseek-v4-pro also dipped to 10.5 token/s at 23:00.
  • No missing-data limitation exists: all 8 models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0% across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model with an average throughput of 112.08 token/s, ahead of glm-5.3-flash at 107.43 token/s and glm-5.3 at 100.5 token/s; nemotron-3-ultra is the weakest at 22.48 token/s average, peaking at only 50.73 token/s.
  • The most significant volatility is deepseek-v4-pro, which swung from a 10.5 token/s low at 23:00 to 126.0 token/s at 21:40, while glm-5.3-flash shows the highest coefficient of variation at 41.7%; deepseek-v4-flash posted the steepest upward trend at 34.4%.
  • No missing-data limitation exists: all eight models have 12 of 12 samples, and overall coverage is 100.0% across 96 valid points.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b delivered the highest average throughput at 110.03 token/s, peaking at 165.05 token/s, while nemotron-3-ultra was weakest at 18.57 token/s average, never exceeding 30.12 token/s.
  • The most operationally significant volatility is deepseek-v4-pro's collapse to 10.5 token/s at 23:00 after reaching 126.0 token/s at 21:40; glm-5.3-flash also swung sharply (cv 45.4%), ranging 57.72 to 181.6 token/s.
  • No missing-data limitation exists: all eight models report 12 of 12 samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 110.75 token/s average throughput, while nemotron-3-ultra is weakest at 17.42 token/s average, roughly six times slower.
  • glm-5.3-flash shows the sharpest volatility, spiking from 53.99 token/s at 19:00 to 181.6 token/s at 20:40 before falling back to 58.43 token/s at 21:00; its coefficient of variation is 47.3%.
  • No missing-data limitation applies: all 8 models delivered 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput belongs to gemma4:31b at 115.24 token/s (peak 154.14 token/s), while nemotron-3-ultra is weakest at 17.40 token/s, never exceeding 30.12 token/s.
  • The sharpest operational volatility is glm-5.3-flash, which jumped from 59.72 token/s at 20:00 to 181.60 token/s at 20:40, then fell to 58.43 token/s at 21:00; glm-5.2 also swung between 7.07 and 87.64 token/s with 58.0% coefficient of variation.
  • No missing-data limitation exists: all eight models recorded 12 of 12 expected samples, totaling 96 valid points at 100.0% coverage, so the four-hour window is fully represented.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput was gemma4:31b at 106.21 token/s (peaking at 151.86 token/s), while nemotron-3-ultra was weakest at 16.57 token/s, never exceeding 33.72 token/s.
  • deepseek-v4-pro showed the most operationally significant volatility, with a 59.2% coefficient of variation and swings from 5.68 to 95.96 token/s; glm-5.2 was similarly unstable at 54.5%, dropping to 7.07 token/s at 19:00 UTC.
  • No missing-data limitation applies: all eight models recorded 12 of 12 expected samples, with 96 valid points and 100.0% coverage, so results reflect complete twenty-minute observations across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model with an average throughput of 97.33 token/s (peak 146.95 token/s), while nemotron-3-ultra is the weakest at 15.79 token/s average, never exceeding 33.72 token/s across the window.
  • Volatility is the dominant operational signal: deepseek-v4-pro swung from 106.62 token/s at 15:20 to 5.68 token/s at 19:00 (cv 56.6%), and glm-5.3 ranged between 160.6 and 55.89 token/s, indicating severe throughput instability on several models.
  • No missing-data limitation applies: all 8 models delivered 12 of 12 expected samples, with 96 valid points and 100.0% coverage, so the four-hour window is fully populated despite sharp single-interval drops.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 97.43 token/s average throughput, ahead of glm-5.3 at 92.47 token/s; nemotron-3-ultra is the weakest at 16.28 token/s, never exceeding 33.72 token/s in any sample.
  • deepseek-v4-flash shows the steepest trend, up 54.6% to a latest 64.8 token/s, while deepseek-v4-pro is the most volatile, with a 54.9% coefficient of variation, swings between 9.11 and 106.62 token/s, and a 17.9% decline.
  • No missing-data limitation applies: all eight models have 12 of 12 expected samples, totaling 96 valid points at 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput was glm-5.3 at 102.37 token/s (peaking at 160.60 token/s at 16:40); weakest was nemotron-3-ultra at 16.33 token/s, ending at 1.83 token/s at 17:00.
  • Most significant volatility: glm-5.3-flash dropped from 156.48 token/s at 13:20 to 6.19 token/s at 14:40, with a 50.8% coefficient of variation and a -31.7% trend; gemma4:31b also declined -29.1%, closing at 38.99 token/s.
  • No missing-data limitation applies: all eight models recorded 12 of 12 samples, yielding 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput was gemma4:31b at 112.33 token/s (peak 159.45 token/s); weakest was nemotron-3-ultra at 14.22 token/s, never exceeding 22.69 token/s.
  • Most operationally significant volatility: glm-5.3-flash (cv 48.6%) fell to 6.19 token/s at 14:40 after peaking at 156.48 token/s at 13:20, and gemma4:31b declined 28.0% overall, closing at 59.99 token/s.
  • No missing data: all eight models report 12 of 12 samples, 96 valid points, and 100.0% coverage; the only limitation is that the window spans just four hours of observations.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b delivered the highest average throughput at 120.24 token/s (peak 159.45 token/s), while nemotron-3-ultra was weakest at 14.02 token/s average, peaking at only 22.69 token/s.
  • glm-5.3-flash showed the sharpest volatility, ranging from 6.19 to 156.48 token/s with a 49.2% coefficient of variation, and several models dropped sharply in the final observation: glm-5.2 fell to 8.61 token/s and deepseek-v4-pro to 9.89 token/s at 15:00 UTC.
  • No missing-data limitation applies: all eight models recorded 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • Strongest average output throughput was gemma4:31b at 130.05 token/s (peak 159.45 token/s), while nemotron-3-ultra was weakest at 17.49 token/s, never exceeding 74.74 token/s.
  • The most operationally significant volatility was nemotron-3-ultra, with a 106.7% coefficient of variation, a -50.4% trend, and a fall from 74.74 token/s at 10:20 to 13.37 token/s at 13:00; deepseek-v4-flash also swung between 128.56 and 27.43 token/s.
  • No missing-data limitation exists: all 8 models delivered 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 119.16 token/s average throughput, ahead of glm-5.3 at 103.05 token/s, while nemotron-3-ultra is the weakest at 22.5 token/s average, never exceeding 74.74 token/s in any interval.
  • nemotron-3-ultra shows the most extreme volatility, with a coefficient of variation of 88.5% and swings from 2.46 token/s at 10:00 to 74.74 token/s at 10:20; minimax-m3 also spiked to 169.37 token/s at 11:00 against a 73.64 token/s average.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0% across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model by average throughput at 115.36 token/s, ahead of glm-5.3 at 100.75 token/s, while nemotron-3-ultra is the weakest at 21.18 token/s, with a minimum of just 2.46 token/s.
  • The most operationally significant volatility comes from nemotron-3-ultra, whose coefficient of variation is 96.5%, swinging between 2.46 and 74.74 token/s; minimax-m3 also spiked to 169.37 token/s at 11:00 after averaging 64.19 token/s.
  • No missing-data limitation applies: all eight models have 12 of 12 expected samples, valid_point_count is 96, and coverage is 100.0% for the full 07:03 to 11:03 UTC window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 110.67 token/s average throughput, peaking at 158.15 token/s, while nemotron-3-ultra is the weakest at 16.61 token/s average, never exceeding 45.55 token/s.
  • deepseek-v4-flash shows the sharpest trend, climbing 74.7% from a 30.8 token/s low to 128.56 token/s at 09:40, and deepseek-v4-pro is the most volatile among large models, swinging between 4.58 and 111.1 token/s (cv 51.9%).
  • No missing-data limitation: all 8 models have 12 of 12 samples, 96 valid points, 100% coverage.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput belongs to gemma4:31b at 109.02 token/s (peak 158.15 token/s), while nemotron-3-ultra is weakest at 15.80 token/s average, never exceeding 45.55 token/s.
  • The most operationally significant volatility is nemotron-3-ultra's 78.9% coefficient of variation, alongside deepseek-v4-pro collapsing from 111.10 token/s at 07:20 to 4.58 token/s at 09:00; deepseek-v4-flash shows the steepest decline at -30.9%.
  • No missing-data limitation exists: all eight models report 12 of 12 samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput was gemma4:31b at 111.84 token/s (peaking at 158.15 token/s); weakest was nemotron-3-ultra at 17.12 token/s, which peaked at only 64.11 token/s and ended at 7.27 token/s.
  • Every model declined over the window; deepseek-v4-flash fell hardest at -54.3%, dropping from 118.59 to 32.55 token/s, while glm-5.2 fell -35.2%. nemotron-3-ultra showed the greatest volatility with a 90.4% coefficient of variation across a 5.09 to 64.11 token/s range.
  • No missing-data limitation applies: all eight models recorded 12 of 12 expected samples, giving 96 valid points and 100.0% coverage for the four-hour period.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 124.92 token/s average throughput, while nemotron-3-ultra is the weakest at 24.53 token/s average, roughly a fifth of the leader's pace.
  • The most operationally significant volatility is nemotron-3-ultra, whose throughput fell 65.4% over the window (coefficient of variation 81.5%), dipping to 5.09 token/s at 06:00 UTC; glm-5.3-flash also swung between 187.08 and 53.74 token/s.
  • No missing-data limitation exists: all 8 models report 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model by average throughput at 117.4 token/s (p95 170.54 token/s); nemotron-3-ultra is the weakest at 24.53 token/s average, never exceeding 68.35 token/s.
  • nemotron-3-ultra shows the most operationally significant volatility, ranging 3.91 to 68.35 token/s with 85.9% CV and ending near 5 token/s; deepseek-v4-flash posted the sharpest rise (+80.2% trend, 30.44 to 118.59 token/s), while glm-5.3 declined 25.1%.
  • No missing-data limitation exists: every model delivered 12 of 12 samples, giving 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 posted the strongest average throughput at 129.1 token/s (peaking at 179.12 token/s), while nemotron-3-ultra was weakest at 32.65 token/s average, dipping as low as 3.91 token/s at 02:20.
  • The most significant trend is glm-5.2's 37.9% decline, from 100.34 token/s at 00:20 to 43.0 token/s at 04:00; deepseek-v4-flash fell 35.9% to 51.66 token/s. nemotron-3-ultra showed the highest volatility, with a 61.1% coefficient of variation.
  • No missing-data limitation applies: all eight models recorded 12 of 12 expected samples, and the dataset reports 96 valid points with 100.0% coverage.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model by average throughput at 125.11 token/s, narrowly ahead of gemma4:31b at 122.71 token/s; nemotron-3-ultra is the weakest at 34.33 token/s average, with a low of 3.91 token/s at 02:20.
  • The most operationally significant volatility is nemotron-3-ultra's coefficient of variation of 78.9%, swinging between 3.91 and 102.82 token/s; deepseek-v4-flash shows the steepest decline at -45.9%, falling from 119.47 token/s at 23:40 to 54.18 token/s at 03:00.
  • No missing-data limitation exists: all eight models have 12 of 12 samples, and valid_point_count is 96 with 100.0% coverage.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 128.06 token/s average throughput, while nemotron-3-ultra is the weakest at 37.38 token/s, roughly 29% of gemma4:31b's average.
  • glm-5.3 shows the most significant movement, rising 103.7% over the window to a latest 179.12 token/s, but with 53.3% coefficient of variation; minimax-m3 fell 34.3% to 47.53 token/s, and nemotron-3-ultra's 76.0% coefficient of variation marks it as the most volatile.
  • No missing-data limitation exists: all eight models have 12 of 12 samples and 100.0% coverage, with 96 valid points; the only constraint is the four-hour window itself.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput was gemma4:31b at 133.9 token/s (p95 159.85, max 160.99 token/s); weakest was nemotron-3-ultra at 31.85 token/s, ranging from 2.62 to 102.82 token/s across the window.
  • The most operationally significant volatility is nemotron-3-ultra's 93.5% coefficient of variation, swinging between 2.62 and 102.82 token/s; glm-5.3 shows a 63.9% upward trend, closing at 172.22 token/s, while glm-5.3-flash declined 35.4% from its 193.02 token/s peak.
  • No missing-data limitation applies: all eight models recorded 12 of 12 expected samples, totaling 96 valid points at 100.0% coverage for the four-hour period.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 133.33 token/s average throughput, while nemotron-3-ultra is the weakest at 24.38 token/s, roughly 5.5x lower.
  • nemotron-3-ultra shows the most volatility, ranging from 2.62 to 102.82 token/s with a 121.0% coefficient of variation; glm-5.3 has the steepest decline at -34.1% trend, falling from a 141.89 token/s peak to 69.56 token/s.
  • No missing-data limitation exists: all eight models have 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 127.66 token/s average throughput (peak 161.53 token/s), while nemotron-3-ultra is the weakest at 14.26 token/s average, never exceeding 70.12 token/s in any observation.
  • The most operationally significant volatility is nemotron-3-ultra's coefficient of variation of 126.1%, swinging between 2.62 and 70.12 token/s; glm-5.3-flash and deepseek-v4-flash show the steepest upward trends at 51.9% and 51.3%.
  • No missing-data limitation exists: all 8 models report 12 of 12 samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 130.49 token/s average throughput, while nemotron-3-ultra is the weakest at 11.83 token/s, roughly a tenth of the leader's pace.
  • The most operationally significant volatility is nemotron-3-ultra's 105.9% coefficient of variation, swinging between 2.62 and 48.14 token/s. minimax-m3 also jumped from 74.10 token/s at 20:40 to 170.35 token/s at 21:00, and glm-5.3-flash spiked to 193.02 token/s at 21:20.
  • There is no missing-data limitation: all eight models report 12 of 12 samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model by average throughput at 131.19 token/s (peaking at 161.53 token/s), while nemotron-3-ultra is the weakest at 12.51 token/s average, never exceeding 48.14 token/s and ending at 5.21 token/s.
  • The most operationally significant movement is minimax-m3's sustained decline: from 183.32 token/s at 17:20 it fell to a 56.41 token/s low at 20:00, holding near 70-76 token/s for hours before recovering to 170.35 token/s at 21:00, with 43.3% coefficient of variation and an 18.6% negative trend.
  • No missing-data limitation exists in this window: all eight models delivered 12 of 12 expected samples, and the dataset reports 96 valid points with 100.0% coverage.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput was gemma4:31b at 123.86 token/s; weakest was nemotron-3-ultra at 14.49 token/s, roughly 8.5x lower, with a latest reading of 8.63 token/s.
  • The most operationally significant volatility is nemotron-3-ultra's 98.7% coefficient of variation, swinging between 2.98 and 48.14 token/s; deepseek-v4-flash shows the steepest trend at +100.2%, while minimax-m3 declined 31.9% to a latest 56.41 token/s.
  • No missing-data limitation applies: 96 of 96 expected points are valid, coverage is 100.0%, and every model reports 12 of 12 samples across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 115.74 token/s average throughput, ahead of minimax-m3 at 93.62 token/s; nemotron-3-ultra is the weakest at 14.49 token/s average, with a low of 2.55 token/s.
  • glm-5.3-flash held near 65 token/s for most of the window then jumped to 174.61 token/s at 19:00, its maximum; glm-5.3 showed the widest swings, ranging 15.28 to 148.30 token/s with 58.1% coefficient of variation.
  • No missing-data limitation: all 8 models delivered 12 of 12 expected samples, giving 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput was minimax-m3 at 108.59 token/s (peak 183.32 token/s); weakest was nemotron-3-ultra at 16.23 token/s, with a maximum of only 47.12 token/s.
  • The most operationally significant volatility appears in nemotron-3-ultra (cv 86.9%) and glm-5.2 (cv 66.5%); deepseek-v4-pro swung between 112.57 token/s at 16:40 and 4.33 token/s at 16:00, ending at 5.23 token/s.
  • No missing-data limitation exists in this window: all eight models recorded 12 of 12 expected samples, giving 96 valid points and 100.0% coverage.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 103.55 token/s average throughput, ahead of minimax-m3 at 98.72 token/s; nemotron-3-ultra is the weakest at 15.96 token/s average, with a maximum of only 47.12 token/s.
  • The most significant movement is late-window degradation: minimax-m3 fell 31.7% to 56.41 token/s at 17:00 and glm-5.3 fell 48.9% to 15.28 token/s; glm-5.2 was the most volatile, ranging 13.06 to 173.59 token/s with an 84.7% coefficient of variation.
  • No missing-data limitation applies: all 8 models recorded 12 of 12 expected samples, giving 96 valid points and 100.0% coverage over the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 107.84 token/s average throughput, narrowly ahead of minimax-m3 at 106.76 token/s; nemotron-3-ultra is the weakest at 16.53 token/s average, with a maximum of only 51.54 token/s.
  • glm-5.2 shows the most extreme volatility, ranging from 8.92 to 173.59 token/s with a 96.9% coefficient of variation, including a spike to 173.59 token/s at 15:00 UTC; deepseek-v4-pro also fell to 4.33 token/s at 16:00 UTC.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0% across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput was gemma4:31b at 113.26 token/s (peak 153.26 token/s at 12:40), while nemotron-3-ultra was weakest at 25.61 token/s, never exceeding 55.49 token/s in any interval.
  • glm-5.2 showed the most extreme volatility, with a coefficient of variation of 115.0%: it averaged 47.79 token/s but spiked to 173.59 token/s at 15:00 after sitting near 13-26 token/s for most of the window; nemotron-3-ultra also collapsed to 1.78 token/s at 13:00.
  • No missing-data limitation applies: all eight models delivered 12 of 12 expected samples, giving 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model by average throughput at 120.59 token/s, while nemotron-3-ultra is the weakest at 19.93 token/s average.
  • glm-5.2 shows the most operationally significant trend, declining 61.7% over the window to a latest 13.06 token/s, with extreme volatility (cv 100.6%, range 8.92 to 161.74 token/s); deepseek-v4-pro also fell to 7.63 token/s at 14:00.
  • No missing-data limitation exists in this window: all eight models have 12 of 12 samples, 96 valid points, and 100% coverage.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model with an average throughput of 108.88 token/s (range 55.13–164.94 token/s), while nemotron-3-ultra is the weakest at 32.75 token/s average (range 2.56–72.75 token/s).
  • The most operationally significant movement is nemotron-3-ultra's declining trend of -41.5%, including a low of 2.56 token/s at 11:00 UTC. glm-5.2 shows the highest volatility (cv 70.3%), swinging from 12.4 token/s at 11:40 to 161.74 token/s at 12:00.
  • No missing-data limitation applies: all eight models have 12 of 12 expected samples, valid_point_count is 96, and coverage is 100.0% across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 114.5 token/s average throughput, well ahead of second-place glm-5.3 at 90.24 token/s; nemotron-3-ultra is the weakest at 30.84 token/s average, with a minimum of 2.56 token/s.
  • The most operationally significant movement is nemotron-3-ultra's decline of 43.5 percent, falling from a 72.75 token/s peak to 2.56 token/s at 11:00 UTC; glm-5.3-flash also shows high volatility, with a 58.7 percent coefficient of variation and a dip to 9.94 token/s at 08:40.
  • No missing-data limitation applies: all 8 models have 12 of 12 samples, and the dataset reports 96 valid points with 100.0 percent coverage.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput was gemma4:31b at 120.08 token/s, peaking at 164.94 token/s; weakest was nemotron-3-ultra at 32.14 token/s, dipping as low as 2.98 token/s.
  • Volatility is the main operational concern: glm-5.3-flash swung between 9.94 and 163.78 token/s (57.7% CV) and nemotron-3-ultra showed 69.9% CV; at 10:00 UTC deepseek-v4-pro fell to 9.21 token/s and glm-5.3-flash to 14.79 token/s.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput was gemma4:31b at 130.59 token/s (peak 164.94 token/s), while nemotron-3-ultra was weakest at 27.76 token/s (minimum 2.98 token/s).
  • minimax-m3 showed the steepest decline, trending -35.9% and falling from 87.86 token/s at 05:20 to 60.56 token/s at 09:00, with a low of 27.03 token/s at 08:40; nemotron-3-ultra was most volatile with a 79.2% coefficient of variation.
  • No missing-data limitation applies: all eight models delivered 12 of 12 expected samples, giving 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput was gemma4:31b at 138.29 token/s (peak 162.26 token/s); weakest was nemotron-3-ultra at 27.15 token/s, with a minimum of just 1.93 token/s.
  • The most operationally significant movement is minimax-m3's 18.9% decline, ending at 46.88 token/s, alongside glm-5.3's 25.5% rise to 123.32 token/s; nemotron-3-ultra showed extreme volatility with a coefficient of variation of 89.7% and a 1.93 to 77.99 token/s range.
  • No missing-data limitation applies: all eight models recorded 12 of 12 samples, 96 valid points total, and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput is gemma4:31b at 134.21 token/s across 12 samples (range 99.74 to 156.96 token/s); weakest is nemotron-3-ultra at 22.89 token/s, which peaked at only 77.99 token/s and ended the window at 4.51 token/s.
  • The most operationally significant volatility is nemotron-3-ultra, swinging between 1.93 and 77.99 token/s with a 96.0% coefficient of variation; deepseek-v4-pro shows the steepest decline at -31.1% trend, hitting 10.99 token/s at 06:00, and glm-5.3-flash is erratic at 53.4% CV with a 194.31 token/s spike.
  • No missing-data limitation applies: all eight models report 12 of 12 expected samples, 96 valid points, and 100.0% coverage over the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model with average throughput of 136.02 token/s (p95 160.99 token/s), while nemotron-3-ultra is the weakest at 32.72 token/s average, never exceeding 77.99 token/s.
  • The most operationally significant volatility is in glm-5.3-flash, which swings between 45.94 and 194.31 token/s with a 55.5% coefficient of variation; glm-5.3 also fell to 5.72 token/s at 05:20 and deepseek-v4-pro dropped to 10.99 token/s at 06:00.
  • No missing-data limitation exists: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0% across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model by average throughput at 132.43 token/s, while nemotron-3-ultra is the weakest at 33.89 token/s; deepseek-v4-flash and glm-5.2 also sit low at 56.95 and 59.02 token/s respectively.
  • The most operationally significant volatility is nemotron-3-ultra, whose throughput swings between 1.93 and 77.99 token/s with a 74.4% coefficient of variation; minimax-m3 shows the clearest upward trend at 31.4%, rising from 44.23 to 88.4 token/s, while deepseek-v4-pro declines 25.5%.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, 96 valid points total, and 100.0% coverage across the four-hour window.
4-hour window · 88 points · 91.7% coverage
  • Strongest average throughput belongs to gemma4:31b at 130.97 token/s (peak 167.44 token/s), while nemotron-3-ultra is weakest at 34.27 token/s, never exceeding 66.28 token/s.
  • glm-5.3-flash shows the widest swings, ranging 57.4 to 188.92 token/s with 42.8% coefficient of variation, and deepseek-v4-flash declined 30.0% over the window, ending at 30.73 token/s after reaching 120.54 token/s at 01:00.
  • Each of the eight models recorded 11 of 12 expected samples, giving 88 valid points and 91.7% coverage, so one 20-minute observation is missing per model, limiting trend reliability.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput was gemma4:31b at 127.28 token/s (peak 167.44 token/s); weakest was nemotron-3-ultra at 37.37 token/s, with a low of 3.87 token/s.
  • Nemotron-3-ultra showed the most volatility, with a coefficient of variation of 86.5% and swings between 3.87 and 112.61 token/s; glm-5.3 fell from 158.15 token/s at 23:20 to 18.43 token/s at 02:00, a 31.6% negative trend.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, 96 valid points, and 100.0% coverage, though the window covers only four hours.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 124.79 token/s average throughput; nemotron-3-ultra is the weakest at 37.77 token/s average, with a minimum of 3.19 token/s.
  • glm-5.3 shows the steepest decline, trending -32.4% from 142.52 token/s at 21:20 to 92.24 token/s at 01:00, while nemotron-3-ultra is the most volatile (cv 81.9%, range 3.19–112.61 token/s); deepseek-v4-flash rose 27.5% to 120.54 token/s.
  • No missing-data limitation: all 8 models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0%.