← Performance dashboard

Hourly performance insights

Summaries of rolling four-hour performance data

1195 retained summaries
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash posted the strongest average throughput at 120.09 token/s, ahead of gemma4:31b at 100.31 token/s, while nemotron-3-ultra was weakest at 24.03 token/s, never exceeding 52.53 token/s in any interval.
  • glm-5.3-flash showed the most operationally significant volatility, with a coefficient of variation of 52.7% and swings from 43.26 token/s at 18:00 to 183.59 token/s at 17:00 and 182.29 token/s at 18:40; glm-5.3 also declined 30.0% overall, ending at 57.67 token/s.
  • No missing-data limitation applies: all eight models recorded 12 of 12 expected samples, totaling 96 valid points at 100.0% coverage for the window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model with an average throughput of 108.31 token/s (peaking at 153.96 token/s), while nemotron-3-ultra is the weakest at 26.38 token/s average, never exceeding 69.26 token/s.
  • The most operationally significant volatility comes from nemotron-3-ultra, whose coefficient of variation is 80.5% with swings between 4.77 and 69.26 token/s; deepseek-v4-flash shows the strongest upward trend at 45.1%, ending at 132.22 token/s.
  • No missing-data limitation exists in this window: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0%, so every twenty-minute observation from 15:20 to 19:00 UTC is present.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 101.65 token/s average throughput; nemotron-3-ultra is the weakest at 24.06 token/s, roughly a quarter of gemma4:31b's pace.
  • nemotron-3-ultra shows the most operationally significant volatility, with a coefficient of variation of 87.1%, swings between 4.77 and 69.26 token/s, and a -41.7% trend ending at 4.81 token/s; deepseek-v4-flash trended up 63.4%, peaking at 153.96 token/s.
  • No missing-data limitation applies: all eight models delivered 12 of 12 expected samples, 96 valid points, and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3-flash is the strongest model at 102.26 token/s average throughput, peaking at 183.59 token/s; nemotron-3-ultra is the weakest at 22.22 token/s average, with a maximum of only 69.26 token/s.
  • glm-5.3 shows the steepest upward trend at +92.2%, climbing from 60.13 token/s at 13:20 to 154.61 token/s at 16:40, while nemotron-3-ultra is the most volatile with a 91.8% coefficient of variation, swinging between 2.23 and 69.26 token/s.
  • No missing-data limitation exists: all 8 models delivered 12 of 12 expected samples, giving 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3-flash is the strongest model at 102.91 token/s average throughput, ahead of deepseek-v4-flash at 83.31 token/s; nemotron-3-ultra is the weakest at 26.21 token/s average, with a maximum of only 69.26 token/s.
  • The most operationally significant volatility is nemotron-3-ultra's 80.2% coefficient of variation, swinging between 2.23 and 69.26 token/s; gemma4:31b shows the steepest trend, rising 95.3% to a 152.34 token/s peak at 15:20 before dropping to 63.13 token/s.
  • No missing-data limitation applies: all eight models report 12 of 12 samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3-flash is the strongest model at 101.2 token/s average throughput, ahead of deepseek-v4-flash at 87.52 token/s; nemotron-3-ultra is the weakest at 27.27 token/s average, with a minimum of 2.23 token/s.
  • glm-5.3 shows the most operationally significant volatility, swinging between 9.35 and 134.99 token/s (cv 54.8%), including a drop to 9.35 token/s at 14:20 followed by 134.99 token/s at 14:40; nemotron-3-ultra is also erratic with cv 69.3%.
  • There is no missing-data limitation: every model has 12 of 12 samples, and the dataset contains 96 valid points at 100.0% coverage, so the only constraint is the four-hour window itself.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3-flash is the strongest model at 101.97 token/s average throughput, while nemotron-3-ultra is the weakest at 32.99 token/s average.
  • The most operationally significant movement is nemotron-3-ultra's decline of 49.9 percent, from 70.05 token/s at 10:20 to 21.32 token/s at 14:00, with high volatility (cv 62.2 percent); deepseek-v4-pro also fell 36.5 percent to 32.11 token/s, while glm-5.3-flash rose 52.7 percent to 107.31 token/s.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, and valid_point_count is 96 with 100.0 percent coverage.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3-flash is the strongest model at 101.02 token/s average throughput, with a peak of 188.73 token/s; nemotron-3-ultra is the weakest at 39.53 token/s average, never exceeding 71.72 token/s.
  • Volatility is the dominant operational signal: glm-5.3-flash swings between 15.54 and 188.73 token/s (54.5% CV), glm-5.2 drops to 2.45 token/s at 12:00 UTC, and seven of eight models trend downward, with deepseek-v4-pro the only gainer at +21.2%.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, 96 valid points, and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput was glm-5.3 at 95.12 token/s; weakest was nemotron-3-ultra at 52.13 token/s, with glm-5.3-flash second at 92.65 token/s and minimax-m3 the steadiest at 78.78 token/s (CV 16.4%).
  • glm-5.3-flash was the most volatile model, spanning 15.54 to 188.73 token/s (CV 55.9%), and the 12:00 UTC window saw sharp drops: glm-5.2 at 2.45 token/s and glm-5.3-flash at 15.54 token/s, while deepseek-v4-pro trended up 10.6% to 104.79 token/s.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0%, so the four-hour window is fully represented.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3-flash is the strongest model at 103.53 token/s average throughput, peaking at 188.73 token/s; nemotron-3-ultra is the weakest at 50.97 token/s average, with a low of 13.37 token/s.
  • Volatility is the dominant operational concern: nemotron-3-ultra swings between 13.37 and 105.2 token/s (55.1% CV), deepseek-v4-flash between 36.93 and 146.11 token/s (47.3% CV), and glm-5.3-flash between 52.87 and 188.73 token/s despite its high average.
  • No missing-data limitation applies: all 8 models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0% across the 07:20 to 11:00 UTC observations.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3-flash is the strongest model at 100.78 token/s average throughput, ahead of gemma4:31b at 90.95 token/s; nemotron-3-ultra is the weakest at 44.91 token/s average, with a minimum of just 6.15 token/s.
  • Volatility is the dominant operational concern: nemotron-3-ultra shows a 73.1% coefficient of variation and deepseek-v4-flash 51.1%, with deepseek-v4-flash swinging from 146.11 token/s at 09:20 to 36.93 token/s at 09:40.
  • No missing-data limitation applies in this window: all 8 models have 12 of 12 samples, 96 valid points, and 100.0% coverage.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3-flash is the strongest model by average throughput at 88.37 token/s, while nemotron-3-ultra is the weakest at 57.49 token/s; gemma4:31b follows closely at 86.15 token/s.
  • Volatility is the main operational concern: nemotron-3-ultra swings between 6.15 and 111.63 token/s (CV 66.9%), and glm-5.3-flash ranges 16.04 to 174.3 token/s. Most models declined over the window, with deepseek-v4-flash down 20.1% and minimax-m3 down 14.3%; only glm-5.3-flash rose, up 11.1%.
  • No missing-data limitation exists: all 8 models delivered 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3-flash is the strongest model at 93.76 token/s average throughput, narrowly ahead of deepseek-v4-flash at 93.05 token/s; nemotron-3-ultra is the weakest at 49.5 token/s average.
  • The most operationally significant movement is nemotron-3-ultra's 57.6 percent decline, from a 111.63 token/s peak at 06:00 to 13.37 token/s at 08:00, including lows of 6.15 and 7.15 token/s near 06:40–07:00; glm-5.3-flash is the most volatile, ranging 16.04 to 174.3 token/s with a 51.8 percent coefficient of variation.
  • No missing-data limitation applies: all 8 models delivered 12 of 12 expected samples, and the dataset records 96 valid points with 100.0 percent coverage.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model at 99.9 token/s average throughput, edging glm-5.3-flash (91.87 token/s); nemotron-3-ultra is weakest at 48.64 token/s average, roughly half the leader.
  • glm-5.3-flash shows the widest volatility, with a 54.6% coefficient of variation and swings from 16.04 to 174.3 token/s; nemotron-3-ultra climbed 72.6% overall to a 111.63 token/s peak at 06:00 before collapsing to 7.15 token/s by 07:00.
  • No missing-data limitation: all eight models delivered 12 of 12 expected samples, and the dataset reports 96 valid points with 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model by average throughput at 101.62 token/s, while nemotron-3-ultra is the weakest at 49.68 token/s; the other six models average between 74.46 and 89.09 token/s.
  • The most operationally significant movement is nemotron-3-ultra's climb from 15.95 token/s at 04:00 to 111.63 token/s at 06:00, a 132.9% trend. Volatility is highest for nemotron-3-ultra (CV 58.8%) and glm-5.3-flash (54.0%), whose throughput swings between 16.04 and 174.3 token/s.
  • No missing-data limitation applies: all 8 models have 12 of 12 samples, 96 valid points, and 100.0% coverage, so every twenty-minute observation in the window is present.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput was gemma4:31b at 101.43 token/s; weakest was nemotron-3-ultra at 39.12 token/s, roughly 2.6 times lower.
  • gemma4:31b showed the widest volatility, swinging between 30.3 and 184.45 token/s with a 46.6% coefficient of variation, while glm-5.3 posted the sharpest trend, up 46.3% over the window, closing at 139.83 token/s after dipping to 36.23 token/s at 04:40.
  • No missing-data limitation applies: all eight models reported 12 of 12 samples, yielding 96 valid points and 100.0% coverage across the four-hour period.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 120.66 token/s average throughput, while nemotron-3-ultra is the weakest at 34.11 token/s; deepseek-v4-flash follows at 99.6 token/s and glm-5.2 trails near the bottom at 69.46 token/s.
  • Volatility is the dominant operational signal: gemma4:31b swung from a 184.45 token/s peak at 01:40 to 72.95 token/s at 02:00 and a 30.3 token/s low at 03:00, and glm-5.3-flash ranged 37.97–167.8 token/s with a 46.7% coefficient of variation.
  • No missing-data limitation exists: all 8 models delivered 12 of 12 samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 129.64 token/s average throughput, more than double the fleet median; nemotron-3-ultra is the weakest at 40.36 token/s average, never exceeding 84.43 token/s.
  • glm-5.3 shows the sharpest decline, trending down 43.7% from 140.41 to 70.98 token/s, while nemotron-3-ultra and glm-5.3-flash are the most volatile with coefficient-of-variation values of 56.0% and 50.6%; deepseek-v4-flash swung between 25.14 and 137.51 token/s.
  • No missing-data limitation exists: all 8 models delivered 12 of 12 expected samples, yielding 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3-flash is the strongest model at 104.87 token/s average throughput, edging glm-5.3 at 102.0 token/s; nemotron-3-ultra is weakest at 34.45 token/s average, peaking at only 84.43 token/s.
  • The most operationally significant movement is the synchronized decline in the glm family: glm-5.2 fell 30.0% and glm-5.3 fell 29.8% over the window, with latest readings of 25.45 and 51.77 token/s; nemotron-3-ultra is highly volatile (cv 76.1%), swinging from 8.69 to 84.43 token/s.
  • No missing-data limitation: all 8 models report 12 of 12 samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b delivered the strongest average throughput at 126.83 token/s, followed by glm-5.3 at 111.09 token/s; nemotron-3-ultra was weakest at 22.82 token/s average, peaking at only 81.64 token/s.
  • The most operationally significant volatility came from nemotron-3-ultra, whose coefficient of variation was 102.3% with a 409.3% trend, jumping from 4.88 token/s at 22:00 to 81.64 token/s at 23:40; glm-5.3-flash also swung between 21.78 and 158.36 token/s (cv 51.1%).
  • No missing-data limitation exists in this window: all eight models recorded 12 of 12 expected samples, totaling 96 valid points at 100.0% coverage.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model by average throughput at 122.26 token/s, while nemotron-3-ultra is the weakest at 18.59 token/s; gemma4:31b is second at 114.05 token/s.
  • The most operationally significant change is nemotron-3-ultra's late surge: it ran at 4.88-12.12 token/s through 23:20, then jumped to 81.64 token/s at 23:40 and 60.45 token/s at 00:00, driving its 129.0% coefficient of variation. glm-5.3-flash is also volatile, ranging 21.78-157.0 token/s.
  • No missing-data limitation applies: all 8 models have 12 of 12 samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model with an average throughput of 120.37 token/s (peak 167.3 token/s), while nemotron-3-ultra is the weakest at 10.15 token/s average, peaking at only 23.06 token/s.
  • The most significant shift is glm-5.3-flash, which trended up 66.9% over the window, climbing from a low of 21.78 token/s at 21:40 to 150.01 token/s at 23:00; nemotron-3-ultra trended down 41.5% over the same period.
  • No missing-data limitation applies: all 8 models delivered 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model at 124.9 token/s average throughput, ahead of gemma4:31b at 105.55 token/s; nemotron-3-ultra is the weakest at 9.22 token/s average, never exceeding 23.06 token/s.
  • The sharpest operational movement is nemotron-3-ultra's 43.4% decline, from 23.06 token/s at 20:00 to 4.88 token/s at 22:00; minimax-m3 also swung from a 168.39 token/s peak at 20:00 down to 69.66 token/s at 21:40, a 19.3% trend drop.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, 96 valid points overall, and 100.0% coverage, so no conclusions are constrained by gaps.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model at 126.47 token/s average throughput, peaking at 167.30 token/s; nemotron-3-ultra is the weakest at 11.42 token/s average, never exceeding 32.62 token/s.
  • Volatility is the dominant operational signal: gemma4:31b swings between 27.25 and 160.34 token/s with 39.5% CV, nemotron-3-ultra shows 74.4% CV, and minimax-m3 posts the largest trend change at 31.8%, while glm-5.3 remains steadiest at 18.6% CV.
  • No missing-data limitation exists: all eight models report 12 of 12 expected samples, 96 valid points, and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model by average throughput at 106.39 token/s, narrowly ahead of gemma4:31b at 105.43 token/s, while nemotron-3-ultra is the weakest at 13.59 token/s average, never exceeding 32.62 token/s.
  • The most operationally significant movement is minimax-m3's 57.7% upward trend, rising from 67.96 token/s at 16:20 to 168.39 token/s at 20:00; gemma4:31b shows the highest volatility (cv 42.3%), swinging between 27.25 and 160.34 token/s.
  • No missing-data limitation applies: all eight models report 12 of 12 expected samples, totaling 96 valid points with 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model at 98.47 token/s average throughput, peaking at 144.55 token/s, while nemotron-3-ultra is the weakest at 12.14 token/s average and never exceeding 32.62 token/s.
  • Volatility is the dominant operational signal: gemma4:31b swings between 27.25 and 146.37 token/s (cv 47.8%), deepseek-v4-flash ranges 26.31 to 133.55 token/s (cv 52.1%), and nemotron-3-ultra shows the highest relative volatility at cv 75.7%.
  • No missing-data limitation exists in this window: all eight models report 12 of 12 expected samples, with 96 valid points and 100.0% coverage; the only constraint is that the dataset spans just four hours.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model at 89.0 token/s average throughput, ahead of gemma4:31b at 83.15 token/s; nemotron-3-ultra is weakest at 18.19 token/s average, peaking at only 45.96 token/s.
  • nemotron-3-ultra shows the highest volatility (CV 76.3%), oscillating between 3.44 and 45.96 token/s and ending at 6.34 token/s; deepseek-v4-flash has the steepest trend at +84.3% but swings from 26.31 to 133.55 token/s.
  • No missing-data limitation: all 8 models delivered 12 of 12 expected samples, with 96 valid points and 100.0% coverage over the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model at 90.6 token/s average throughput, peaking at 136.75 token/s; nemotron-3-ultra is weakest at 19.49 token/s average, never exceeding 45.96 token/s.
  • deepseek-v4-pro shows the sharpest operational volatility, dropping from 116.96 token/s at 14:20 to 8.02 token/s at 15:00 and again to 6.8 token/s at 16:00; nemotron-3-ultra has the highest coefficient of variation at 66.2%.
  • No missing-data limitation: all 8 models report 12 of 12 samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model by average throughput at 89.64 token/s, while nemotron-3-ultra is the weakest at 24.87 token/s; the spread between them is roughly 64.77 token/s.
  • The most operationally significant movement is deepseek-v4-flash, whose throughput fell 55.5% over the window, from a 143.01 token/s peak at 12:40 to 32.51 token/s at 16:00. deepseek-v4-pro also shows sharp swings, dipping to 6.5 and 6.8 token/s with a 60.3% coefficient of variation.
  • No missing-data limitation applies: all 8 models delivered 12 of 12 expected samples, giving 96 valid points and 100.0% coverage across the four-hour period.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model at 102.04 token/s average throughput, ahead of runner-up glm-5.3-flash at 78.45 token/s; nemotron-3-ultra is the weakest at 28.70 token/s average, peaking at only 59.33 token/s.
  • deepseek-v4-pro shows the sharpest volatility, swinging between 6.5 and 116.96 token/s (48.9% coefficient of variation), with lows of 6.5 token/s at 13:00 UTC and 8.02 token/s at 15:00 UTC; nemotron-3-ultra is also unstable at 61.3% CV with a -24.3% trend.
  • No missing-data limitation: all eight models have 12 of 12 samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model with an average throughput of 99.81 token/s (peaking at 137.96 token/s), while nemotron-3-ultra is the weakest at 33.69 token/s on average, never exceeding 75.87 token/s.
  • deepseek-v4-flash shows the highest volatility, with a coefficient of variation of 45.8% and swings between 35.86 and 143.01 token/s; deepseek-v4-pro also dipped sharply to 6.5 token/s at 13:00 UTC before recovering to 84.71 token/s by 14:00.
  • No missing-data limitation applies: all 8 models delivered 12 of 12 expected samples, giving 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 delivered the strongest average throughput at 89.79 token/s, while nemotron-3-ultra was weakest at 32.04 token/s; gemma4:31b followed closely behind glm-5.3 at 85.95 token/s.
  • The most significant volatility came from deepseek-v4-pro (CV 49.5%), which ranged from 6.5 to 121.1 token/s and closed at just 6.5 token/s; glm-5.3 also swung from 65.18 token/s at 10:00 to 137.96 token/s at 11:20, then back to 63.38 token/s at 12:40.
  • No missing-data limitation applies: all eight models recorded 12 of 12 expected samples, totaling 96 valid points at 100.0% coverage.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model with an average throughput of 100.92 token/s, while nemotron-3-ultra is the weakest at 34.15 token/s; gemma4:31b follows at 90.43 token/s and minimax-m3 at 83.99 token/s.
  • deepseek-v4-pro shows the most operationally significant volatility, ranging from 8.33 to 121.1 token/s with a coefficient of variation of 53.3% and an 82.9% upward trend; nemotron-3-ultra is even less stable at 73.6% CV, dipping to 3.26 token/s.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, and the dataset records 96 valid points with 100.0% coverage over the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b delivered the strongest average throughput at 99.67 token/s, while nemotron-3-ultra was weakest at 34.15 token/s; glm-5.3-flash (52.05 token/s) and deepseek-v4-flash (63.91 token/s) also trailed the fleet average.
  • The most operationally significant volatility came from nemotron-3-ultra (coefficient of variation 67.5%, range 3.26 to 76.24 token/s) and deepseek-v4-pro (53.4%, range 8.33 to 121.1 token/s), while glm-5.3 showed the steepest decline at -26.2% trend.
  • No missing-data limitation exists: all eight models recorded 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 88 points · 91.7% coverage
  • glm-5.3 is the strongest model at 107.56 token/s average throughput, while nemotron-3-ultra is the weakest at 28.08 token/s average, with a latest reading of only 3.26 token/s.
  • The most significant volatility is deepseek-v4-pro, whose coefficient of variation is 55.8%, ranging from 8.33 to 107.36 token/s and trending down 26.4%; glm-5.3-flash shows a similar 27.0% decline, including a dip to 4.95 token/s at 08:00.
  • Every model has 11 of 12 expected samples (91.7% coverage), leaving 88 valid points of 96, so each model is missing one observation.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model with an average throughput of 111.08 token/s (peak 169.24 token/s), while nemotron-3-ultra is the weakest at 32.32 token/s average, ranging from 5.07 to 76.24 token/s.
  • The most operationally significant volatility is nemotron-3-ultra's coefficient of variation of 67.0%, swinging between 5.07 and 76.24 token/s; deepseek-v4-pro also fell sharply from 107.36 token/s at 08:40 to 11.0 token/s at 09:00, its lowest reading of the window.
  • No missing-data limitation applies: all eight models delivered 12 of 12 expected samples, giving 96 valid points and 100.0% coverage across the four-hour period.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model at 107.02 token/s average throughput, peaking at 169.24 token/s; nemotron-3-ultra is the weakest at 38.62 token/s average, with a low of 5.07 token/s.
  • The most operationally significant volatility is nemotron-3-ultra's 65.4% coefficient of variation and -49.0% trend, while glm-5.3-flash dropped to 4.95 token/s at 08:00 against a 135.07 token/s peak; deepseek-v4-flash also swung between 137.88 and 30.21 token/s.
  • No missing-data limitation applies: all 8 models delivered 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model with an average throughput of 114.24 token/s (peak 175.78 token/s), while nemotron-3-ultra is the weakest at 36.38 token/s average, peaking at only 83.24 token/s.
  • The most significant trend is deepseek-v4-flash's decline of 39.4%, falling from 135.67 token/s at 03:20 to 45.17 token/s at 07:00; nemotron-3-ultra also shows extreme volatility with a 74.7% coefficient of variation, swinging between 4.21 and 83.24 token/s.
  • No missing-data limitation applies: all eight models have 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model at 109.99 token/s average throughput, while nemotron-3-ultra is the weakest at 41.06 token/s average, with a minimum of just 4.21 token/s.
  • nemotron-3-ultra shows the most volatility, with a coefficient of variation of 80.3% and swings between 4.21 and 102.83 token/s, plus a 65.2% upward trend; glm-5.3-flash is also unstable (45.2% CV) and declined 29.8%.
  • No missing-data limitation: all 8 models delivered 12 of 12 expected samples, giving 96 valid points and 100.0% coverage over the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model at 122.79 token/s average throughput, peaking at 175.78 token/s, while nemotron-3-ultra is the weakest at 42.53 token/s average, with a low of 4.21 token/s.
  • nemotron-3-ultra shows the most operationally significant volatility, with a coefficient of variation of 79.7% and swings between 4.21 and 102.83 token/s; gemma4:31b, glm-5.3, and both deepseek variants all recorded their window minimums at 03:00 UTC.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, and the dataset reports 96 valid points with 100.0% coverage over the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 delivered the highest average throughput at 116.37 token/s, narrowly ahead of gemma4:31b at 114.0 token/s, while nemotron-3-ultra was weakest at 34.68 token/s, peaking at only 102.83 token/s.
  • deepseek-v4-pro showed the sharpest decline, trending -36.6% and closing at 18.79 token/s after a 20.75 token/s dip at 02:20; nemotron-3-ultra was most volatile with a coefficient of variation of 83.1%, swinging between 7.54 and 102.83 token/s.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, 96 valid points, and 100.0% coverage over the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput was gemma4:31b at 120.89 token/s; weakest was nemotron-3-ultra at 24.48 token/s, roughly one-fifth of the leader's average.
  • The most operationally significant volatility came from nemotron-3-ultra, which swung between 1.0 and 80.97 token/s with a 92.8% coefficient of variation; glm-5.3 showed the strongest upward trend at +61.8%, closing at 138.11 token/s.
  • No missing-data limitation exists in this window: all eight models delivered 12 of 12 expected samples, 96 valid points overall, and 100.0% coverage, so no throughput conclusions are affected by gaps.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b delivered the highest average throughput at 121.31 token/s, peaking at 179.27 token/s, while nemotron-3-ultra was weakest at 15.99 token/s average and never exceeded 46.52 token/s.
  • The most operationally significant volatility came from nemotron-3-ultra, with a coefficient of variation of 105.1%, swinging from 1.0 token/s at 22:40 to 46.52 token/s at 00:20; glm-5.3-flash showed the steepest decline at -14.7% trend.
  • No missing-data limitation applies: all eight models recorded 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 113.88 token/s average throughput, while nemotron-3-ultra is the weakest at 9.11 token/s average, with a minimum of 1.0 token/s and a maximum of only 40.16 token/s.
  • The most operationally significant volatility is nemotron-3-ultra, whose coefficient of variation is 147.6% and whose throughput rose from 1.37 token/s at 20:20 to 40.16 token/s at 23:40; deepseek-v4-pro shows the largest positive trend among the other models at 30.4%.
  • No missing-data limitation exists: all eight models have 12 of 12 expected samples, valid_point_count is 96, and coverage is 100.0%.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput is gemma4:31b at 113.03 token/s (range 68.57 to 163.35 token/s); weakest is nemotron-3-ultra at 2.13 token/s, peaking at only 4.8 token/s, roughly 53 times below the leader.
  • glm-5.3-flash shows the widest volatility, with a 40.0% coefficient of variation and swings between 56.29 and 168.33 token/s, plus a -28.1% trend despite finishing at 142.51 token/s; deepseek-v4-pro also dipped to 23.77 token/s at 22:00.
  • No missing data: all eight models report 12 of 12 samples, 96 valid points, and 100.0% coverage; the only limitation is the four-hour window, which cannot confirm whether these patterns persist.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput is glm-5.3 at 106.36 token/s, followed by glm-5.3-flash at 102.65 token/s; weakest is nemotron-3-ultra at 2.42 token/s, far below the next-lowest glm-5.2 at 70.48 token/s.
  • The most operationally significant volatility is glm-5.3-flash, swinging between 56.29 and 168.33 token/s with a 41.9% coefficient of variation, while deepseek-v4-flash shows the steepest upward trend at +47%, ending near 95.75 token/s after peaking at 123.44 token/s.
  • No missing-data limitation exists: all eight models recorded 12 of 12 expected samples, 96 valid points, and 100% coverage; the only constraint is the four-hour window itself.
4-hour window · 88 points · 91.7% coverage
  • glm-5.3-flash is the strongest model at 103.33 token/s average throughput, narrowly ahead of glm-5.3 at 103.06 token/s; nemotron-3-ultra is the weakest at 2.76 token/s average, with a maximum of only 5.29 token/s.
  • The most operationally significant volatility is glm-5.3-flash, whose throughput swings between 55.56 and 168.33 token/s (coefficient of variation 43.9%), while nemotron-3-ultra shows the steepest decline at -32.0% trend, ending at 1.84 token/s.
  • Each of the eight models has 11 of 12 expected samples (91.7% coverage, 88 valid points), so one twenty-minute observation is missing per model, which limits interval-level comparisons.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model by average throughput at 94.85 token/s (peak 149.70 token/s), while nemotron-3-ultra is the weakest at 2.65 token/s average, never exceeding 5.29 token/s.
  • Volatility is the dominant operational signal: gemma4:31b swung from 142.49 token/s at 17:20 to 19.97 token/s at 18:40, and minimax-m3 fell from a 164.79 token/s peak to 26.38 token/s at 20:00; deepseek-v4-pro shows the widest relative spread (cv 49.6%).
  • No missing-data limitation applies: all 8 models delivered 12 of 12 expected samples, with 96 valid points and 100.0% coverage, so the four-hour window is fully represented.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3-flash is the strongest model by average throughput at 93.51 token/s, ahead of glm-5.3 at 89.6 token/s and minimax-m3 at 86.2 token/s; nemotron-3-ultra is the weakest at 2.6 token/s, far below the next-lowest, deepseek-v4-pro at 50.53 token/s.
  • The most operationally significant movement is minimax-m3's decline, trending -37.6% from a 164.79 token/s peak at 16:40 to 72.8 token/s at 19:00; deepseek-v4-pro shows the widest volatility, with a 52.7% coefficient of variation and swings between 8.89 and 86.06 token/s.
  • No missing-data limitation applies: all eight models report 12 of 12 expected samples, 96 valid points, and 100.0% coverage over the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3-flash is the strongest model at 93.51 token/s average throughput, ahead of minimax-m3 (87.80 token/s) and gemma4:31b (87.16 token/s); nemotron-3-ultra is the weakest at 1.96 token/s average, never exceeding 3.72 token/s.
  • deepseek-v4-pro shows the most operationally significant volatility, with a 57.9% coefficient of variation, a peak of 92.28 token/s at 14:20 UTC, and a collapse to 8.89 token/s at 18:00 UTC; glm-5.3-flash also declined 39.0% across the window.
  • No missing-data limitation applies: all eight models recorded 12 of 12 samples, 96 valid points total, and 100.0% coverage over the four-hour window.