← Performance dashboard

Hourly performance insights

Summaries of rolling four-hour performance data

1194 retained summaries
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model with an average throughput of 146.23 token/s, peaking at 186.3 token/s at 22:00 UTC, while nemotron-3-ultra is the weakest at 35.57 token/s average, never exceeding 94.35 token/s.
  • The most operationally significant volatility comes from nemotron-3-ultra, whose coefficient of variation is 86.1%, swinging between 6.13 and 94.35 token/s and trending up 155.9% over the window; gemma4:31b also fell from a 173.46 token/s peak to 55.4 token/s at 22:00, a 23.4% decline.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0%, so the four-hour window is fully populated.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput was gemma4:31b at 146.45 token/s (peak 173.46 token/s); weakest was nemotron-3-ultra at 19.53 token/s, with a maximum of only 44.67 token/s.
  • deepseek-v4-flash showed the sharpest upward trend at +38.1%, ranging from 64.58 to 187.41 token/s, while glm-5.2 declined 30.0% to 55.5 token/s; nemotron-3-ultra was the most volatile, with a 73.1% coefficient of variation.
  • No missing-data limitation applies: all eight models report 12 of 12 samples at 100.0% coverage, and the 96 valid points match the expected total, so the four-hour window is fully complete.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 143.42 token/s average throughput, ahead of deepseek-v4-flash at 133.39 token/s; nemotron-3-ultra is the weakest at 19.52 token/s average, with a maximum of only 44.67 token/s.
  • glm-5.3 shows the largest trend, rising 52.9% to a 179.38 token/s maximum, while glm-5.3-flash is the most volatile (cv 40.7%), swinging between 39.43 and 193.35 token/s; deepseek-v4-pro fell 31.0% with dips to 22.90 token/s.
  • No missing data: all 96 expected points are present at 100.0% coverage, though the four-hour window with 12 samples per model limits longer-term conclusions.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 131.73 token/s average throughput, ahead of deepseek-v4-flash at 116.37 token/s; nemotron-3-ultra is the weakest at 20.07 token/s average, never exceeding 55.82 token/s in any interval.
  • glm-5.3 shows the most operationally significant volatility, ranging from 16.0 to 179.38 token/s with a 54.3% coefficient of variation and an 80.7% upward trend, while nemotron-3-ultra is even less stable at 89.7% CV, swinging between 2.77 and 55.82 token/s.
  • No missing-data limitation applies: all eight models delivered 12 of 12 expected samples, yielding 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 124.47 token/s average throughput, ahead of deepseek-v4-flash at 119.17 token/s; nemotron-3-ultra is the weakest at 17.91 token/s average, with a maximum of only 55.82 token/s.
  • The most operationally significant volatility is nemotron-3-ultra, whose coefficient of variation is 92.2%, swinging between 2.57 and 55.82 token/s, including readings of 2.77 token/s at 17:00 and 6.59 token/s at 18:00; glm-5.3 also shows instability with a 47.2% coefficient of variation and a dip to 16.0 token/s at 15:20.
  • No missing-data limitation applies: all 8 models have 12 of 12 expected samples, valid_point_count is 96, and coverage is 100.0% across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b posted the highest average throughput at 118.22 token/s (peak 174.17 token/s), while nemotron-3-ultra was weakest at 22.58 token/s average, never exceeding 61.83 token/s and ending at 2.77 token/s.
  • The most significant volatility came from glm-5.3, whose throughput fell 53.7% over the window, from a 194.37 token/s peak at 13:40 to a low of 16.0 token/s at 15:20, with a 51.3% coefficient of variation; nemotron-3-ultra was similarly erratic at 86.7%.
  • No missing-data limitation applies: all eight models delivered 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput was deepseek-v4-flash at 116.0 token/s (range 77.9 to 148.26 token/s), while nemotron-3-ultra was weakest at 23.4 token/s average, peaking at only 61.83 token/s and ending at 5.46 token/s.
  • The most operationally significant volatility came from nemotron-3-ultra, whose coefficient of variation reached 89.1% with swings between 2.57 and 61.83 token/s, and glm-5.3, which fell 24.1% over the window, including a drop to 16.0 token/s at 15:20 UTC.
  • No missing-data limitation applies: all eight models recorded 12 of 12 expected samples, and the dataset reports 96 valid points with 100.0% coverage, so every twenty-minute interval in the four-hour window is represented.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput belongs to gemma4:31b at 118.16 token/s (peak 179.47 token/s), with deepseek-v4-flash close behind at 114.69 token/s; weakest is nemotron-3-ultra at 22.59 token/s, peaking at only 61.83 token/s and ending at 2.57 token/s.
  • The most operationally significant volatility is nemotron-3-ultra, swinging between 2.57 and 61.83 token/s with a cv of 88.0%; glm-5.3 shows the steepest trend, up 63.9% from 64.21 token/s to a 194.37 token/s peak, while gemma4:31b is also volatile at cv 41.1%.
  • No missing-data limitation exists: all eight models report 12 of 12 samples, 96 valid points total, and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b posted the highest average throughput at 119.17 token/s, while nemotron-3-ultra was weakest at 29.12 token/s; deepseek-v4-flash followed at 112.96 token/s.
  • glm-5.3-flash deteriorated most, trending down 39.9% from 189.98 token/s at 10:20 to 74.57 token/s at 14:00 with a 50.9% coefficient of variation; glm-5.3 trended up 25.8%, and nemotron-3-ultra swung between 3.43 and 61.83 token/s (77.8% CV).
  • No missing-data limitation applies: all 8 models returned 12 of 12 expected samples, totaling 96 valid points at 100.0% coverage over the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput was gemma4:31b at 118.2 token/s (p95 168.37 token/s), while nemotron-3-ultra was weakest at 24.86 token/s, never exceeding 56.28 token/s and dipping as low as 3.43 token/s.
  • glm-5.3-flash showed the widest swings, ranging 34.79 to 189.98 token/s with a 52.5% coefficient of variation and a -36.0% trend, closing at its 34.79 token/s minimum; nemotron-3-ultra was less stable still at 83.0% CV, with repeated collapses near 4 token/s around 11:00–12:20 UTC.
  • No missing-data limitation applies: all eight models reported 12 of 12 expected samples, yielding 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput was gemma4:31b at 117.59 token/s (peak 186.58 token/s), while nemotron-3-ultra was weakest at 27.7 token/s, dipping to 4.13 token/s at 12:00.
  • glm-5.3-flash showed the sharpest volatility, swinging between 54.06 and 189.98 token/s with a 45.5% coefficient of variation and a +48.6% trend; glm-5.2 declined 21.6% over the window.
  • No missing-data limitation applies: all 8 models delivered 12 of 12 expected samples, giving 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput is gemma4:31b at 115.98 token/s, ahead of deepseek-v4-flash at 115.05 token/s; weakest is nemotron-3-ultra at 28.10 token/s, roughly a quarter of the leader's average.
  • The most significant movement is glm-5.3-flash's upward trend of +34.1%, ending at 182.00 token/s after swinging between 54.06 and 189.98 token/s; gemma4:31b is highly volatile (CV 42.9%, range 24.95 to 186.58 token/s), and nemotron-3-ultra dropped to 4.28 token/s at 11:00.
  • No missing-data limitation applies: all eight models recorded 12 of 12 expected samples, giving 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput was deepseek-v4-flash at 115.78 token/s, edging gemma4:31b at 107.57 token/s and deepseek-v4-pro at 104.25 token/s; weakest was nemotron-3-ultra at 29.50 token/s, well below glm-5.3-flash at 78.27 token/s.
  • The most operationally significant volatility came from gemma4:31b, which swung between 21.39 and 186.58 token/s with a 52.7% coefficient of variation, including drops to 24.95 token/s at 08:00; nemotron-3-ultra was also erratic, dipping to 2.37 token/s at 07:00 with 69.6% CV.
  • No missing-data limitation applies: all eight models delivered 12 of 12 expected samples, and the dataset reports 96 valid points with 100.0% coverage over the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model by average throughput at 112.88 token/s, narrowly ahead of gemma4:31b at 112.72 token/s, while nemotron-3-ultra is the weakest at 34.1 token/s.
  • nemotron-3-ultra shows the most operationally significant volatility, swinging between 2.37 and 107.9 token/s with a 92.5% coefficient of variation and a -34.7% trend; glm-5.3 also oscillates between 62.02 and 160.22 token/s.
  • No missing-data limitation applies: all eight models report 12 of 12 expected samples, yielding 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model on average throughput at 119.16 token/s (peak 192.08 token/s), while nemotron-3-ultra is the weakest at 38.22 token/s average, with a low of 2.37 token/s.
  • nemotron-3-ultra shows the most operationally significant volatility, with a coefficient of variation of 91.5% and swings between 2.37 and 107.9 token/s; glm-5.3-flash posted the steepest decline, trending -38.3% to finish at 92.89 token/s.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, totaling 96 valid points at 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 134.06 token/s average throughput, while nemotron-3-ultra is the weakest at 44.36 token/s; deepseek-v4-flash leads the remaining models at 107.93 token/s.
  • nemotron-3-ultra shows the most operationally significant volatility, with a coefficient of variation of 84.9% and swings between 2.37 and 107.9 token/s, including a drop to 2.37 token/s at 07:00 UTC; glm-5.3-flash is also unstable at 53.6% CV.
  • No missing-data limitation exists: all eight models report 12 of 12 samples, valid_point_count is 96, and coverage is 100.0% across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput was gemma4:31b at 136.19 token/s; weakest was nemotron-3-ultra at 48.03 token/s, with deepseek-v4-pro (95.25 token/s) and minimax-m3 (97.82 token/s) also below 100 token/s.
  • glm-5.3-flash showed the highest volatility, ranging from 16.85 to 200.22 token/s (CV 53.7%), while glm-3 declined 44.8% overall and hit a low of 6.19 token/s at 04:20; nemotron-3-ultra swung between 2.58 and 107.9 token/s.
  • No missing-data limitation applies: all eight models recorded 12 of 12 samples, totaling 96 valid points at 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b delivered the highest average throughput at 140.0 token/s (peak 195.2 token/s), while nemotron-3-ultra was weakest at 44.95 token/s average, never exceeding 97.05 token/s.
  • The most operationally significant volatility came from glm-5.3-flash, which swung between 200.22 token/s at 02:40 and 16.85 token/s at 03:20 (cv 51.9%), and glm-5.3 fell to 6.19 token/s at 04:20; nemotron-3-ultra was also erratic, dipping to 2.58 token/s at 03:00 with cv 76.4%.
  • No missing-data limitation applies: all 8 models recorded 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model by average throughput at 148.74 token/s, while nemotron-3-ultra is the weakest at 40.63 token/s, roughly a 3.7x gap between the two.
  • The most operationally significant volatility is nemotron-3-ultra, whose throughput swings between 2.58 and 95.27 token/s with a coefficient of variation of 84.4%; glm-5.3-flash also spiked to 200.22 token/s at 02:40 before collapsing to 16.85 token/s at 03:20.
  • No missing-data limitation applies: all eight models recorded 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 142.54 token/s average throughput, ahead of deepseek-v4-flash at 125.79 token/s, while nemotron-3-ultra is weakest at 40.05 token/s average with a floor of 2.58 token/s.
  • nemotron-3-ultra shows the most operationally significant volatility, with a 74.9% coefficient of variation and swings between 2.58 and 87.38 token/s; glm-5.3-flash also spiked to 200.22 token/s at 02:40 before dropping to 80.18 token/s at 03:00.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0%, so every 20-minute observation is present.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 153.69 token/s average throughput, peaking at 183.91 token/s; nemotron-3-ultra is the weakest at 32.36 token/s average, ending at just 4.09 token/s.
  • glm-5.3 shows the most volatility, ranging from 56.69 to 175.08 token/s with a 41.3% coefficient of variation, while glm-5.2 declined 32.8% over the window, from 95.03 to 49.16 token/s.
  • No missing-data limitation exists: all eight models have 12 of 12 samples, and the dataset reports 96 valid points with 100.0% coverage.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model with an average throughput of 148.67 token/s (p95 180.66 token/s), while nemotron-3-ultra is the weakest at 32.25 token/s average, dipping as low as 5.34 token/s.
  • glm-5.3 shows the most volatility, with a coefficient of variation of 34.7% and swings between 56.69 and 175.08 token/s; deepseek-v4-flash dropped sharply in the final observation to 62.52 token/s, its four-hour minimum, versus a 118.71 token/s average.
  • No missing-data limitation applies: all eight models have 12 of 12 expected samples, 96 valid points, and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model by average throughput at 138.02 token/s, while nemotron-3-ultra is the weakest at 31.48 token/s; deepseek-v4-flash ranks second at 122.13 token/s.
  • nemotron-3-ultra shows the most operationally significant volatility, swinging between 5.34 and 67.24 token/s with a coefficient of variation of 74.6%, including repeated drops below 10 token/s at 21:00, 22:00, 22:40, and 23:00 UTC; glm-5.2 also climbed 46.8% over the window, from 60.26 to 99.46 token/s.
  • No missing-data limitation applies: all eight models delivered 12 of 12 expected samples, giving 96 valid points and 100.0% coverage for the four-hour period.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b posted the strongest average throughput at 132.37 token/s (p95 174.76 token/s), while nemotron-3-ultra was the weakest at 26.67 token/s average, never exceeding 67.24 token/s.
  • The most operationally significant volatility came from nemotron-3-ultra, with a coefficient of variation of 84.7% and repeated collapses to roughly 5–9 token/s (latest 5.34 token/s) between spikes near 67 token/s; glm-5.3 also trended up 44.8% to a 155.85 token/s peak.
  • No missing-data limitation applies: all eight models report 12 of 12 samples, 96 valid points total, and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 123.92 token/s average throughput, peaking at 175.31 token/s; nemotron-3-ultra is weakest at 30.38 token/s average, with a low of 5.41 token/s.
  • Volatility is the dominant operational signal: nemotron-3-ultra shows a 76.1% coefficient of variation, and glm-5.3-flash collapsed to 11.36 token/s at 20:40 before recovering to 82.07 token/s at 21:40; deepseek-v4-flash posted the steepest upward trend at 30.3%.
  • No missing-data limitation exists: all eight models delivered 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 123.39 token/s average throughput, with a p95 of 170.5 token/s and a maximum of 174.3 token/s; nemotron-3-ultra is the weakest at 27.76 token/s average, never exceeding 77.12 token/s.
  • glm-5.3 shows the steepest decline, trending -38.8% from 145.17 token/s at 17:20 to 61.91 token/s at 21:00, while deepseek-v4-pro improved 36.3% to 110.81 token/s; nemotron-3-ultra is the most volatile at 69.2% coefficient of variation, ranging 6.03 to 77.12 token/s.
  • No missing-data limitation applies: all 8 models delivered 12 of 12 expected samples, yielding 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 109.44 token/s average throughput, narrowly ahead of glm-5.3 at 108.73 token/s; nemotron-3-ultra is the weakest at 24.88 token/s average, never exceeding 77.12 token/s.
  • Volatility dominates: nemotron-3-ultra swings between 3.77 and 77.12 token/s with a coefficient of variation of 80.7%, and glm-5.3-flash spikes to 190.98 token/s at 17:00 before ending near 81 token/s; gemma4:31b shows the only positive trend at +13.7%.
  • No missing-data limitation applies: all eight models report 12 of 12 samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model at 124.9 token/s average throughput, while nemotron-3-ultra is the weakest at 22.94 token/s average, with a minimum of 3.77 token/s.
  • glm-5.3-flash shows the most operationally significant volatility, with a coefficient of variation of 55.9% and swings between 190.98 and 44.64 token/s; its throughput declined 33.6% over the window, ending at 69.46 token/s.
  • No missing-data limitation exists: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0%, so the four-hour window is fully represented.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model with average throughput of 114.58 token/s (p95 148.98 token/s), while nemotron-3-ultra is the weakest at 21.05 token/s average, peaking at only 62.34 token/s.
  • glm-5.3-flash shows the widest swings, ranging from 44.64 to 190.98 token/s with a coefficient of variation of 56.0%, and nemotron-3-ultra is both volatile (cv 77.3%) and declining 32.0% over the window, ending at 15.81 token/s.
  • No missing-data limitation applies: all eight models delivered 12 of 12 expected samples, and overall coverage is 100.0% with 96 valid points, so the four-hour window is fully represented.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model at 117.39 token/s average throughput, ahead of deepseek-v4-flash at 108.88 token/s and gemma4:31b at 103.81 token/s; nemotron-3-ultra is the weakest at 26.66 token/s average, never exceeding 62.34 token/s.
  • glm-5.3-flash shows the sharpest volatility, with a 57.5% coefficient of variation and swings between 47.18 and 190.98 token/s, while gemma4:31b, minimax-m3, and nemotron-3-ultra declined across the window by 37.5%, 40.0%, and 53.7% respectively.
  • No missing-data limitation applies: all eight models report 12 of 12 expected samples, with 96 valid points and 100.0% coverage, so every twenty-minute observation in the four-hour window is present.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model with an average throughput of 125.95 token/s (p95 151.16 token/s), while nemotron-3-ultra is the weakest at 27.83 token/s average, never exceeding 62.34 token/s.
  • glm-5.3-flash shows the sharpest volatility, spiking from 47.18 token/s at 13:40 to 190.28 token/s at 15:20 before dropping to 63.36 token/s at 16:00, with a 54.7% coefficient of variation; minimax-m3 also declined 34.1% to a period-low 30.79 token/s.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, and the dataset reports 96 valid points at 100.0% coverage.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 delivered the highest average throughput at 121.01 token/s, narrowly ahead of gemma4:31b at 118.11 token/s, while nemotron-3-ultra was weakest at 28.67 token/s, never exceeding 62.34 token/s.
  • nemotron-3-ultra showed the greatest volatility, with a 66.3% coefficient of variation and swings between 2.54 and 62.34 token/s; minimax-m3 posted the strongest trend, rising 48.0% overall and peaking at 163.02 token/s at 13:20.
  • No missing-data limitation applies: all eight models recorded 12 of 12 expected samples, yielding 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model with an average throughput of 124.61 token/s (peaking at 173.25 token/s), while nemotron-3-ultra is the weakest at 29.46 token/s average, never exceeding 55.05 token/s.
  • The most operationally significant volatility is nemotron-3-ultra, which swings between 2.54 and 55.05 token/s with a coefficient of variation of 64.7%, including three observations below 5 token/s. glm-5.3-flash also shows a notable decline, with a trend of -32.2% and a drop from a 152.19 token/s peak to 47.18 token/s.
  • No missing-data limitation exists: all eight models have 12 of 12 expected samples, valid_point_count is 96, and overall coverage is 100.0%, so the four-hour window is fully represented.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 125.83 token/s average throughput, peaking at 181.44 token/s; nemotron-3-ultra is the weakest at 24.61 token/s average, with a floor of 2.54 token/s.
  • nemotron-3-ultra shows the most operationally significant volatility, swinging between 2.54 and 55.05 token/s with a coefficient of variation of 81.7%, while gemma4:31b declined 24.3% over the window despite its lead.
  • No missing-data limitation applies: all eight models delivered 12 of 12 expected samples, 96 valid points, and 100% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model by average throughput at 124.66 token/s (peaking at 181.44 token/s), while nemotron-3-ultra is the weakest at 25.58 token/s on average, with a floor of just 3.33 token/s.
  • The most operationally significant volatility is nemotron-3-ultra, whose coefficient of variation is 75.4% with swings between 3.33 and 55.05 token/s across the window; by contrast deepseek-v4-pro is the steadiest performer, holding 84.36 to 137.21 token/s with a 15.2% coefficient of variation.
  • No missing-data limitation applies: coverage is 100.0% with 96 valid points, and every model recorded 12 of 12 expected samples, so the four-hour picture is complete.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model by average throughput at 127.51 token/s (p95 174.95 token/s), while nemotron-3-ultra is the weakest at 24.55 token/s average, never exceeding 55.05 token/s in any interval.
  • The most operationally significant movement is minimax-m3's late surge, jumping from 89.26 token/s at 10:40 to 168.2 token/s at 11:00, driving a +32.3% trend; nemotron-3-ultra is the most volatile, with a 70.9% coefficient of variation and swings between 2.73 and 55.05 token/s.
  • No missing-data limitation applies: every model has 12 of 12 samples, valid_point_count is 96, and coverage is 100.0% across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 127.58 token/s average throughput, with a peak of 181.44 token/s and a p95 of 172.6 token/s; nemotron-3-ultra is the weakest at 26.42 token/s average, never exceeding 77.53 token/s.
  • glm-5.3-flash shows the most operationally significant volatility, dropping from 186.06 token/s at 06:20 to 53.05 token/s at 10:00, a -40.5% trend with 48.0% coefficient of variation; nemotron-3-ultra is even more erratic at 79.6% CV, swinging between 2.73 and 77.53 token/s.
  • No missing-data limitation applies: all eight models report 12 of 12 samples, valid_point_count is 96, and coverage is 100.0% across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput was gemma4:31b at 124.21 token/s (peaking at 186.77 token/s), while nemotron-3-ultra was weakest at 30.20 token/s, never exceeding 86.78 token/s and averaging under a quarter of gemma4:31b's rate.
  • glm-5.3-flash showed the sharpest volatility, swinging between 51.98 and 186.06 token/s with a 50.4% coefficient of variation and a -29.7% trend; nemotron-3-ultra was even less stable at 87.7% CV, dropping to 2.73 token/s at 07:20.
  • No missing-data limitation exists: all eight models delivered 12 of 12 samples, and the dataset reports 96 valid points with 100.0% coverage over the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 143.9 token/s average throughput, ahead of deepseek-v4-flash at 102.15 token/s; nemotron-3-ultra is the weakest at 29.02 token/s average, with a minimum of 2.26 token/s.
  • glm-5.3-flash shows the only positive trend at +26.3%, ending at 136.05 token/s, but swings between 51.98 and 186.06 token/s; nemotron-3-ultra is the most volatile at 100.4% coefficient of variation, and gemma4:31b declined 25.0% to 113.24 token/s.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0% across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput is gemma4:31b at 133.15 token/s (peaking at 194.06 token/s); weakest is nemotron-3-ultra at 36.86 token/s, with a low of 2.26 token/s.
  • The most operationally significant volatility is nemotron-3-ultra, which swings between 2.26 and 86.78 token/s (cv 78.6%) with near-zero readings from 04:40 through 06:20; glm-5.3 is also unstable (cv 52.0%, range 25.38 to 180.01 token/s), and most models dip near 06:00.
  • No missing-data limitation applies: all 8 models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0%, so the four-hour window is fully populated.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 142.24 token/s average throughput, peaking at 194.06 token/s; nemotron-3-ultra is the weakest at 32.19 token/s average, with a low of 2.26 token/s.
  • glm-5.3 shows the sharpest decline, falling 39.1% from 177.13 token/s at 02:20 to 25.38 token/s at 06:00, while nemotron-3-ultra is highly volatile (89.0% CV), swinging between 86.78 and 2.26 token/s.
  • No missing-data limitation applies: all 8 models recorded 12 of 12 samples, giving 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 133.36 token/s average throughput, peaking at 194.06 token/s, while nemotron-3-ultra is the weakest at 29.19 token/s average, never exceeding 66.55 token/s.
  • The most significant trend is glm-5.3's decline of 31.9%, falling from a 177.13 token/s peak around 02:20 to 52.05–64.66 token/s in the final three observations; glm-5.3-flash fell 22.9% over the same window. nemotron-3-ultra is highly volatile with a 77.9% coefficient of variation, swinging between 2.26 and 66.55 token/s.
  • No missing-data limitation applies: all 8 models have 12 of 12 samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model at 130.42 token/s average throughput, ahead of gemma4:31b at 120.81 token/s; nemotron-3-ultra is the weakest at 28.63 token/s average, with a maximum of only 66.55 token/s.
  • The most operationally significant movement is glm-5.3-flash, which fell from 190.15 token/s at 00:20 to 78.96 token/s at 04:00, a -49.3% trend with 46.6% coefficient of variation; deepseek-v4-pro improved 33.1% to 109.4 token/s.
  • No missing-data limitation applies: all 8 models have 12 of 12 samples, 96 valid points, and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model with an average throughput of 136.27 token/s, edging gemma4:31b at 130.13 token/s, while nemotron-3-ultra is weakest at 23.65 token/s, far below the next-lowest glm-5.2 at 74.15 token/s.
  • deepseek-v4-flash shows the sharpest decline, trending -25.9% from roughly 150 token/s early in the window to a 75.2 token/s low at 02:40; glm-5.3-flash is the most volatile, with a 46.5% coefficient of variation and swings between 57.89 and 190.15 token/s.
  • No missing-data limitation applies: all eight models report 12 of 12 expected samples, 96 valid points, and 100.0% coverage, so the four-hour window is complete.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model at 132.30 token/s average throughput, edging glm-5.3 at 130.87 token/s; nemotron-3-ultra is the weakest at 30.43 token/s average, with a low of 2.92 token/s at 00:00.
  • The most operationally significant volatility is nemotron-3-ultra, whose coefficient of variation is 91.7% with swings between 2.92 and 104.07 token/s; glm-5.3-flash also oscillates sharply, ranging from 70.62 to 190.15 token/s (cv 42.4%).
  • No missing-data limitation applies: all eight models report 12 of 12 samples, valid_point_count is 96, and coverage is 100.0% across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model at 140.39 token/s average throughput, while nemotron-3-ultra is the weakest at 39.37 token/s average, with a floor of 2.92 token/s.
  • The most operationally significant volatility is nemotron-3-ultra's decline: its throughput fell 58.8% over the window, with a 75.3% coefficient of variation and readings as low as 2.92 token/s near 00:00 UTC, versus deepseek-v4-flash's steadier 19.0% variation.
  • No missing-data limitation exists: all eight models delivered 12 of 12 expected samples, and the dataset reports 96 valid points with 100.0% coverage, so the four-hour window is fully represented.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model at 132.14 token/s average throughput (12 samples, range 93.09 to 165.04 token/s), while nemotron-3-ultra is the weakest at 48.94 token/s average (range 3.55 to 104.07 token/s).
  • The most operationally significant volatility is nemotron-3-ultra's 61.5% coefficient of variation, including dips to 3.55 token/s at 22:20 and 5.47 token/s at 19:40; glm-5.3 also swung from 14.06 to 149.7 token/s within the window.
  • No missing-data limitation applies: all eight models have 12 of 12 expected samples, with 96 valid points and 100.0% coverage.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model at 132.59 token/s average throughput (range 93.09 to 165.04 token/s), while nemotron-3-ultra is the weakest at 45.43 token/s average, dipping as low as 5.47 token/s at 19:40.
  • The most operationally significant volatility is glm-5.3, with a coefficient of variation of 50.7%: it fell to 14.06 token/s at 21:00, then spiked to 147.69 token/s at 21:20 and 146.02 token/s at 21:40. glm-5.3-flash shows comparable swings, from 50.41 to 188.94 token/s.
  • No missing-data limitation applies: all 8 models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0% across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model with an average throughput of 126.59 token/s (p95 150.99 token/s), while nemotron-3-ultra is the weakest at 37.24 token/s average, dipping to a minimum of 4.81 token/s.
  • The most operationally significant movement is glm-5.3's decline of 47.5%, falling from 134.65 token/s at 19:00 to 14.06 token/s at 21:00; glm-5.3-flash also shows high volatility (cv 55.0%), spiking to 188.94 token/s at 20:40.
  • No missing-data limitation exists: all eight models recorded 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash posted the strongest average throughput at 120.09 token/s, ahead of gemma4:31b at 100.31 token/s, while nemotron-3-ultra was weakest at 24.03 token/s, never exceeding 52.53 token/s in any interval.
  • glm-5.3-flash showed the most operationally significant volatility, with a coefficient of variation of 52.7% and swings from 43.26 token/s at 18:00 to 183.59 token/s at 17:00 and 182.29 token/s at 18:40; glm-5.3 also declined 30.0% overall, ending at 57.67 token/s.
  • No missing-data limitation applies: all eight models recorded 12 of 12 expected samples, totaling 96 valid points at 100.0% coverage for the window.