← Performance dashboard

Hourly performance insights

Summaries of rolling four-hour performance data

1189 retained summaries
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model at 146.9 token/s average throughput, peaking at 237.72 token/s; nemotron-3-ultra is the weakest at 24.72 token/s average, never exceeding 40.01 token/s.
  • glm-5.2 shows the most operationally significant volatility, with a 79.2% coefficient of variation, a single 203.28 token/s spike at 19:00 against a 61.38 token/s average, and a -21.8% trend; deepseek-v4-pro is steadiest at 13.7% CV.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0% across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput is deepseek-v4.1-flash at 155.96 token/s; weakest is nemotron-3-ultra at 18.87 token/s, ranging from 3.72 to 40.01 token/s across the window.
  • glm-5.2 shows the most extreme volatility (CV 90.7%), spiking from 34.37 token/s at 18:20 to 203.28 token/s at 19:00; deepseek-v4.1-flash also dipped to 41.63 token/s at 18:00 before recovering to 144.78 token/s.
  • No missing-data limitation exists: all 8 models report 12 of 12 samples, with 96 valid points and 100.0% coverage over the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model by average throughput at 164.58 token/s, while nemotron-3-ultra is the weakest at 19.98 token/s, roughly an eighth of the leader's rate.
  • The most operationally significant volatility is glm-5.2, whose throughput fell 41.0% over the window with a 76.6% coefficient of variation, including a one-off spike to 197.41 token/s at 16:00 UTC; deepseek-v4.1-flash also dropped to 41.63 token/s in its final sample.
  • No missing-data limitation applies: all eight models have 12 of 12 expected samples, and the dataset reports 96 valid points with 100.0% coverage.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model at 163.78 token/s average throughput, peaking at 216.17 token/s, while nemotron-3-ultra is the weakest at 16.0 token/s average, never exceeding 33.36 token/s.
  • glm-5.2 shows the highest volatility, with a 70.6% coefficient of variation and swings from 28.68 to 197.41 token/s; nemotron-3-ultra shows the steepest decline, trending down 57.0% to a 3.72 token/s low at 16:00 UTC.
  • No missing-data limitation applies: all eight models report 12 of 12 samples, and the dataset shows 96 valid points with 100.0% coverage over the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model at 154.58 token/s average throughput, while nemotron-3-ultra is the weakest at 17.08 token/s average.
  • deepseek-v4.1-flash shows the only positive trend, up 28.9% to 186.52 token/s at 16:00; glm-5.2 is the most volatile at 60.9% coefficient of variation, swinging between 29.99 and 197.41 token/s, and minimax-m3's latest reading fell to 10.31 token/s, its minimum.
  • No missing-data limitation exists: all eight models have 12 of 12 samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model at 151.75 token/s average throughput (peaking at 216.17 token/s), while nemotron-3-ultra is the weakest at 17.26 token/s average (minimum 7.6 token/s).
  • glm-5.3 shows the steepest decline, trending -27.2% with a drop from 142.48 token/s at 12:00 to 19.2 token/s at 14:20; nemotron-3-ultra rose 84.4% to 33.36 token/s. glm-5.2 is the most volatile, with a 49.1% coefficient of variation and a range of 6.92 to 139.3 token/s.
  • No missing-data limitation: all eight models report 12 of 12 samples, 96 valid points, and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model at 157.46 token/s average throughput, peaking at 203.72 token/s, while nemotron-3-ultra is the weakest at 18.85 token/s average, never exceeding 71.14 token/s.
  • The most operationally significant movement is nemotron-3-ultra's collapse from 71.14 token/s at 10:20 to single digits (7.6 to 10.56 token/s) for most of the window, a 45.0% decline; glm-5.2 also swung between 6.92 and 169.78 token/s with 52.7% coefficient of variation.
  • No missing-data limitation exists: all eight models have 12 of 12 samples, 96 valid points, and 100.0% coverage, though conclusions rest on only four hours of observations.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model at 168.36 token/s average throughput, while nemotron-3-ultra is the weakest at 18.53 token/s average.
  • glm-5.2 shows the most volatility, with a coefficient of variation of 60.0% and a swing from 169.78 token/s at 11:00 down to 6.92 token/s at 12:00; nemotron-3-ultra also declined 51.3% across the window.
  • No missing-data limitation: all 8 models have 12 of 12 samples, 96 valid points, and 100.0% coverage across 20-minute observations from 09:20 to 13:00 UTC.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model at 179.55 token/s average throughput (peak 207.47 token/s), while nemotron-3-ultra is the weakest at 20.19 token/s average, never exceeding 71.14 token/s.
  • glm-5.2 shows the most operationally significant volatility, with a coefficient of variation of 72.8%: it spiked to 169.78 token/s at 11:00 UTC, then collapsed to 6.92 token/s at 12:00 UTC, a swing of roughly 163 token/s within forty minutes.
  • No missing-data limitation exists in this window: all 8 models delivered 12 of 12 expected samples, yielding 96 valid points and 100.0% coverage across the four-hour period.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model at 183.51 token/s average throughput (range 150.44 to 219.87 token/s), while nemotron-3-ultra is the weakest at 20.71 token/s average, including a low of 3.81 token/s.
  • glm-5.2 shows the most operationally significant volatility, with a 64.0% coefficient of variation, swings between 18.09 and 169.78 token/s, and a final reading of 169.78 token/s far above its 62.55 token/s average; nemotron-3-ultra is also erratic at 79.3% CV.
  • No missing-data limitation exists: all 8 models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0%, so the four-hour window is fully populated.
4-hour window · 95 points · 99.0% coverage
  • deepseek-v4.1-flash is the strongest model at 185.35 token/s average throughput (peak 219.87 token/s), while nemotron-3-ultra is the weakest at 15.89 token/s average, dipping as low as 3.81 token/s.
  • glm-5.2 shows the most operationally significant volatility, with a 58.3% coefficient of variation and a -46.6% trend, falling from a 133.66 token/s peak at 07:00 to 18.09 token/s at 08:40; glm-5.3-flash also swings between 224.47 and 62.21 token/s.
  • deepseek-v4.1-flash has only 11 of 12 expected samples (91.7% coverage), missing the 06:20 observation, so its average may be less comparable than fully covered models.
4-hour window · 92 points · 95.8% coverage
  • Highest average throughput was deepseek-v4.1-flash at 184.42 token/s (peak 219.87 token/s), while nemotron-3-ultra was lowest at 18.97 token/s, never exceeding 30.56 token/s.
  • glm-5.3-flash showed the sharpest decline, trending -36.9% and falling from a 224.47 token/s peak at 06:20 to 65.08 token/s at 09:00; glm-5.2 also deteriorated late, ending at 19.08 token/s versus its 133.66 token/s peak.
  • deepseek-v4.1-flash reported only 8 of 12 expected samples (66.7% coverage), with no observations from 05:20 to 06:20, so its average is based on partial data.
4-hour window · 89 points · 92.7% coverage
  • deepseek-v4.1-flash leads on average throughput at 192.64 token/s (peak 219.87 token/s), while nemotron-3-ultra is weakest at 22.47 token/s average, peaking at only 43.96 token/s.
  • nemotron-3-ultra shows the steepest decline, down 46.2% across the window to a low of 3.81 token/s at 07:20, and glm-5.3-flash swung from a 224.47 token/s peak at 06:20 to 62.21 token/s at 08:00.
  • deepseek-v4.1-flash has only 5 of 12 expected samples (41.7% coverage), all after 06:40, so its leading average rests on partial data; overall coverage is 92.7% with 89 valid points.
4-hour window · 86 points · 89.6% coverage
  • glm-5.3-flash leads full-coverage models at 142.74 token/s average (peak 224.47 token/s); nemotron-3-ultra is weakest at 29.26 token/s average, ending at 14.59 token/s. deepseek-v4.1-flash averages 209.81 token/s but from only two samples.
  • glm-5.3-flash rose from 61.38 token/s at 03:20 to a 224.47 token/s peak at 06:20 (+51.9% trend), while glm-5.2 spiked to 133.66 token/s at 07:00; glm-5.2 and glm-5.3 are the most volatile, with CVs of 52.6% and 46.2%.
  • deepseek-v4.1-flash has only 2 of 12 expected samples (16.7% coverage), limiting confidence in its 209.81 token/s average; overall dataset coverage is 89.6% with 86 valid points.
4-hour window · 84 points · 87.5% coverage
  • glm-5.3-flash is the strongest model at 138.75 token/s average throughput, peaking at 215.92 token/s at 06:00 UTC, while nemotron-3-ultra is the weakest at 28.58 token/s average.
  • glm-5.2 shows the largest upward trend at +72.7%, climbing from 41.79 token/s at 02:20 to 63.95 token/s at 06:00; nemotron-3-ultra is the most volatile with a 54.7% coefficient of variation, ranging 4.76 to 61.81 token/s, and glm-5.3 dipped to 9.93 token/s at 03:20.
  • deepseek-v4-flash reported zero samples across all 12 expected intervals, reducing overall coverage to 87.5% (84 of 96 valid points), so its throughput is unknown for this window.
4-hour window · 84 points · 87.5% coverage
  • glm-5.3-flash is the strongest model with an average throughput of 125.18 token/s (range 61.38–175.56 token/s), while nemotron-3-ultra is the weakest at 33.74 token/s average, dipping as low as 4.76 token/s.
  • glm-5.2 shows the most operationally significant volatility, with a coefficient of variation of 77.9% and swings between 19.28 and 205.45 token/s, plus a -46.9% trend; deepseek-v4-pro also declined -24.1% to a latest reading of 49.34 token/s.
  • deepseek-v4-flash has no samples for the entire window (0 of 12 expected), leaving overall coverage at 87.5% (84 valid points), so its throughput cannot be assessed.
4-hour window · 84 points · 87.5% coverage
  • glm-5.3-flash is the strongest model with an average throughput of 129.13 token/s (p95 173.69 token/s), while nemotron-3-ultra is the weakest at 31.33 token/s average, peaking at only 61.81 token/s.
  • The most significant movement is glm-5.2's decline of 63.7 percent, falling from a 205.45 token/s peak at 01:40 to a 19.28 token/s low at 03:40; a broad dip also hit all reporting models at 03:20, with glm-5.3 dropping to 9.93 token/s.
  • deepseek-v4-flash reported 0 of 12 expected samples, so its throughput is unknown and overall coverage is 87.5 percent (84 valid points), limiting comparability across the full model set.
4-hour window · 84 points · 87.5% coverage
  • glm-5.3-flash is the strongest model at 131.66 token/s average throughput, edging out gemma4:31b at 126.91 token/s, while nemotron-3-ultra is the weakest at 28.24 token/s average, ending at just 4.76 token/s.
  • glm-5.2 shows the most volatility, with a 72.1% coefficient of variation and swings from 22.0 to 205.45 token/s; gemma4:31b also dipped sharply to 43.48 token/s at 02:00 before recovering to 133.89 token/s.
  • deepseek-v4-flash reported zero samples across all 12 expected intervals, leaving its throughput unmeasured and reducing overall coverage to 87.5% (84 of 96 valid points).
4-hour window · 84 points · 87.5% coverage
  • glm-5.3-flash is the strongest model with an average throughput of 134.02 token/s (peaking at 187.3 token/s), while nemotron-3-ultra is the weakest at 30.99 token/s average, never exceeding 53.76 token/s.
  • glm-5.2 shows the most operationally significant volatility, with a coefficient of variation of 68.6% and a swing from 29.68 token/s at 00:00 to 205.45 token/s at 01:40, a 114.3% trend increase over the window.
  • deepseek-v4-flash reported zero samples across all 12 expected 20-minute intervals, so overall coverage is only 87.5% (84 of 96 expected points), and no throughput figures can be assessed for that model.
4-hour window · 84 points · 87.5% coverage
  • glm-5.3-flash is the strongest model with an average throughput of 142.28 token/s (peaking at 222.81 token/s), while nemotron-3-ultra is the weakest at 24.42 token/s average, roughly one-sixth of the leader's rate.
  • The most significant movement is glm-5.3's decline from 164.66 token/s at 23:40 to 53.7 token/s at 00:40, a -12.1% trend; glm-5.2 also shows high volatility with a 50.3% coefficient of variation and a late spike to 131.68 token/s at 01:00.
  • deepseek-v4-flash reported 0 of 12 expected samples (0% coverage), leaving its performance unknown and reducing overall coverage to 87.5% (84 valid points).
4-hour window · 84 points · 87.5% coverage
  • glm-5.3 is the strongest model with an average throughput of 138.84 token/s (peak 178.39 token/s), while nemotron-3-ultra is the weakest at 26.73 token/s average, peaking at only 39.56 token/s.
  • glm-5.2 shows the most operationally significant volatility: a 52.2% coefficient of variation and a -29.5% trend, falling from a 132.56 token/s peak at 20:40 to 29.68 token/s at 00:00, with a low of 28.66 token/s at 21:00.
  • deepseek-v4-flash has zero samples out of 12 expected (0.0% coverage), so its throughput is entirely unknown; overall dataset coverage is 87.5% with 84 valid points, limiting any fleet-wide comparison for that model.
4-hour window · 84 points · 87.5% coverage
  • glm-5.3-flash is the strongest model by average throughput at 134.38 token/s, peaking at 222.81 token/s, while nemotron-3-ultra is the weakest at 25.11 token/s average, never exceeding 39.56 token/s.
  • glm-5.3-flash shows the largest upward trend at 43.3%, closing at 187.3 token/s, but glm-5.2 is the most volatile with a 52.7% coefficient of variation, swinging from 22.07 to 132.56 token/s within the window.
  • deepseek-v4-flash reported zero samples across all 12 expected observations, so its throughput is unknown, and overall coverage is only 87.5% (84 of 96 valid points), limiting window-wide comparisons.
4-hour window · 84 points · 87.5% coverage
  • glm-5.3 is the strongest model with an average throughput of 128.26 token/s (latest 159.0 token/s), while nemotron-3-ultra is the weakest at 23.07 token/s average, peaking at only 39.56 token/s.
  • glm-5.3 shows the largest upward trend at 24.4% over the window, and glm-5.3-flash spiked to 222.81 token/s at 21:40 after dipping to 52.34 token/s at 21:00; glm-5.2 is the most volatile with a 65.3% coefficient of variation, ranging 22.07 to 182.47 token/s.
  • deepseek-v4-flash reported 0 of 12 samples (0% coverage), reducing overall coverage to 87.5% (84 valid points), so its throughput cannot be assessed for this period.
4-hour window · 84 points · 87.5% coverage
  • glm-5.3-flash is the strongest model by average throughput at 110.21 token/s, narrowly ahead of glm-5.3 at 109.60 token/s, while nemotron-3-ultra is the weakest at 21.75 token/s, roughly one-fifth of the leader's rate.
  • glm-5.2 shows the most volatility, with a coefficient of variation of 67.7% and swings from a 182.47 token/s spike at 19:00 down to 22.07 token/s at 20:00, plus a -28.8% trend; gemma4:31b also dipped to 31.40 token/s at 19:00.
  • deepseek-v4-flash reported zero samples against 12 expected, leaving overall coverage at 87.5% (84 of 96 valid points), so its throughput cannot be assessed in this window.
4-hour window · 84 points · 87.5% coverage
  • glm-5.3-flash is the strongest model by average throughput at 110.26 token/s, narrowly ahead of gemma4:31b at 109.48 token/s; nemotron-3-ultra is weakest at 17.35 token/s, with a maximum of only 32.86 token/s.
  • glm-5.2 shows the most volatility, with a coefficient of variation of 73.5% and a swing from 13.41 token/s at 16:40 to a spike of 182.47 token/s at 19:00, followed by a drop to 22.07 token/s at 20:00.
  • deepseek-v4-flash has no samples (0 of 12 expected, 0.0% coverage), so its throughput is unknown; overall coverage is 87.5% with 84 valid points out of 96 expected.
4-hour window · 84 points · 87.5% coverage
  • glm-5.3 is the strongest model by average throughput at 107.81 token/s, narrowly ahead of glm-5.3-flash (102.40 token/s) and gemma4:31b (103.36 token/s); nemotron-3-ultra is the weakest at 17.24 token/s, roughly one-sixth of the leader.
  • glm-5.2 shows the most operationally significant volatility, with a coefficient of variation of 84.8% and swings between 13.41 and 216.33 token/s, including a jump to 182.47 token/s at 19:00; deepseek-v4-pro also dipped to 29.95 token/s at 17:20.
  • deepseek-v4-flash reported 0 of 12 expected samples, reducing overall coverage to 87.5% (84 valid points), so fleet-wide comparisons exclude that model entirely.
4-hour window · 84 points · 87.5% coverage
  • gemma4:31b is the strongest model by average throughput at 107.47 token/s, while nemotron-3-ultra is the weakest at 16.2 token/s; glm-5.3-flash follows at 94.38 token/s and minimax-m3 at 30.98 token/s.
  • glm-5.2 shows the most operationally significant volatility, with a coefficient of variation of 101.6% and a spread from 13.41 to 216.33 token/s, including a single spike to 216.33 token/s at 16:00 UTC against a 55.4 token/s average.
  • deepseek-v4-flash has no data, with 0 of 12 expected samples and 0.0% coverage, reducing overall coverage to 87.5% (84 valid points), so its throughput cannot be assessed in this window.
4-hour window · 84 points · 87.5% coverage
  • glm-5.3-flash is the strongest model at 95.4 token/s average throughput, while nemotron-3-ultra is the weakest at 17.07 token/s average, with minimax-m3 also low at 25.34 token/s.
  • glm-5.2 shows the most operationally significant volatility, spiking to 216.33 token/s at 16:00 UTC against a 47.23 token/s average and a 117.0% coefficient of variation; gemma4:31b also swung between 47.07 and 166.23 token/s across the window.
  • deepseek-v4-flash returned no samples (0 of 12 expected, 0.0% coverage), so its throughput is unknown; overall coverage is 87.5% with 84 valid points, limiting fleet-wide conclusions.
4-hour window · 84 points · 87.5% coverage
  • glm-5.3-flash is the strongest model at 100.74 token/s average throughput, ahead of deepseek-v4-pro at 96.21 token/s; nemotron-3-ultra is the weakest at 18.55 token/s average, with minimax-m3 close behind at 25.38 token/s.
  • glm-5.2 shows the most extreme volatility, with a coefficient of variation of 109.3% and a final 16:00 spike to 216.33 token/s against a 49.04 token/s average; glm-5.3-flash is the steadiest, ranging 76.42 to 133.83 token/s.
  • deepseek-v4-flash returned zero samples across all 12 expected observations, leaving 0% coverage and reducing overall dataset coverage to 87.5% (84 valid points), so its throughput cannot be assessed.
4-hour window · 84 points · 87.5% coverage
  • glm-5.3-flash is the strongest model at 112.04 token/s average throughput (76.47–183.17 token/s range), while nemotron-3-ultra is the weakest at 18.73 token/s average, peaking at only 32.08 token/s.
  • Every model shows declining throughput over the window; glm-5.2 fell 56.5% overall and is the most volatile, swinging between 9.93 and 117.83 token/s with a 66.7% coefficient of variation, including a spike at 12:00 UTC.
  • deepseek-v4-flash has no data (0 of 12 samples, 0% coverage), reducing dataset coverage to 87.5% (84 of 96 expected points), so its throughput cannot be assessed.
4-hour window · 84 points · 87.5% coverage
  • glm-5.3-flash is the strongest model at 117.35 token/s average throughput (peak 183.17 token/s), while nemotron-3-ultra is the weakest at 17.63 token/s average (minimum 2.52 token/s).
  • Most models declined late in the window: gemma4:31b fell 31.6%, glm-5.2 fell 32.3%, and minimax-m3 dropped to 5.55 token/s at 14:00. glm-5.2 was also the most volatile, with a 58.1% coefficient of variation spanning 9.93 to 117.83 token/s.
  • deepseek-v4-flash reported 0 of 12 samples (0% coverage), reducing overall coverage to 87.5% (84 valid points), so its throughput cannot be assessed.
4-hour window · 84 points · 87.5% coverage
  • glm-5.3-flash is the strongest model at 117.08 token/s average throughput, edging glm-5.3 at 111.85 token/s; nemotron-3-ultra is the weakest at 17.37 token/s average, below minimax-m3's 31.45 token/s.
  • gemma4:31b shows the sharpest volatility, swinging between 40.17 and 156.18 token/s with a 40.0% coefficient of variation, including a 44.93 token/s reading at 12:00 followed by 121.26 token/s at 12:20; nemotron-3-ultra also fluctuates widely, from 2.24 to 32.08 token/s.
  • deepseek-v4-flash has no samples (0 of 12 expected, 0.0% coverage), so its throughput cannot be assessed; overall coverage is 87.5% with 84 valid points.
4-hour window · 84 points · 87.5% coverage
  • glm-5.3-flash is the strongest model with an average throughput of 120.26 token/s (peaking at 183.17 token/s), while nemotron-3-ultra is the weakest at 15.62 token/s average, never exceeding 28.36 token/s.
  • The most operationally significant volatility comes from nemotron-3-ultra, whose coefficient of variation is 58.8%, swinging between 2.24 and 28.36 token/s, including repeated drops below 6 token/s at 10:00, 10:40, and 11:00 UTC; deepseek-v4-pro shows the clearest upward trend at 22.0%.
  • deepseek-v4-flash contributed zero samples across all 12 expected observations, so its throughput is entirely unknown; overall dataset coverage is 87.5% with 84 valid points, limiting any fleet-wide conclusions for the four-hour window.
4-hour window · 84 points · 87.5% coverage
  • glm-5.3-flash is the strongest model by average throughput at 113.79 token/s (peaking at 154.62 token/s), narrowly ahead of glm-5.3 at 112.32 token/s, while nemotron-3-ultra is the weakest at 13.82 token/s average, never exceeding 28.46 token/s.
  • deepseek-v4-pro shows the most operationally significant volatility, swinging between 15.75 and 114.25 token/s across the window while posting the largest upward trend at 40.0%; nemotron-3-ultra is also erratic, with a 65.5% coefficient of variation and repeated drops below 6 token/s.
  • deepseek-v4-flash contributed no samples (0 of 12 expected, 0.0% coverage), so its throughput is unknown; overall dataset coverage is 87.5% with 84 valid points, limiting conclusions about fleet-wide performance.
4-hour window · 84 points · 87.5% coverage
  • glm-5.3-flash is the strongest model at 131.0 token/s average throughput, peaking at 232.65 token/s at 06:40 UTC; nemotron-3-ultra is weakest at 18.31 token/s average, falling to 2.24 token/s at 10:00 UTC.
  • The most significant volatility is nemotron-3-ultra's 72.9% coefficient of variation, swinging between 54.01 and 2.24 token/s; glm-5.3-flash also shows a -19.2% trend, dropping from 232.65 token/s at 06:40 to 63.95 token/s at 09:40.
  • deepseek-v4-flash has zero samples (0 of 12 expected, 0.0% coverage), so its throughput is unknown; overall dataset coverage is 87.5% with 84 valid points, limiting completeness of comparisons.
4-hour window · 84 points · 87.5% coverage
  • glm-5.3-flash is the strongest model at 129.76 token/s average throughput, peaking at 232.65 token/s; nemotron-3-ultra is the weakest at 20.39 token/s average, never exceeding 54.01 token/s.
  • glm-5.2 shows the most operationally significant volatility, swinging from 145.7 token/s at 06:00 to 13.03 token/s at 07:20, with a coefficient of variation of 57.4% and a -32.3% trend; nemotron-3-ultra also declined 43.7% across the window.
  • deepseek-v4-flash reported zero samples against 12 expected intervals (0% coverage), lowering overall coverage to 87.5% with 84 of 96 valid points, so its throughput cannot be assessed in this window.
4-hour window · 84 points · 87.5% coverage
  • glm-5.3-flash is the strongest model with an average throughput of 133.57 token/s, peaking at 232.65 token/s at 06:40 UTC; nemotron-3-ultra is the weakest at 25.10 token/s average, never exceeding 54.01 token/s.
  • glm-5.2 shows the most volatility, swinging from 155.72 token/s at 05:00 to 13.03 token/s at 07:20, with a coefficient of variation of 66.4% and a 36.5% downward trend; glm-5.3-flash stayed steadier, never dropping below 93.11 token/s.
  • deepseek-v4-flash reported zero of 12 expected samples, contributing to overall coverage of 87.5% (84 valid points), so its throughput cannot be assessed for this window.
4-hour window · 84 points · 87.5% coverage
  • glm-5.3-flash is the strongest model at 134.77 token/s average throughput, while nemotron-3-ultra is the weakest at 30.90 token/s, with minimax-m3 close behind at 37.07 token/s.
  • glm-5.3 shows the largest upward trend at +51.9%, climbing from roughly 66 token/s early in the window to peaks of 186.71 token/s, though its 58.3% coefficient of variation makes it highly volatile; nemotron-3-ultra declined 26.9%.
  • deepseek-v4-flash reported 0 of 12 expected samples, leaving 84 valid points and 87.5% overall coverage, so its throughput is unknown and period-wide averages exclude it.
4-hour window · 84 points · 87.5% coverage
  • glm-5.3-flash is the strongest model at 121.98 token/s average throughput, peaking at 211.11 token/s; nemotron-3-ultra is the weakest at 32.18 token/s average, never exceeding 63.34 token/s.
  • glm-5.2 shows the highest volatility with a 57.5% coefficient of variation, swinging between 30.04 and 165.82 token/s; glm-5.3 has the largest upward trend at +22.6%, while nemotron-3-ultra declined -18.0%.
  • deepseek-v4-flash has zero samples out of 12 expected, so its throughput is unknown; overall dataset coverage is 87.5% with 84 valid points across the eight models.
4-hour window · 84 points · 87.5% coverage
  • glm-5.3-flash is the strongest model by average throughput at 124.95 token/s, while nemotron-3-ultra is the weakest at 33.41 token/s; gemma4:31b and deepseek-v4-pro fall between at 97.79 and 92.99 token/s.
  • glm-5.2 shows the sharpest decline, with a -23.0% trend and throughput falling from a 165.82 token/s peak at 03:00 to 30.04 token/s at 04:00; it also has the highest volatility at 54.8% coefficient of variation, and deepseek-v4-pro trends -22.3%.
  • deepseek-v4-flash reported zero samples (0 of 12 expected), so its throughput is unknown; overall coverage is 87.5% with 84 valid points, limiting window-wide comparisons.
4-hour window · 84 points · 87.5% coverage
  • glm-5.3-flash is the strongest model at 130.84 token/s average (peak 226.83 token/s), while nemotron-3-ultra is the weakest at 34.73 token/s average (minimum 10.17 token/s).
  • glm-5.3 shows the sharpest decline, falling from 171.01 token/s at 00:40 to 66.19 token/s at 04:00, a -25.6% trend; glm-5.2 spiked to 165.82 token/s at 03:00 before dropping to 30.04 token/s by 04:00.
  • deepseek-v4-flash reported no samples across all 12 expected intervals (0% coverage), leaving overall coverage at 84 of 96 points (87.5%), so its throughput cannot be assessed.
4-hour window · 84 points · 87.5% coverage
  • glm-5.3-flash is the strongest model at 134.4 token/s average throughput (peak 226.83 token/s), while nemotron-3-ultra is the weakest at 31.88 token/s average (low 10.17 token/s).
  • glm-5.3 shows the sharpest operational swing, falling 44.8% overall from a 172.43 token/s peak at 23:40 to 51.49 token/s at 01:40 before recovering to 152.08 token/s at 03:00; glm-5.2 spiked 138% to 165.82 token/s in its latest sample.
  • deepseek-v4-flash reported 0 of 12 expected samples (0% coverage), reducing overall coverage to 87.5% (84 valid points), so its throughput is unknown for this window.
4-hour window · 84 points · 87.5% coverage
  • glm-5.3-flash is the strongest model with an average throughput of 135.22 token/s (peaking at 226.83 token/s), while nemotron-3-ultra is the weakest at 28.97 token/s average, never exceeding 54.11 token/s.
  • The most operationally significant movement is glm-5.3's late-window collapse: it held 141.30–173.41 token/s from 22:20 through 01:00, then fell to 72.56, 51.49, and 64.29 token/s in the final three observations. Conversely, glm-5.2 recovered from a 9.42 token/s low at 00:00 to 89.62 token/s by 02:00.
  • deepseek-v4-flash reported zero of 12 expected samples (0% coverage), reducing overall valid points to 84 of 96 (87.5%); its throughput cannot be assessed this period.
4-hour window · 84 points · 87.5% coverage
  • glm-5.3-flash is the strongest model with an average throughput of 144.15 token/s (peaking at 226.83 token/s), while nemotron-3-ultra is the weakest at 25.36 token/s average, never exceeding 54.11 token/s.
  • nemotron-3-ultra shows the most volatility, with a coefficient of variation of 60.9% and swings between 9.88 and 54.11 token/s; glm-5.3 is the steadiest high performer, averaging 132.11 token/s with a 21.6% upward trend.
  • deepseek-v4-flash has no samples (0 of 12 expected), so its throughput is unknown; overall coverage is 87.5% (84 of 96 valid points), limiting comparability.
4-hour window · 84 points · 87.5% coverage
  • glm-5.3-flash is the strongest model at 127.67 token/s average throughput, edging glm-5.3 at 124.99 token/s; nemotron-3-ultra is the weakest at 21.11 token/s average, below minimax-m3's 56.78 token/s.
  • glm-5.2 shows the sharpest decline, trending -26.9% and falling from 74.07 token/s at 20:20 to 9.42 token/s at 00:00; minimax-m3 is the most volatile, ranging 17.91 to 127.68 token/s with a 47.8% coefficient of variation.
  • deepseek-v4-flash reported zero of 12 expected samples (0% coverage), leaving its throughput unknown; overall coverage is 87.5% with 84 valid points, which limits cross-model comparison.
4-hour window · 84 points · 87.5% coverage
  • glm-5.3 is the strongest model by average throughput at 117.92 token/s, narrowly ahead of glm-5.3-flash at 115.26 token/s, while nemotron-3-ultra is the weakest at 19.10 token/s, never exceeding 28.63 token/s in any observation.
  • The most significant volatility comes from minimax-m3, which swings between 3.26 and 127.68 token/s with a 52.8% coefficient of variation; glm-5.3-flash shows the steepest upward trend at 48.7%, including a spike to 211.11 token/s at 22:00 UTC.
  • deepseek-v4-flash reported zero samples across all 12 expected intervals, so its throughput is unknown and overall coverage drops to 87.5% (84 of 96 valid points).
4-hour window · 84 points · 87.5% coverage
  • Strongest average throughput was glm-5.3-flash at 105.05 token/s, peaking at 211.11 token/s in the final observation; weakest was nemotron-3-ultra at 18.14 token/s, never exceeding 28.63 token/s.
  • The most operationally significant volatility came from glm-5.2, which swung between 7.14 and 196.32 token/s with a coefficient of variation of 86.2% and a 37.7% downward trend, while glm-5.3-flash jumped from 101.28 to 211.11 token/s at 22:00.
  • deepseek-v4-flash reported zero of 12 expected samples (0% coverage), reducing overall coverage to 87.5% (84 valid points), so its throughput cannot be assessed for this window.
4-hour window · 84 points · 87.5% coverage
  • gemma4:31b delivered the strongest average throughput at 103.04 token/s, while nemotron-3-ultra was weakest at 18.62 token/s, never exceeding 28.63 token/s across the window.
  • glm-5.2 showed the most operationally significant volatility, swinging between 21.02 and 196.32 token/s with an 80.0% coefficient of variation and a -47.3% trend; glm-5.3 also dipped to 8.52 token/s at 18:40 before recovering to 160.43 token/s at 20:40.
  • deepseek-v4-flash reported zero samples against 12 expected, leaving its throughput unmeasured and reducing overall coverage to 87.5% (84 of 96 expected points), so model-level comparisons exclude it entirely.
4-hour window · 84 points · 87.5% coverage
  • glm-5.3 is the strongest model by average throughput at 90.31 token/s, narrowly ahead of gemma4:31b (89.87 token/s) and deepseek-v4-pro (89.79 token/s); nemotron-3-ultra is the weakest at 19.01 token/s, peaking at only 27.56 token/s.
  • The most operationally significant movement is glm-5.3's decline of 35.8% over the window, from 132.83 token/s at 16:20 to 76.22 token/s at 20:00, including a low of 8.52 token/s at 18:40; glm-5.2 shows the widest volatility, ranging 11.59 to 203.8 token/s with a 93.6% coefficient of variation.
  • deepseek-v4-flash reported 0 of 12 expected samples (0.0% coverage), so its throughput is unknown; overall coverage is 87.5% with 84 valid points, limiting any fleet-wide comparison.
4-hour window · 84 points · 87.5% coverage
  • gemma4:31b is the strongest model on average throughput at 101.61 token/s, while nemotron-3-ultra is the weakest at 22.13 token/s, roughly a fifth of the leader's pace.
  • glm-5.2 shows the most operationally significant volatility, with a coefficient of variation of 76.1% and swings from 11.59 to 203.8 token/s within the window; glm-5.3 also declined 21.7% and nemotron-3-ultra fell 37.0% over the period.
  • deepseek-v4-flash reported zero samples across all 12 expected observations, leaving 0% coverage and reducing overall dataset coverage to 87.5% (84 of 96 expected points), so its throughput is unassessed.