← Performance dashboard

Hourly performance insights

Summaries of rolling four-hour performance data

1190 retained summaries
4-hour window · 84 points · 87.5% coverage
  • gemma4:31b is the strongest model on average throughput at 101.61 token/s, while nemotron-3-ultra is the weakest at 22.13 token/s, roughly a fifth of the leader's pace.
  • glm-5.2 shows the most operationally significant volatility, with a coefficient of variation of 76.1% and swings from 11.59 to 203.8 token/s within the window; glm-5.3 also declined 21.7% and nemotron-3-ultra fell 37.0% over the period.
  • deepseek-v4-flash reported zero samples across all 12 expected observations, leaving 0% coverage and reducing overall dataset coverage to 87.5% (84 of 96 expected points), so its throughput is unassessed.
4-hour window · 84 points · 87.5% coverage
  • glm-5.3 is the strongest model with an average throughput of 103.63 token/s (range 65.39–139.09 token/s), while nemotron-3-ultra is the weakest at 22.80 token/s average, peaking at only 43.98 token/s.
  • glm-5.2 shows the most operationally significant volatility, with a coefficient of variation of 81.3% and swings from 11.59 token/s at 17:00 to 203.80 token/s at 16:40; minimax-m3 also fluctuated widely, from 17.72 to 100.11 token/s.
  • deepseek-v4-flash reported no samples (0 of 12 expected), reducing overall coverage to 87.5% (84 valid points of 96 expected), so its throughput is unknown for this window.
4-hour window · 84 points · 87.5% coverage
  • glm-5.3 is the strongest model with an average throughput of 100.63 token/s (range 65.39–134.52 token/s), while nemotron-3-ultra is the weakest at 23.71 token/s average, peaking at only 43.98 token/s.
  • glm-5.2 shows the most operationally significant volatility, swinging between 11.59 and 203.80 token/s with a coefficient of variation of 85.0% and a +91.0% trend, ending at 11.59 token/s at 17:00 UTC; minimax-m3 also declined 16.6% over the window.
  • deepseek-v4-flash reported no samples (0 of 12 expected), so its throughput is unknown; overall coverage is 87.5% with 84 of 96 expected points, limiting comparability.
4-hour window · 84 points · 87.5% coverage
  • glm-5.3 is the strongest model with an average throughput of 108.94 token/s (range 65.39 to 143.34 token/s), while nemotron-3-ultra is the weakest at 23.49 token/s average, peaking at only 43.98 token/s.
  • The most significant volatility is glm-5.2, with a coefficient of variation of 85.6% and swings from 2.36 to 199.19 token/s; minimax-m3 shows the steepest decline, trending down 46.1% from a 150.39 token/s peak to 17.72 token/s at 16:00 UTC.
  • deepseek-v4-flash reported zero samples across all 12 expected intervals, leaving 84 of 96 expected points (87.5% coverage), so its throughput is unassessed for this window.
4-hour window · 84 points · 87.5% coverage
  • glm-5.3 is the strongest model by average throughput at 103.08 token/s, while nemotron-3-ultra is the weakest at 23.17 token/s; the next-lowest averages are glm-5.2 at 51.46 token/s and glm-5.3-flash at 63.81 token/s.
  • glm-5.2 shows the most extreme volatility, ranging from 2.36 to 199.19 token/s with a 101.2% coefficient of variation, and it declined 25.9% over the window; minimax-m3 fell 26.2%, dropping from 145.59 token/s at 11:20 to 28.12 token/s at 15:00.
  • deepseek-v4-flash has no data (0 of 12 samples, 0.0% coverage), and overall coverage is 87.5% with 84 valid points, so model comparisons exclude it entirely.
4-hour window · 84 points · 87.5% coverage
  • Strongest average throughput was minimax-m3 at 108.61 token/s, edging glm-5.3 at 107.45 token/s; weakest was nemotron-3-ultra at 21.63 token/s, far below the next-lowest glm-5.2 at 55.63 token/s.
  • glm-5.2 showed the sharpest volatility, with a coefficient of variation of 93.7%, spiking to 199.19 token/s at 12:20 then collapsing to 2.36 token/s at 13:00; glm-5.3-flash fell 39.6% overall, from a 230.44 token/s peak to 61.78 token/s at 14:00.
  • deepseek-v4-flash reported 0 of 12 samples (0% coverage), leaving 84 valid points and 87.5% overall coverage, so fleet-wide comparisons exclude that model entirely.
4-hour window · 84 points · 87.5% coverage
  • glm-5.3 is the strongest model by average throughput at 116.52 token/s (p95 154.09 token/s), while nemotron-3-ultra is the weakest at 22.93 token/s average, never exceeding 54.52 token/s.
  • The most operationally significant volatility is the 13:00 UTC collapse affecting two models simultaneously: minimax-m3 fell from 150.39 token/s at 12:40 to 7.6 token/s, and glm-5.2 dropped from 43.89 token/s to 2.36 token/s. glm-5.3-flash also swung from a 230.44 token/s peak at 10:40 down to 56.15 token/s at 11:20, with 54.8% coefficient of variation.
  • deepseek-v4-flash reported zero samples across all 12 expected intervals (0% coverage), reducing overall valid points to 84 of 96 (87.5% coverage), so its throughput is entirely unmeasured in this window.
4-hour window · 84 points · 87.5% coverage
  • glm-5.3 is the strongest model by average throughput at 102.88 token/s, narrowly ahead of gemma4:31b (101.79 token/s) and minimax-m3 (101.76 token/s); nemotron-3-ultra is the weakest at 23.01 token/s average, peaking at only 54.52 token/s.
  • The most operationally significant volatility is glm-5.3-flash, which spiked to 230.44 token/s at 10:40 before falling back to 56.15 token/s by 11:20, with a 54.7% coefficient of variation; minimax-m3 also swung from a 20.8 token/s low to 145.59 token/s.
  • deepseek-v4-flash reported zero samples across all 12 expected intervals, so its throughput is unknown and overall coverage drops to 87.5% (84 of 96 expected points).
4-hour window · 84 points · 87.5% coverage
  • glm-5.3 is the strongest model by average throughput at 115.55 token/s, while nemotron-3-ultra is the weakest at 19.02 token/s, roughly a sixfold gap.
  • minimax-m3 shows the most operationally significant trend, rising 97.2% from 60.95 token/s at 07:20 to 127.65 token/s at 11:00; glm-5.2 is the most volatile, with a 91.0% coefficient of variation and swings between 10.15 and 220.6 token/s.
  • deepseek-v4-flash reported zero samples across all 12 expected observations, so its throughput is unknown; overall coverage is 87.5% with 84 valid points, limiting conclusions about fleet-wide performance.
4-hour window · 84 points · 87.5% coverage
  • gemma4:31b is the strongest model with an average throughput of 113.06 token/s (p95 150.0 token/s), while nemotron-3-ultra is the weakest at 17.92 token/s average, never exceeding 51.03 token/s in any interval.
  • glm-5.2 shows the most operationally significant volatility, with a coefficient of variation of 94.8% and throughput swinging between 10.15 and 220.6 token/s; its trend fell 58.1% across the window, and glm-5.3 dropped to 6.76 token/s at 08:20.
  • deepseek-v4-flash has no samples (0 of 12 expected, 0% coverage), so its throughput is unknown; overall coverage is 87.5% with 84 valid points, limiting fleet-wide comparison completeness.
4-hour window · 77 points · 80.2% coverage
  • gemma4:31b is the strongest model at 116.51 token/s average throughput, while nemotron-3-ultra is the weakest at 17.12 token/s average, never exceeding 34.13 token/s in any observation.
  • glm-5.2 shows the most operationally significant volatility, swinging between 10.15 and 220.60 token/s with a coefficient of variation of 88.8%, including a drop to 10.15 token/s at 07:40; minimax-m3 also declined 50.0% overall, ending at 20.80 token/s.
  • deepseek-v4-flash has zero samples (0.0% coverage), so its throughput is unknown; every other model has 11 of 12 samples (91.7% coverage), and overall coverage is 80.2% with 77 valid points, limiting completeness.
4-hour window · 84 points · 87.5% coverage
  • gemma4:31b is the strongest model by average throughput at 115.03 token/s, narrowly ahead of glm-5.3 at 114.81 token/s; nemotron-3-ultra is the weakest at 19.44 token/s, with a maximum of only 45.24 token/s.
  • glm-5.2 shows the most extreme volatility, swinging between 10.15 and 220.6 token/s with a coefficient of variation of 82.8%, including a drop from 220.6 to 10.15 token/s between 07:20 and 07:40; minimax-m3 also declined 46.9% overall, ending at 26.94 token/s.
  • deepseek-v4-flash has no samples for the entire window (0 of 12 expected, 0% coverage), so its throughput is unknown; overall dataset coverage is 87.5% with 84 valid points, limiting fleet-wide conclusions.
4-hour window · 84 points · 87.5% coverage
  • glm-5.3 is the strongest model with an average throughput of 123.38 token/s (peak 176.0 token/s), while nemotron-3-ultra is the weakest at 18.09 token/s average, never exceeding 45.24 token/s.
  • glm-5.2 shows the most volatility, with a coefficient of variation of 74.5%; it spiked to 211.72 token/s at 06:40 UTC before collapsing to 13.27 token/s at 07:00. deepseek-v4-pro also dipped to 4.56 token/s at 06:00.
  • deepseek-v4-flash has no data (0 of 12 samples, 0.0% coverage), lowering overall coverage to 87.5% (84 of 96 valid points), so fleet-wide averages exclude that model entirely.
4-hour window · 84 points · 87.5% coverage
  • glm-5.3 is the strongest model with an average throughput of 133.2 token/s (peak 176.0 token/s at 03:20), while nemotron-3-ultra is the weakest at 25.71 token/s average, never exceeding 69.7 token/s.
  • The most significant volatility is glm-5.2, with a 62.4% coefficient of variation, spiking to 180.18 token/s at 03:00 before falling to 34.91 token/s at 03:20; minimax-m3 shows the largest upward trend at 111.3%, climbing from 47.01 token/s at 02:40 to 158.55 token/s at 05:20.
  • deepseek-v4-flash reported 0 of 12 expected samples (0% coverage), reducing overall coverage to 87.5% (84 valid points), so its throughput cannot be assessed this period.
4-hour window · 84 points · 87.5% coverage
  • glm-5.3 is the strongest model with an average throughput of 130.85 token/s (p95 174.05 token/s), while nemotron-3-ultra is the weakest at 28.88 token/s average, peaking at only 70.85 token/s.
  • glm-5.2 shows the most operationally significant volatility, swinging between 27.78 and 197.48 token/s with a coefficient of variation of 69.8% and a trend of -30.9%; nemotron-3-ultra is similarly unstable (cv 75.7%, trend -45.9%).
  • deepseek-v4-flash has no data for the entire window (0 of 12 samples, 0% coverage), reducing overall coverage to 87.5% (84 of 96 expected points), so its throughput cannot be assessed.
4-hour window · 84 points · 87.5% coverage
  • glm-5.3 is the strongest model with an average throughput of 128.69 token/s (latest 172.46 token/s), while nemotron-3-ultra is the weakest at 23.79 token/s average, peaking at only 70.85 token/s.
  • glm-5.3 shows the clearest upward trend, rising 39.8% to a sustained 166-176 token/s after 03:20, while minimax-m3 fell 39.8% from a 177.39 token/s peak at 00:20 to 48.1 token/s at 04:00; glm-5.2 is the most volatile at 70.8% coefficient of variation, swinging between 27.78 and 197.48 token/s.
  • deepseek-v4-flash reported zero of 12 expected samples (0% coverage), reducing overall dataset coverage to 87.5% (84 valid points), so its throughput cannot be assessed in this window.
4-hour window · 84 points · 87.5% coverage
  • gemma4:31b is the strongest model on average throughput at 122.08 token/s, narrowly ahead of glm-5.3 at 120.05 token/s, while nemotron-3-ultra is the weakest at 30.35 token/s average, peaking at only 70.85 token/s.
  • The most operationally significant movement is minimax-m3's decline of 34.6% across the window, falling from 177.39 token/s at 00:20 to 57.07 token/s at 03:00; glm-5.2 also swung between 26.85 and 197.48 token/s with 76.7% coefficient of variation.
  • deepseek-v4-flash reported zero of 12 expected samples, contributing to overall coverage of 87.5% (84 valid points), so its throughput cannot be assessed for this period.
4-hour window · 84 points · 87.5% coverage
  • gemma4:31b delivered the strongest average throughput at 118.7 token/s across all 12 samples, narrowly ahead of minimax-m3 at 118.24 token/s, while nemotron-3-ultra was weakest at 24.97 token/s average, peaking at only 70.85 token/s.
  • The most operationally significant volatility came from glm-5.2, whose coefficient of variation reached 76.1%, swinging between 26.85 and 197.48 token/s and ending on its maximum; minimax-m3 also fell sharply from 177.39 token/s at 00:20 to 41.22 token/s at 01:00.
  • deepseek-v4-flash reported 0 of 12 expected samples, reducing overall coverage to 87.5% (84 valid points), so its throughput is entirely unassessed and period-wide comparisons exclude it.
4-hour window · 84 points · 87.5% coverage
  • minimax-m3 is the strongest model with an average throughput of 120.66 token/s, narrowly ahead of gemma4:31b at 118.87 token/s, while nemotron-3-ultra is the weakest at 18.98 token/s, roughly six times slower than the leader.
  • glm-5.2 shows the most extreme volatility, with a coefficient of variation of 83.5% and swings from 26.85 to 208.08 token/s, including a drop to 31.88 token/s at 00:40 followed by a spike to 167.48 token/s at 01:00.
  • deepseek-v4-flash has no samples for the entire window (0 of 12 expected), reducing overall coverage to 87.5% (84 valid points), so its throughput cannot be assessed.
4-hour window · 84 points · 87.5% coverage
  • Highest average throughput belongs to minimax-m3 at 134.75 token/s (peak 171.93 token/s, latest 137.34 token/s); lowest is nemotron-3-ultra at 22.64 token/s, whose maximum of 63.19 token/s falls below most models' averages.
  • glm-5.2 is the most volatile model, with an 80.7% coefficient of variation, values ranging from 26.85 to 208.08 token/s, and a -41.6% trend ending at its minimum; minimax-m3 was steadiest at 15.9% CV.
  • deepseek-v4-flash reported zero of 12 expected samples (0% coverage), so its throughput is unknown; overall coverage is 87.5% with 84 valid points, limiting fleet-wide comparison.
4-hour window · 84 points · 87.5% coverage
  • minimax-m3 is the strongest model by average throughput at 132.15 token/s, while nemotron-3-ultra is the weakest at 17.89 token/s, roughly 7.4x lower.
  • glm-5.2 shows the highest volatility with a coefficient of variation of 65.2%, swinging from a 29.2 token/s low to a 208.08 token/s spike at 22:00 UTC; nemotron-3-ultra also declined 30.0% over the window to 14.49 token/s.
  • deepseek-v4-flash reported 0 of 12 expected samples (0.0% coverage), so overall coverage is only 84 valid points, or 87.5%, limiting fleet-wide conclusions.
4-hour window · 84 points · 87.5% coverage
  • Strongest average throughput was minimax-m3 at 129.9 token/s (12 of 12 samples); weakest was nemotron-3-ultra at 17.23 token/s, peaking at only 31.2 token/s.
  • glm-5.2 showed the sharpest volatility, with a 69.2% coefficient of variation, a low of 10.37 token/s at 19:00, and a spike to 208.08 token/s at 22:00 versus its 69.8 token/s average; minimax-m3 also fell to 88.57 token/s at 22:00 from a 159.58 token/s high.
  • deepseek-v4-flash reported 0 of 12 expected samples (0.0% coverage), and overall coverage was 87.5% with 84 valid points, so per-model comparisons rest on incomplete data.
4-hour window · 84 points · 87.5% coverage
  • minimax-m3 is the strongest model at 133.24 token/s average throughput; nemotron-3-ultra is the weakest at 19.59 token/s average, peaking at only 32.41 token/s.
  • glm-5.2 shows the largest upward trend at 98.3%, climbing from 21.72 token/s at 17:20 to a maximum of 87.45 token/s at 20:00, while gemma4:31b and glm-5.3-flash are highly volatile with coefficient-of-variation values of 42.8% and 47.7%.
  • deepseek-v4-flash reported zero samples across all 12 expected intervals, leaving its throughput unmeasured and reducing overall dataset coverage to 87.5% (84 valid points).
4-hour window · 84 points · 87.5% coverage
  • Highest average throughput was minimax-m3 at 129.66 token/s (peak 166.0 token/s, CV 15.5%), while nemotron-3-ultra was weakest at 18.81 token/s average, never exceeding 44.36 token/s.
  • glm-5.2 showed the strongest upward trend at +37.3%, climbing from a 10.37 token/s low at 19:00 to 87.45 token/s at 20:00; nemotron-3-ultra was the most volatile (CV 62.3%) and declined 28.4%, and glm-5.3-flash spiked to 199.21 token/s at 19:20 before dropping to 54.79 token/s.
  • deepseek-v4-flash reported 0 of 12 expected samples (0% coverage), leaving its throughput unknown; overall coverage was 87.5% with 84 valid points.
4-hour window · 84 points · 87.5% coverage
  • minimax-m3 is the strongest model at 128.25 token/s average throughput (range 86.3 to 166.0 token/s), while nemotron-3-ultra is the weakest at 22.47 token/s average, peaking at only 45.82 token/s.
  • glm-5.2 shows the steepest decline, trending -49.1% and ending at 10.37 token/s, its four-hour minimum; gemma4:31b is highly volatile with a 40.9% coefficient of variation, swinging between 27.01 and 162.98 token/s.
  • deepseek-v4-flash has no samples (0 of 12 expected, 0.0% coverage), so its throughput is unknown; overall coverage is 87.5% with 84 of 96 expected points, limiting fleet-wide conclusions.
4-hour window · 84 points · 87.5% coverage
  • minimax-m3 is the strongest model by average throughput at 123.22 token/s (peaking at 166.0 token/s at 18:00), while nemotron-3-ultra is the weakest at 24.92 token/s average, never exceeding 45.82 token/s across the window.
  • glm-5.2 shows the most operationally significant volatility and decline: a 43.3% downward trend, coefficient of variation of 61.6%, a drop from 152.68 token/s at 14:40 to 15.95 token/s at 15:00, and a final reading of 34.84 token/s, far below its 63.39 token/s average.
  • deepseek-v4-flash has no data for the entire period (0 of 12 samples, 0% coverage), reducing overall dataset coverage to 87.5% (84 of 96 expected points), so its throughput cannot be assessed.
4-hour window · 84 points · 87.5% coverage
  • minimax-m3 is the strongest model at 121.14 token/s average throughput (74.21–154.28 token/s range), while nemotron-3-ultra is the weakest at 23.44 token/s average, peaking at only 45.82 token/s.
  • glm-5.2 shows the most operationally significant volatility, with a 56.7% coefficient of variation and sharp drops to 12.75 token/s at 14:00 and 15.95 token/s at 15:00, plus a -18.2% trend; gemma4:31b also fell to 27.01 token/s in its final sample.
  • deepseek-v4-flash reported 0 of 12 expected samples (0% coverage), leaving its throughput unknown; overall coverage is 84 of 96 points (87.5%), so conclusions rest on incomplete data.
4-hour window · 84 points · 87.5% coverage
  • Highest average throughput was minimax-m3 at 114.13 token/s (peaking at 154.28 token/s), while nemotron-3-ultra was weakest at 21.69 token/s average and never exceeded 45.82 token/s.
  • glm-5.2 showed the sharpest volatility, with a 59.0% coefficient of variation and swings between 12.75 and 155.93 token/s within the window; deepseek-v4-pro posted the strongest upward trend at 39.6%.
  • deepseek-v4-flash returned no data (0 of 12 samples), leaving overall coverage at 87.5% (84 of 96 expected points), so any cross-model comparison excludes that model entirely.
4-hour window · 84 points · 87.5% coverage
  • glm-5.3 is the strongest model at 108.94 token/s average throughput, ahead of minimax-m3 (104.65 token/s) and gemma4:31b (102.36 token/s); nemotron-3-ultra is weakest at 17.25 token/s average, never exceeding 29.39 token/s.
  • glm-5.2 shows the most volatility, with a 71.0% coefficient of variation and swings between 12.75 and 155.93 token/s, including a drop from 155.93 token/s at 13:40 to 12.75 token/s at 14:00; glm-5.3-flash declined 30.3% over the window to 54.46 token/s.
  • deepseek-v4-flash reported zero samples across all 12 expected intervals, leaving its throughput unknown and reducing overall coverage to 87.5% (84 of 96 expected points).
4-hour window · 84 points · 87.5% coverage
  • Strongest average throughput is minimax-m3 at 109.1 token/s (peak 144.37 token/s); weakest is nemotron-3-ultra at 14.48 token/s average, never exceeding 27.89 token/s.
  • glm-5.2 shows the sharpest volatility (CV 72.1%), ranging 12.75 to 155.93 token/s; it trended up 69.4% overall, reaching 155.93 token/s at 13:40 before dropping to 12.75 token/s at 14:00. glm-5.3-flash also swung from a 174.2 token/s peak down to 60.05 token/s.
  • deepseek-v4-flash reported 0 of 12 samples (0% coverage), so it is excluded from comparisons; overall coverage is 87.5% with 84 valid points, which limits model-to-model conclusions.
4-hour window · 84 points · 87.5% coverage
  • Strongest average throughput was gemma4:31b at 102.47 token/s; weakest was nemotron-3-ultra at 18.9 token/s, roughly one-fifth of the leader's rate.
  • glm-5.2 showed the most volatility, with a 101.5% coefficient of variation and swings between 16.88 and 196.55 token/s; minimax-m3 posted the largest gain, trending up 85.3% from 47.62 token/s at 08:20 to a 144.37 token/s peak at 10:40.
  • deepseek-v4-flash returned 0 of 12 expected samples (0% coverage), leaving its throughput unknown and reducing overall dataset coverage to 87.5% (84 valid points), so comparisons exclude that model entirely.
4-hour window · 84 points · 87.5% coverage
  • gemma4:31b is the strongest model with an average throughput of 104.39 token/s (peak 151.14 token/s), while nemotron-3-ultra is the weakest at 18.59 token/s average, never exceeding 55.68 token/s in any observation.
  • The most operationally significant movement is minimax-m3's upward trend of 127.7%, climbing from a low of 15.55 token/s at 07:40 to 144.37 token/s at 10:40 and 135.43 token/s at 11:00. Volatility is also high for glm-5.2, whose coefficient of variation is 98.7% with swings between 16.88 and 196.55 token/s.
  • deepseek-v4-flash has no data for the entire window, with 0 of 12 expected samples and 0.0% coverage, reducing overall dataset coverage to 87.5% (84 valid points), so its throughput cannot be assessed.
4-hour window · 84 points · 87.5% coverage
  • gemma4:31b is the strongest model by average throughput at 110.47 token/s, while nemotron-3-ultra is the weakest at 17.36 token/s, with a minimum of 4.91 token/s and a maximum of 55.68 token/s.
  • glm-5.2 shows the most extreme volatility, with a coefficient of variation of 99.5% and swings between 16.88 and 196.55 token/s, including a spike to 196.55 token/s at 08:40 followed by a drop to 16.88 token/s at 09:00.
  • deepseek-v4-flash has zero samples for the period, contributing to overall coverage of 87.5% (84 of 96 expected points), so its throughput cannot be assessed from this dataset.
4-hour window · 84 points · 87.5% coverage
  • gemma4:31b is the strongest model at 114.2 token/s average throughput, while nemotron-3-ultra is the weakest at 12.37 token/s average, never exceeding 26.66 token/s across the window.
  • glm-5.2 shows the most extreme volatility, ranging from 16.88 to 196.55 token/s with a coefficient of variation of 100.1%; glm-5.3 shows the steepest trend, rising 43.2% and peaking at 166.83 token/s.
  • deepseek-v4-flash has no samples (0 of 12 expected, 0% coverage), so its throughput is unknown; overall coverage is 87.5% with 84 valid points, limiting fleet-wide conclusions.
4-hour window · 84 points · 87.5% coverage
  • gemma4:31b is the strongest model with an average throughput of 119.48 token/s (peak 157.29 token/s), while nemotron-3-ultra is the weakest at 18.48 token/s average, peaking at only 69.25 token/s.
  • The most significant volatility is nemotron-3-ultra's 96.9% coefficient of variation, collapsing from 69.25 token/s at 04:20 to 3.55 token/s at 06:00. minimax-m3 also declined 55.1%, from 140.98 token/s at 04:20 to 53.88 token/s at 08:00, and glm-5.2 spiked to 188.19 token/s at 06:40.
  • deepseek-v4-flash has zero samples out of 12 expected, so its throughput is unknown; overall coverage is 87.5% (84 of 96 expected points), limiting model comparisons.
4-hour window · 84 points · 87.5% coverage
  • Highest average throughput was glm-5.3-flash at 110.94 token/s, narrowly ahead of gemma4:31b at 108.54 token/s; lowest was nemotron-3-ultra at 21.52 token/s.
  • The sharpest operational movement is nemotron-3-ultra, down 65.5% from a 69.25 token/s peak to 13.81 token/s at 07:00 with 82.8% coefficient of variation; minimax-m3 also fell 35.0%, ending at 12.24 token/s.
  • deepseek-v4-flash reported 0 of 12 samples (0.0% coverage), so its throughput is unknown; overall coverage was 87.5% with 84 valid points, limiting fleet-wide conclusions.
4-hour window · 84 points · 87.5% coverage
  • glm-5.3-flash is the strongest model by average throughput at 111.9 token/s, edging gemma4:31b at 107.05 token/s, while nemotron-3-ultra is the weakest at 26.68 token/s average.
  • The most significant volatility is glm-5.3, which swung between 186.54 and 7.86 token/s with a 54.2% coefficient of variation and a -37.6% trend, ending the window at its minimum; minimax-m3 trended up 61.3% to an 84.97 token/s average.
  • deepseek-v4-flash reported zero samples of its expected 12, contributing to overall coverage of 87.5% with 84 valid points, so its throughput cannot be assessed for this period.
4-hour window · 84 points · 87.5% coverage
  • gemma4:31b delivered the strongest average throughput at 119.03 token/s (peak 172.67 token/s), while nemotron-3-ultra was weakest at 29.94 token/s average, never exceeding 69.25 token/s.
  • minimax-m3 showed the largest swing, climbing 68.5% overall from a 39.84 token/s low at 03:00 to a 140.98 token/s peak at 04:20; glm-5.3-flash was the most volatile model, with a 46.0% coefficient of variation spanning 59.63 to 202.24 token/s.
  • deepseek-v4-flash reported zero samples against 12 expected, contributing to overall coverage of only 87.5% (84 valid points), so fleet-wide comparisons exclude that model entirely.
4-hour window · 84 points · 87.5% coverage
  • gemma4:31b delivered the highest average throughput at 128.05 token/s (peaking at 178.0 token/s), while nemotron-3-ultra was weakest at 23.98 token/s average, never exceeding 40.17 token/s.
  • glm-5.2 showed the most extreme volatility, with a coefficient of variation of 68.5% and swings from 11.36 token/s at 01:40 to 196.21 token/s at 00:20; gemma4:31b also declined from 172.67 token/s at 02:00 to 69.61 token/s at 03:40.
  • deepseek-v4-flash reported zero samples across all 12 expected observations, reducing overall coverage to 87.5% (84 of 96 valid points), so its throughput is unmeasured this period.
4-hour window · 84 points · 87.5% coverage
  • gemma4:31b is the strongest model at 144.54 token/s average throughput, while nemotron-3-ultra is the weakest at 27.73 token/s average.
  • The most operationally significant movement is minimax-m3's sustained decline, falling 49.7% from 170.28 token/s at 23:20 to 39.84 token/s at 03:00; glm-5.2 also shows extreme swings, ranging from 11.36 to 196.21 token/s with a 79.8% coefficient of variation.
  • deepseek-v4-flash returned zero of its 12 expected samples, so its throughput is unknown; overall coverage is 87.5% (84 valid points), which limits conclusions about fleet-wide performance for the period.
4-hour window · 84 points · 87.5% coverage
  • Strongest average throughput was gemma4:31b at 143.06 token/s across 12 samples; weakest was nemotron-3-ultra at 29.81 token/s, roughly a fifth of the leader.
  • The most operationally significant movement is minimax-m3's sustained decline, falling 37.9% from a 173.55 token/s peak at 22:40 to 48.08 token/s at 02:00; glm-5.2 also swung widely, ranging 11.36 to 196.21 token/s with 74.3% coefficient of variation.
  • deepseek-v4-flash reported zero samples against 12 expected, contributing to overall coverage of 87.5% (84 valid points), so its throughput is unmeasured and fleet-wide averages may be incomplete.
4-hour window · 84 points · 87.5% coverage
  • gemma4:31b is the strongest model at 136.43 token/s average throughput, while nemotron-3-ultra is the weakest at 26.97 token/s average, peaking at only 63.1 token/s.
  • glm-5.2 shows the most operationally significant volatility, with a 63.4% coefficient of variation, swinging from 168.57 token/s at 21:20 down to 22.36 token/s at 23:20, then spiking to 196.21 token/s at 00:20 before falling to 34.43 token/s.
  • deepseek-v4-flash reported 0 of 12 expected samples, leaving it entirely unmeasured; overall coverage is 87.5% with 84 valid points, so fleet-wide averages exclude that model.
4-hour window · 84 points · 87.5% coverage
  • Highest average throughput was minimax-m3 at 129.67 token/s (peaking at 173.55 token/s at 22:40), while nemotron-3-ultra was weakest at 26.84 token/s average, never exceeding 63.1 token/s.
  • glm-5.2 showed the most volatility, ranging 22.36 to 168.57 token/s with a 66.3% coefficient of variation, including a one-off spike to 168.57 token/s at 21:20; glm-5.3-flash jumped from 67.47 token/s at 23:00 to 187.53 token/s at 23:20.
  • deepseek-v4-flash reported zero of 12 expected samples (0% coverage), lowering overall coverage to 87.5% with 84 valid points, so its throughput cannot be assessed for this window.
4-hour window · 84 points · 87.5% coverage
  • minimax-m3 is the strongest model by average throughput at 129.96 token/s (range 89.3–173.55 token/s), while nemotron-3-ultra is the weakest at 23.53 token/s average, peaking at only 43.05 token/s.
  • glm-5.2 shows the most operationally significant volatility: throughput collapsed to 8.73 token/s at 20:00, stayed below 33 token/s through 21:00, then spiked to 168.57 token/s at 21:20, yielding a 63.0% coefficient of variation versus 15.2% for minimax-m3.
  • deepseek-v4-flash reported zero samples against 12 expected, contributing to overall coverage of 87.5% (84 valid points), so its throughput is unmeasured and period-wide averages exclude it.
4-hour window · 84 points · 87.5% coverage
  • Strongest average throughput was minimax-m3 at 121.11 token/s (12 of 12 samples, 100% coverage); weakest was nemotron-3-ultra at 17.04 token/s, peaking at only 37.9 token/s.
  • glm-5.2 showed the most operationally significant volatility: a 76.8% coefficient of variation, swings from 8.73 to 168.57 token/s, and a 57.7% upward trend, including a 168.57 token/s spike at 21:20 after sub-30 token/s readings between 20:00 and 20:40.
  • deepseek-v4-flash reported zero samples against 12 expected, leaving its throughput unmeasured; overall dataset coverage is 84 valid points, or 87.5%, so model comparisons exclude that model entirely.
4-hour window · 84 points · 87.5% coverage
  • minimax-m3 is the strongest model at 121.63 token/s average throughput, while nemotron-3-ultra is the weakest at 18.29 token/s average.
  • glm-5.3-flash shows the most operationally significant volatility, ranging from 10.61 to 176.59 token/s with a 58.6% coefficient of variation and a -39.4% trend; glm-5.3 also swung between 41.44 and 159.54 token/s across the window.
  • deepseek-v4-flash reported no samples (0 of 12 expected, 0% coverage), so its throughput is unknown; overall coverage is 87.5%, with 84 of 96 expected points valid.
4-hour window · 84 points · 87.5% coverage
  • minimax-m3 is the strongest model with an average throughput of 118.96 token/s (12 of 12 samples, 100% coverage), while nemotron-3-ultra is the weakest at 19.41 token/s average, peaking at only 37.9 token/s.
  • The most operationally significant movement is glm-5.3-flash, which fell 40.0% over the window to a latest reading of 10.61 token/s, with high volatility (52.3% coefficient of variation) and a maximum of 176.59 token/s.
  • deepseek-v4-flash reported 0 of 12 expected samples (0.0% coverage), so its throughput is unknown; overall dataset coverage is 87.5% with 84 valid points, limiting fleet-wide conclusions.
4-hour window · 84 points · 87.5% coverage
  • gemma4:31b is the strongest model at 122.8 token/s average throughput, ahead of minimax-m3 at 117.98 token/s; nemotron-3-ultra is the weakest at 17.36 token/s average, never exceeding 33.57 token/s.
  • glm-5.3-flash shows the widest volatility, swinging between 51.81 and 176.59 token/s (cv 44.1%), while glm-5.2 declined 19.2% over the window, dipping to 8.66 token/s at 17:40 and closing at 11.08 token/s.
  • deepseek-v4-flash reported no samples (0 of 12, 0.0% coverage), so its throughput is unknown; overall dataset coverage is 87.5% with 84 valid points.
4-hour window · 84 points · 87.5% coverage
  • Strongest average throughput was gemma4:31b at 113.43 token/s, ahead of glm-5.3 (111.99 token/s) and minimax-m3 (109.78 token/s); weakest was nemotron-3-ultra at 17.8 token/s, well below every other reporting model.
  • glm-5.3-flash showed the sharpest swing, rising from a 3.01 token/s low at 14:40 to a 176.59 token/s high at 18:00, a 122.0% trend with 55.6% coefficient of variation; glm-5.2 also ranged from 8.66 to 162.78 token/s.
  • deepseek-v4-flash reported 0 of 12 samples (0.0% coverage), leaving its throughput unknown; overall coverage was 87.5% with 84 valid points, limiting conclusions for that model.
4-hour window · 84 points · 87.5% coverage
  • Highest average throughput was gemma4:31b at 108.26 token/s across 12 samples, peaking at 159.92 token/s; lowest was nemotron-3-ultra at 19.05 token/s, never exceeding 27.42 token/s.
  • glm-5.3-flash showed the sharpest trend, rising 147.3% from roughly 50 token/s early in the window to 151.0–160.09 token/s by 16:40–17:00 UTC, while glm-5.2 was the most volatile at 61.3% coefficient of variation, swinging between 11.12 and 162.78 token/s.
  • deepseek-v4-flash returned 0 of 12 expected samples, so its throughput is unknown; overall coverage is 87.5% (84 of 96 expected points), limiting fleet-wide conclusions for the period.