← Performance dashboard

Hourly performance insights

Summaries of rolling four-hour performance data

1191 retained summaries
4-hour window · 84 points · 87.5% coverage
  • Highest average throughput was gemma4:31b at 108.26 token/s across 12 samples, peaking at 159.92 token/s; lowest was nemotron-3-ultra at 19.05 token/s, never exceeding 27.42 token/s.
  • glm-5.3-flash showed the sharpest trend, rising 147.3% from roughly 50 token/s early in the window to 151.0–160.09 token/s by 16:40–17:00 UTC, while glm-5.2 was the most volatile at 61.3% coefficient of variation, swinging between 11.12 and 162.78 token/s.
  • deepseek-v4-flash returned 0 of 12 expected samples, so its throughput is unknown; overall coverage is 87.5% (84 of 96 expected points), limiting fleet-wide conclusions for the period.
4-hour window · 84 points · 87.5% coverage
  • Highest average throughput was minimax-m3 at 104.41 token/s; lowest was nemotron-3-ultra at 18.53 token/s, with gemma4:31b second-highest at 97.57 token/s.
  • glm-5.2 showed the widest volatility, with a coefficient of variation of 71.6% and throughput swinging between 11.12 and 162.78 token/s; glm-5.3 alternated between peaks near 144 token/s and dips near 27 token/s across the window.
  • deepseek-v4-flash reported no samples (0 of 12 expected, 0% coverage), so its throughput is unknown; overall coverage was 84 of 96 expected points, or 87.5%.
4-hour window · 84 points · 87.5% coverage
  • Highest average throughput was minimax-m3 at 108.03 token/s; lowest was nemotron-3-ultra at 21.77 token/s, with gemma4:31b second-highest at 95.78 token/s.
  • glm-5.2 showed the sharpest swing, rising 68.3% across the window to a 162.78 token/s peak at 14:40 after dropping to 11.12 token/s at 14:00, and it had the highest volatility at 67.1% CV; minimax-m3 declined 20.5% to 35.94 token/s at 15:00.
  • deepseek-v4-flash reported 0 of 12 expected samples (0% coverage), lowering overall coverage to 87.5% with 84 valid points, so no throughput assessment is possible for that model.
4-hour window · 84 points · 87.5% coverage
  • minimax-m3 is the strongest model by average throughput at 114.76 token/s (12 samples, peak 168.84 token/s), while nemotron-3-ultra is the weakest at 32.23 token/s average, never exceeding 77.25 token/s.
  • The most operationally significant movement is nemotron-3-ultra's decline of 50.2 percent, falling from a 77.25 token/s peak at 10:40 to 15.75 token/s at 12:20 and closing at 26.07 token/s; gemma4:31b also swung widely between 12.58 and 173.81 token/s.
  • deepseek-v4-flash has no samples for the entire window (0 of 12 expected, 0.0 percent coverage), and overall coverage is 87.5 percent with 84 of 96 expected points, so its throughput cannot be assessed.
4-hour window · 84 points · 87.5% coverage
  • gemma4:31b is the strongest model by average throughput at 114.14 token/s (p95 165.5 token/s), while nemotron-3-ultra is the weakest at 32.23 token/s average, peaking at only 77.25 token/s.
  • The most operationally significant volatility is the late-window collapse: gemma4:31b fell from 173.81 token/s at 12:20 to 12.58 token/s at 13:00, and glm-5.2 shows the highest relative volatility with a 78.4% coefficient of variation, ranging 11.63 to 139.45 token/s; deepseek-v4-pro is the only clear gainer, trending up 56.0%.
  • deepseek-v4-flash reported zero samples across all 12 expected intervals (0% coverage), leaving its throughput unknown; overall dataset coverage is 87.5% with 84 valid points, so model-level comparisons exclude that model entirely.
4-hour window · 84 points · 87.5% coverage
  • gemma4:31b delivered the strongest average throughput at 128.66 token/s (peak 174.57 token/s), while nemotron-3-ultra was weakest at 35.2 token/s average, never exceeding 77.25 token/s.
  • minimax-m3 showed the sharpest recovery, climbing from a 15.35 token/s low at 09:00 to a 168.84 token/s peak at 11:40, a 96.3% trend gain; glm-5.2 was the most volatile, with a 69.1% coefficient of variation and a drop to 11.63 token/s at 11:00.
  • deepseek-v4-flash reported zero of 12 expected samples (0% coverage), reducing overall coverage to 87.5% with 84 valid points, so its throughput cannot be assessed this period.
4-hour window · 84 points · 87.5% coverage
  • gemma4:31b is the strongest model on average throughput at 142.06 token/s (range 72.03 to 180.87 token/s), while nemotron-3-ultra is the weakest at 31.72 token/s average, peaking at only 77.25 token/s.
  • The most operationally significant volatility is glm-5.2, with a coefficient of variation of 81.9% and swings from 139.45 token/s at 09:20 down to 11.63 token/s at 11:00; minimax-m3 shows the steepest recovery, rising 101.5% overall from 58.69 to 104.63 token/s despite dips to 10.61 token/s at 08:00.
  • deepseek-v4-flash reported 0 of 12 expected samples (0.0% coverage), leaving its throughput unmeasured this period; overall dataset coverage is 87.5% with 84 valid points, so model comparisons exclude it entirely.
4-hour window · 84 points · 87.5% coverage
  • gemma4:31b is the strongest model at 143.68 token/s average throughput, peaking at 180.87 token/s; nemotron-3-ultra is the weakest at 28.07 token/s average, never exceeding 49.47 token/s.
  • glm-5.2 shows the most operationally significant volatility, with a 75.9% coefficient of variation: it spiked to 139.45 token/s at 09:20 UTC, then fell to 20.05 token/s by 10:00 UTC, while deepseek-v4-pro dropped from 105.18 token/s at 08:40 to 14.14 token/s at 10:00.
  • deepseek-v4-flash returned zero of 12 expected samples (0% coverage), lowering overall coverage to 87.5% (84 of 96 points), so its throughput cannot be assessed in this window.
4-hour window · 84 points · 87.5% coverage
  • gemma4:31b is the strongest model at 149.81 token/s average throughput, peaking at 180.87 token/s, while nemotron-3-ultra is the weakest at 32.29 token/s average with a low of 13.47 token/s.
  • minimax-m3 shows the highest volatility at 57.0% coefficient of variation, swinging between 124.9 and 10.61 token/s, and glm-5.3-flash fell 44.5% over the window, dropping to 11.9 token/s at 08:40.
  • deepseek-v4-flash has no samples (0 of 12 expected, 0% coverage), so its throughput is unknown; overall coverage is 87.5% with 84 valid points, limiting conclusions about fleet-wide performance.
4-hour window · 84 points · 87.5% coverage
  • gemma4:31b is the strongest model with an average throughput of 147.07 token/s (range 101.4 to 180.87 token/s), while nemotron-3-ultra is the weakest at 35.14 token/s average, peaking at only 64.47 token/s.
  • The most operationally significant movement is minimax-m3's sustained decline, falling from 162.06 token/s at 04:40 to 10.61 token/s at 08:00, a 60.3% downward trend with 57.6% coefficient of variation, the highest volatility in the dataset.
  • deepseek-v4-flash reported zero samples out of 12 expected, leaving its throughput unmeasured for the full window; overall dataset coverage is 87.5% with 84 valid points, so fleet-wide comparisons exclude that model entirely.
4-hour window · 84 points · 87.5% coverage
  • gemma4:31b is the strongest model at 137.04 token/s average throughput, while nemotron-3-ultra is the weakest at 43.35 token/s average, just below glm-5.2 at 44.13 token/s.
  • minimax-m3 shows the sharpest operational swing, falling from a 162.06 token/s peak at 04:40 to 35.98 token/s at 06:00 and ending at 45.51 token/s, a -38.4% trend; glm-5.3-flash also swung between 196.79 and 53.1 token/s.
  • deepseek-v4-flash has no samples (0 of 12 expected, 0.0% coverage), so its throughput cannot be assessed; overall coverage is 87.5% with 84 valid points across the window.
4-hour window · 84 points · 87.5% coverage
  • gemma4:31b is the strongest model by average throughput at 139.01 token/s, while nemotron-3-ultra is the weakest at 45.51 token/s.
  • minimax-m3 shows the most operationally significant volatility, climbing 94.7% from a low of 16.66 token/s at 03:00 to a peak of 162.06 token/s at 04:40, with a 54.5% coefficient of variation, before falling back to 35.98 token/s at 06:00.
  • deepseek-v4-flash returned no samples (0 of 12 expected), leaving its throughput unmeasured and reducing overall coverage to 87.5% (84 valid points), so comparisons involving that model are not possible for this window.
4-hour window · 84 points · 87.5% coverage
  • gemma4:31b is the strongest model on average throughput at 135.94 token/s (peak 159.32 token/s), while nemotron-3-ultra is the weakest at 48.30 token/s average, never exceeding 93.06 token/s.
  • The most operationally significant movement is minimax-m3, which fell from roughly 97-102 token/s early in the window to a low of 16.66 token/s at 03:00, then recovered to a peak of 162.06 token/s at 04:40, a swing reflected in its 53.1% coefficient of variation and +132.3% trend.
  • deepseek-v4-flash reported zero of its 12 expected samples (0% coverage), so its throughput is unknown; overall dataset coverage is 87.5% with 84 valid points, limiting window-wide comparisons.
4-hour window · 84 points · 87.5% coverage
  • gemma4:31b is the strongest model at 132.87 token/s average throughput, while nemotron-3-ultra is the weakest at 48.8 token/s; glm-5.2 is also low at 68.51 token/s.
  • minimax-m3 shows the most operationally significant volatility: throughput fell from 125.3 token/s at 00:20 to 16.66 token/s at 03:00 before recovering to 124.1 token/s at 04:00, a -33.8% trend with 47.9% coefficient of variation.
  • deepseek-v4-flash reported 0 of 12 expected samples (0% coverage), leaving its throughput unmeasured; overall coverage is 84 of 96 valid points (87.5%), so conclusions exclude that model.
4-hour window · 84 points · 87.5% coverage
  • gemma4:31b is the strongest model by average throughput at 136.8 token/s, while nemotron-3-ultra is the weakest at 55.47 token/s; the other models average between 71.07 and 114.0 token/s.
  • The most operationally significant movement is minimax-m3's sustained decline from 219.82 token/s at 23:20 to 16.66 token/s at 03:00, a -58.3% trend with the highest volatility (60.1% coefficient of variation), ending as the slowest live model.
  • deepseek-v4-flash has no samples for the entire window (0 of 12 expected, 0.0% coverage), so its throughput is unknown; overall coverage is 87.5% with 84 valid points, limiting any comparison involving that model.
4-hour window · 84 points · 87.5% coverage
  • gemma4:31b is the strongest model at 140.9 token/s average throughput, while nemotron-3-ultra is the weakest at 53.59 token/s average.
  • minimax-m3 shows the most operationally significant volatility, with a 40.8% coefficient of variation, swings between 219.82 and 49.99 token/s, and a 35.7% downward trend ending at its minimum of 49.99 token/s.
  • deepseek-v4-flash has no data for the period (0 of 12 samples, 0.0% coverage), leaving overall coverage at 87.5% with 84 valid points, so its throughput cannot be assessed.
4-hour window · 84 points · 87.5% coverage
  • minimax-m3 is the strongest model by average throughput at 143.91 token/s (peaking at 219.82 token/s), while nemotron-3-ultra is the weakest at 51.45 token/s average, below glm-5.2's 60.55 token/s.
  • The most operationally significant movement is deepseek-v4-pro's decline to 12.74 token/s at 01:00, its minimum, with a -21.6% trend over the window; glm-5.3-flash also shows the highest volatility (CV 47.4%), swinging between 46.65 and 187.02 token/s.
  • deepseek-v4-flash returned zero samples (0% coverage), leaving 84 of 96 expected points valid (87.5% overall coverage), so its throughput could not be measured in this period.
4-hour window · 84 points · 87.5% coverage
  • minimax-m3 is the strongest model by average throughput at 153.29 token/s, with a peak of 219.82 token/s, while nemotron-3-ultra is the weakest at 46.49 token/s average, dipping as low as 7.8 token/s.
  • The most operationally significant volatility is glm-5.3-flash's spike to 187.02 token/s at 22:00 UTC against a 78.08 token/s average, and minimax-m3 swung between 65.34 and 219.82 token/s within the window.
  • deepseek-v4-flash reported zero samples of its expected 12, so its throughput is unknown; overall coverage is 87.5% with 84 valid points, limiting window-wide comparisons.
4-hour window · 84 points · 87.5% coverage
  • minimax-m3 is the strongest model by average throughput at 154.07 token/s (peak 197.54 token/s), while nemotron-3-ultra is the weakest at 40.32 token/s average, peaking at only 72.82 token/s.
  • glm-5.3-flash shows the sharpest volatility: it fell from 192.96 token/s at 19:40 to 63.27 token/s at 20:00 and mostly stayed near 60-77 token/s afterward, a -21.2% trend with 55.0% coefficient of variation.
  • deepseek-v4-flash reported 0 of 12 expected samples (0.0% coverage), lowering overall coverage to 87.5% (84 of 96 valid points), so its throughput cannot be assessed for this window.
4-hour window · 84 points · 87.5% coverage
  • minimax-m3 is the strongest model at 153.01 token/s average throughput, while nemotron-3-ultra is the weakest at 40.45 token/s average, with glm-5.2 close behind at 50.50 token/s.
  • The most significant volatility is glm-5.3-flash, with a coefficient of variation of 49.9% and swings between 60.84 and 192.96 token/s; glm-5.3 also collapsed to 15.76 token/s at 19:00 UTC before recovering to 151.61 token/s by 22:00.
  • deepseek-v4-flash reported 0 of 12 expected samples (0% coverage), reducing overall coverage to 87.5% with 84 valid points, so its throughput is entirely unmeasured in this window.
4-hour window · 84 points · 87.5% coverage
  • minimax-m3 is the strongest model with an average throughput of 149.17 token/s (range 128.46–188.76 token/s), while nemotron-3-ultra is the weakest at 35.8 token/s average, ending at just 7.8 token/s.
  • The most significant volatility is in gemma4:31b, which swings between 40.02 and 175.55 token/s (CV 40.0%), and glm-5.3, which dropped to 15.76 token/s at 19:00 before recovering to 124.12 token/s by 21:00.
  • deepseek-v4-flash has no samples (0 of 12 expected, 0.0% coverage), so its throughput cannot be assessed; overall dataset coverage is 87.5% (84 valid points).
4-hour window · 84 points · 87.5% coverage
  • minimax-m3 is the strongest model with an average throughput of 149.7 token/s (range 110.56 to 188.76 token/s), while nemotron-3-ultra is the weakest at 34.38 token/s average (range 10.96 to 63.31 token/s).
  • The most operationally significant volatility is glm-5.3-flash, which held roughly 60 to 73 token/s for the first six intervals, then swung to 192.96 token/s at 19:40 before falling to 63.27 token/s at 20:00; its coefficient of variation is 50.2 percent. glm-5.3 also dropped to 15.76 token/s at 19:00.
  • deepseek-v4-flash reported zero samples out of 12 expected, so its throughput is unknown; overall coverage is 84 of 96 expected points, or 87.5 percent, limiting conclusions about fleet-wide performance.
4-hour window · 84 points · 87.5% coverage
  • gemma4:31b delivered the highest average throughput at 115.24 token/s, ahead of minimax-m3 (111.35 token/s) and glm-5.3 (105.84 token/s); nemotron-3-ultra was weakest at 23.01 token/s average, never exceeding 52.24 token/s.
  • minimax-m3 showed the sharpest trend, rising 117% from 79.55 token/s at 14:20 to a 188.76 token/s peak at 17:20; gemma4:31b was the most volatile, swinging between 40.02 and 180.26 token/s with a 44.6% coefficient of variation.
  • deepseek-v4-flash reported zero of 12 expected samples (0% coverage), leaving 84 valid points and 87.5% overall coverage, so its throughput cannot be assessed for this period.
4-hour window · 84 points · 87.5% coverage
  • gemma4:31b is the strongest model with an average throughput of 124.04 token/s, while nemotron-3-ultra is the weakest at 28.39 token/s, roughly a quarter of the leader's pace.
  • The most operationally significant volatility is gemma4:31b's collapse from 159.64 token/s at 14:20 to 58.4 token/s at 15:00, a drop of about 63%, followed by recovery to 180.26 token/s at 16:40; nemotron-3-ultra also shows the highest relative volatility with a 57.7% coefficient of variation.
  • deepseek-v4-flash reported 0 of 12 expected samples (0% coverage), leaving overall coverage at 87.5% with 84 valid points, so its throughput cannot be assessed for this window.
4-hour window · 84 points · 87.5% coverage
  • Strongest average throughput was gemma4:31b at 118.55 token/s (peak 159.64 token/s); weakest was nemotron-3-ultra at 30.84 token/s, ending the window at 9.85 token/s.
  • The most operationally significant event was gemma4:31b dropping from 159.64 token/s at 14:20 to 58.4 token/s at 15:00, driving its -40.5% trend; glm-5.3 also swung between 184.48 and 61.44 token/s with 41.0% coefficient of variation.
  • deepseek-v4-flash reported 0 of 12 expected samples (0% coverage), lowering overall coverage to 87.5% (84 of 96 valid points), so its throughput cannot be assessed for this window.
4-hour window · 84 points · 87.5% coverage
  • glm-5.3 is the strongest model with an average throughput of 134.86 token/s, narrowly ahead of gemma4:31b at 129.99 token/s, while nemotron-3-ultra is the weakest at 30.87 token/s average, peaking at only 59.11 token/s.
  • The most operationally significant event is gemma4:31b's late collapse: it held 143-160 token/s from 12:20 through 14:20, then fell to 59.94 and 58.4 token/s in the final two observations, a drop of roughly 63% from its 14:20 level.
  • deepseek-v4-flash reported zero of 12 expected samples (0% coverage), reducing overall valid points to 84 of 96 (87.5% coverage), so its throughput cannot be assessed in this window.
4-hour window · 84 points · 87.5% coverage
  • glm-5.3 is the strongest model by average throughput at 135.62 token/s, narrowly ahead of gemma4:31b at 135.35 token/s; nemotron-3-ultra is the weakest at 33.66 token/s average, with a low of 4.41 token/s.
  • glm-5.3 shows the most operationally significant volatility, swinging between 59.94 and 184.48 token/s with a 36.3% coefficient of variation, including sharp drops at 12:00 and 13:00; minimax-m3 also spiked from 58.87 to 167.43 token/s within 40 minutes.
  • deepseek-v4-flash reported zero of 12 expected samples, reducing overall coverage to 87.5% (84 valid points), so its throughput is unmeasured this period.
4-hour window · 84 points · 87.5% coverage
  • gemma4:31b is the strongest model at 127.72 token/s average throughput (p95 165.36 token/s), while nemotron-3-ultra is the weakest at 28.03 token/s average, never exceeding 55.27 token/s.
  • glm-5.3 shows the most operationally significant volatility, with a 44.1% coefficient of variation and swings between 59.94 and 184.48 token/s; its +50.4% trend includes a drop from 176.83 token/s at 11:40 to 73.39 token/s at 12:00 before recovering to 184.48 token/s at 12:20.
  • deepseek-v4-flash has no data (0 of 12 samples, 0.0% coverage), lowering overall coverage to 87.5% with 84 valid points, so its throughput cannot be assessed in this window.
4-hour window · 84 points · 87.5% coverage
  • gemma4:31b is the strongest model with average throughput of 118.33 token/s, while nemotron-3-ultra is weakest at 30.93 token/s; glm-5.3 follows closely at 117.44 token/s.
  • glm-5.3 shows the most operationally significant volatility, swinging between 59.94 and 176.83 token/s with a 41.4% coefficient of variation and a +26.6% trend, including a final-interval drop from 176.83 to 73.39 token/s; nemotron-3-ultra also ranges from 4.41 to 75.22 token/s.
  • deepseek-v4-flash reported zero samples against 12 expected, so its throughput is unknown; overall coverage is 87.5% with 84 valid points, and any model comparison excludes it.
4-hour window · 84 points · 87.5% coverage
  • gemma4:31b is the strongest model by average throughput at 117.8 token/s (12 of 12 samples, 100% coverage), while nemotron-3-ultra is the weakest at 30.16 token/s, with a minimum of 4.41 token/s and a maximum of 75.22 token/s.
  • glm-5.3 shows the most operationally significant volatility, swinging between 59.94 and 171.09 token/s with a coefficient of variation of 41.9% and a trend of -23.4%; nemotron-3-ultra is even less stable at 68.1% coefficient of variation.
  • deepseek-v4-flash has no samples (0 of 12 expected, 0% coverage), so its throughput is unknown; overall coverage is 87.5% with 84 valid points, limiting fleet-wide conclusions.
4-hour window · 84 points · 87.5% coverage
  • gemma4:31b is the strongest model at 127.04 token/s average throughput, while nemotron-3-ultra is the weakest at 34.86 token/s average, with samples ranging from 5.56 to 83.91 token/s.
  • glm-5.3 shows the most operationally significant volatility, alternating between roughly 65-80 token/s and 166-171 token/s across the window, with a 42.8% coefficient of variation; nemotron-3-ultra is also unstable, dipping to 5.56 token/s at 07:40.
  • deepseek-v4-flash reported no samples (0 of 12 expected, 0.0% coverage), reducing overall coverage to 87.5% (84 of 96 expected points), so its throughput cannot be assessed for this period.
4-hour window · 84 points · 87.5% coverage
  • Strongest average throughput was gemma4:31b at 135.85 token/s (peaking at 178.41 token/s), while the weakest was nemotron-3-ultra at 41.14 token/s, which also hit the lowest single reading of 5.56 token/s.
  • The most operationally significant volatility is glm-5.3, which alternates between roughly 170 token/s and roughly 75 token/s every 20 minutes, with a stddev of 51.05 token/s and cv of 45.7%; nemotron-3-ultra also swung from 83.91 token/s down to 5.56 token/s within the window.
  • deepseek-v4-flash has zero samples (0 of 12 expected, 0% coverage), so its throughput is unknown; overall coverage is 87.5% with 84 valid points, limiting fleet-wide conclusions.
4-hour window · 84 points · 87.5% coverage
  • gemma4:31b is the strongest model by average throughput at 139.1 token/s (peak 178.41 token/s), while nemotron-3-ultra is the weakest at 42.86 token/s average, never exceeding 83.91 token/s.
  • glm-5.3-flash shows the most operationally significant volatility, with a coefficient of variation of 52.2% and swings from 186.31 token/s at 04:40 down to 13.76 token/s at 06:20; glm-5.3 also oscillates sharply between roughly 28 and 176 token/s.
  • deepseek-v4-flash has zero samples out of 12 expected (0% coverage), reducing overall valid points to 84 of 96 (87.5% coverage), so its throughput cannot be assessed for this window.
4-hour window · 84 points · 87.5% coverage
  • gemma4:31b is the strongest model by average throughput at 138.39 token/s (peak 176.04 token/s), while nemotron-3-ultra is the weakest at 59.01 token/s average, with a low of 24.6 token/s.
  • glm-5.3 shows the most operationally significant volatility, with a coefficient of variation of 45.4% and a swing from 176.62 token/s at 05:00 to 28.76 token/s at 06:00; glm-5.3-flash also spiked to 186.31 token/s before falling back to 61.32 token/s.
  • deepseek-v4-flash reported zero samples across all 12 expected intervals, reducing overall coverage to 87.5% (84 of 96 valid points), so its throughput is unmeasured in this window.
4-hour window · 84 points · 87.5% coverage
  • Strongest average throughput was gemma4:31b at 134.3 token/s (peak 161.46 token/s), while glm-5.2 was weakest at 58.35 token/s average, dipping as low as 29.48 token/s.
  • glm-5.3-flash showed the sharpest volatility, with a 57.3% trend gain and 45.8% coefficient of variation, spiking from roughly 75 token/s to 186.31 token/s at 04:40 before collapsing to 71.57 token/s at 05:00; glm-5.3 swung between 180.64 and 63.42 token/s.
  • deepseek-v4-flash reported zero samples against 12 expected, contributing to overall coverage of 87.5% (84 valid points), so its throughput is unmeasured and fleet-wide averages may be incomplete.
4-hour window · 84 points · 87.5% coverage
  • gemma4:31b is the strongest model at 132.51 token/s average throughput, peaking at 160.05 token/s, while nemotron-3-ultra is the weakest at 60.83 token/s average, with a low of 33.46 token/s.
  • glm-5.3 shows the most operationally significant volatility, with a 47.2% coefficient of variation and swings between 63.42 and 180.64 token/s, including a drop from 180.64 to 74.29 token/s between 02:40 and 03:00 UTC.
  • deepseek-v4-flash reported no samples (0 of 12 expected), leaving its throughput unknown; overall coverage is 87.5% with 84 valid points, so conclusions exclude that model.
4-hour window · 84 points · 87.5% coverage
  • gemma4:31b delivered the highest average throughput at 123.08 token/s (peak 165.03 token/s), while glm-5.2 was the weakest at 61.36 token/s average, dipping to a low of 26.91 token/s at 00:00 UTC.
  • glm-5.3 showed the most operationally significant volatility, with a 101.9% upward trend and 48.1% coefficient of variation, swinging between 60.89 and 180.64 token/s across the window.
  • deepseek-v4-flash reported zero samples (0% coverage), so its throughput is unknown; overall coverage was 87.5% with 84 of 96 expected points, limiting window-wide comparisons.
4-hour window · 84 points · 87.5% coverage
  • gemma4:31b posted the highest average throughput at 121.02 token/s (peak 166.22 token/s), while nemotron-3-ultra was the weakest at 55.70 token/s average and never exceeded 75.15 token/s.
  • glm-5.3-flash showed the most operationally significant volatility, spiking to 201.76 token/s at 22:40 and 200.73 token/s at 23:00 before settling at 76.60 token/s by 02:00, a -36.9% trend with 55.3% coefficient of variation; glm-5.3 likewise swung between 60.89 and 178.38 token/s.
  • deepseek-v4-flash returned zero of 12 expected samples (0% coverage), leaving 84 valid points overall (87.5% coverage), so its throughput cannot be evaluated for this window.
4-hour window · 84 points · 87.5% coverage
  • gemma4:31b is the strongest model at 115.25 token/s average throughput, ahead of deepseek-v4-pro at 105.64 token/s; glm-5.2 is the weakest at 55.65 token/s, just below nemotron-3-ultra at 59.02 token/s.
  • glm-5.3-flash shows the most volatility, with a coefficient of variation of 58.5% and swings from 201.76 token/s at 22:40 down to 53.22 token/s at 23:40; glm-5.3 also declined 42.3% over the window, from 165.72 token/s at 21:20 to 79.05 token/s at 01:00.
  • deepseek-v4-flash has no samples (0 of 12 expected, 0% coverage), so its throughput cannot be assessed; overall coverage is 87.5% with 84 of 96 expected samples.
4-hour window · 84 points · 87.5% coverage
  • gemma4:31b is the strongest model by average throughput at 122.75 token/s (peak 181.21 token/s), while glm-5.2 is the weakest at 45.66 token/s average, never exceeding 106.35 token/s.
  • The most operationally significant movement is glm-5.3's decline of 37.6%, falling from a 187.11 token/s peak at 20:40 to roughly 60-73 token/s for the final two hours; glm-5.3-flash also swung from 50.96 token/s to 201.76 token/s and back within hours.
  • deepseek-v4-flash reported zero of 12 expected samples (0% coverage), leaving its throughput unknown; overall the dataset holds 84 valid points against 96 expected, or 87.5% coverage.
4-hour window · 84 points · 87.5% coverage
  • glm-5.3 posted the strongest average throughput at 129.17 token/s, peaking at 187.11 token/s, while glm-5.2 was the weakest at 40.30 token/s average and never exceeded 75.38 token/s.
  • The most operationally significant shift is glm-5.3-flash, which jumped from 75.79 token/s at 22:20 to 201.76 token/s at 22:40 and held 200.73 token/s at 23:00, a 63.9% trend with 55.8% coefficient of variation; nemotron-3-ultra also swung between 6.65 and 100.71 token/s.
  • deepseek-v4-flash reported 0 of 12 expected samples (0% coverage), leaving overall coverage at 87.5% with 84 valid points, so no throughput assessment is possible for that model this period.
4-hour window · 84 points · 87.5% coverage
  • glm-5.3 is the strongest model by average throughput at 134.15 token/s, narrowly ahead of gemma4:31b at 133.03 token/s, while nemotron-3-ultra is the weakest at 39.55 token/s and ended the window at just 6.65 token/s.
  • glm-5.2 shows the most operationally significant volatility, with a 73.4% coefficient of variation, a one-off spike to 169.82 token/s at 18:00 against a 53.0 token/s average, and a -43.9% trend; glm-5.3 also swung between 64.1 and 187.11 token/s.
  • deepseek-v4-flash reported 0 of 12 expected samples (0.0% coverage), so overall coverage is 87.5% with 84 valid points, and no throughput comparison is possible for that model.
4-hour window · 84 points · 87.5% coverage
  • gemma4:31b posted the highest average throughput at 135.82 token/s, while nemotron-3-ultra was the weakest at 41.17 token/s; glm-5.3 ranked second at 125.49 token/s.
  • glm-5.2 showed the most operationally significant volatility, with a coefficient of variation of 83.1% and swings between 28.58 and 232.61 token/s across the window; glm-5.3-flash also declined 49.6% from its early peak of 188.36 token/s.
  • deepseek-v4-flash reported zero samples against 12 expected, so its throughput is unknown; overall coverage was 87.5% with 84 valid points, limiting fleet-wide conclusions.
4-hour window · 84 points · 87.5% coverage
  • gemma4:31b is the strongest model by average throughput at 147.61 token/s (p95 170.36 token/s), while nemotron-3-ultra is the weakest at 38.16 token/s average, peaking at only 74.59 token/s.
  • glm-5.2 shows the most operationally significant volatility, swinging between 28.58 and 232.61 token/s with a 74.1% coefficient of variation and a -39.2% trend; glm-5.3-flash also fell -31.3% to 54.69 token/s by 19:00 UTC.
  • deepseek-v4-flash reported 0 of 12 expected samples (0% coverage), leaving its throughput unknown; overall dataset coverage is 87.5% with 84 of 96 expected points, so fleet-wide comparisons are incomplete.
4-hour window · 84 points · 87.5% coverage
  • gemma4:31b is the strongest model by average throughput at 148.72 token/s (p95 166.18 token/s), while nemotron-3-ultra is the weakest at 34.45 token/s average, never exceeding 74.59 token/s.
  • glm-5.2 shows the most volatility, with a coefficient of variation of 70.5% and swings between 28.58 and 232.61 token/s within the four hours; glm-5.3-flash posted the steepest trend at +131.3%, rising from 48.62 token/s at 14:20 to 188.36 token/s at 16:40.
  • deepseek-v4-flash reported zero of 12 expected samples (0.0% coverage), so no throughput can be assessed for it; overall coverage is 87.5% with 84 valid points, limiting fleet-wide comparison.
4-hour window · 84 points · 87.5% coverage
  • gemma4:31b is the strongest model at 145.45 token/s average throughput, while nemotron-3-ultra is the weakest at 30.88 token/s average, never exceeding 62.0 token/s.
  • glm-5.3-flash shows the sharpest shift, climbing from 59.8 token/s at 16:00 to 188.36 token/s at 16:40 (97.9% trend); glm-5.2 is the most volatile, ranging from 31.39 to 232.61 token/s with a 63.9% coefficient of variation.
  • deepseek-v4-flash reported 0 of 12 samples (0% coverage), so its throughput is unknown; overall coverage is 87.5% with 84 valid points.
4-hour window · 84 points · 87.5% coverage
  • Highest average throughput was gemma4:31b at 129.74 token/s (peaking at 164.92 token/s), while nemotron-3-ultra was weakest at 28.24 token/s average, ranging from 6.84 to 58.48 token/s.
  • glm-5.2 showed the most volatility, with a coefficient of variation of 60.1% and swings from 31.39 token/s at 14:20 to 194.35 token/s at 15:20; glm-5.3 declined 25.0% over the window, ending at 59.85 token/s.
  • deepseek-v4-flash reported 0 of 12 expected samples, leaving its throughput unknown and reducing overall coverage to 87.5% (84 of 96 valid points), so any model comparison excludes it entirely.
4-hour window · 84 points · 87.5% coverage
  • glm-5.3 is the strongest model with average throughput of 128.43 token/s (p95 178.85 token/s), while nemotron-3-ultra is the weakest at 25.75 token/s average, never exceeding 58.48 token/s in any observation.
  • glm-5.2 shows the most operationally significant volatility, with a coefficient of variation of 61.4% and swings from 5.89 token/s at 12:00 to 181.99 token/s at 15:00; gemma4:31b also dipped to 58.26 token/s at 13:00 before recovering to 157.22 token/s at 14:20.
  • deepseek-v4-flash has no samples (0 of 12 expected), so its throughput is unknown; overall coverage is 87.5% (84 of 96 valid points), limiting conclusions about fleet-wide performance.
4-hour window · 84 points · 87.5% coverage
  • glm-5.3 is the strongest model by average throughput at 123.58 token/s (peak 181.11 token/s), while nemotron-3-ultra is the weakest at 32.30 token/s average, never exceeding 59.43 token/s.
  • The most significant volatility comes from glm-5.3-flash, which fell from 172.07 token/s at 10:20 to 67.85 token/s at 14:00, a 25.8% decline, with a 48.9% coefficient of variation; nemotron-3-ultra swung between 5.24 and 59.43 token/s.
  • deepseek-v4-flash reported 0 of 12 expected samples (0% coverage), and overall coverage is 87.5% with 84 valid points, so its throughput cannot be assessed this period.
4-hour window · 84 points · 87.5% coverage
  • Strongest average throughput is gemma4:31b at 114.49 token/s (peak 163.4 token/s); weakest is nemotron-3-ultra at 34.40 token/s, dipping as low as 5.24 token/s.
  • The most significant trend is glm-5.3 rising 59.3 percent, from 65.63 token/s at 10:20 to a 181.11 token/s peak at 12:40; glm-5.3-flash shows the sharpest volatility, with a 51.8 percent coefficient of variation and a swing from 172.07 token/s down to 27.66 token/s.
  • deepseek-v4-flash returned zero samples (0 percent coverage), leaving overall coverage at 87.5 percent with 84 of 96 expected points, so its throughput is unmeasured this window.