← Performance dashboard

Hourly performance insights

Summaries of rolling four-hour performance data

1200 retained summaries
4-hour window · 56 points · 58.3% coverage
  • gemma4:31b is the strongest model at 155.75 token/s average throughput, peaking at 189.38 token/s, while nemotron-3-ultra is the weakest at 9.56 token/s average and just 4.81 token/s in its latest sample.
  • glm-5.3 shows the most operationally significant volatility: it fell from 195.89 token/s at 00:20 to roughly 60-69 token/s across 00:40-01:40 before rebounding to 171.36 token/s, a 50.9% coefficient of variation. deepseek-v4-flash similarly swung between 209.95 and 53.9 token/s.
  • Every model recorded only 7 of 12 expected samples, yielding 58.3% coverage and 56 valid points overall, so the 22:03-23:40 UTC window is absent and averages may not represent the full four hours.
4-hour window · 8 points · 8.3% coverage
  • Strongest average throughput is deepseek-v4-flash at 209.95 token/s, followed by glm-5.3 at 183.7 token/s; the weakest is nemotron-3-ultra at 19.47 token/s, roughly eleven times slower than the leader, with minimax-m3 also low at 52.9 token/s.
  • No trend or volatility can be assessed: every model has only a single observation at 00:00 UTC, so stddev_tps and cv_pct are 0.0 and trend_pct is 0.0 for all eight models; the only operational signal is the wide spread between models, from 19.47 to 209.95 token/s.
  • Data coverage is severely limited: each model has 1 sample against 12 expected samples, or 8.3% coverage, with a total of 8 valid points, so all averages, minima, maxima, and p95 values rest on one measurement each.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3-flash is the strongest model at 143.89 token/s average throughput, slightly ahead of gemma4:31b at 141.11 token/s; nemotron-3-ultra is the weakest at 15.08 token/s average, never exceeding 46.23 token/s.
  • glm-5.3 shows the most operationally significant shift, rising 72.0% across the window from roughly 61-77 token/s before 06:00 to 168.8-179.96 token/s afterward; nemotron-3-ultra is the most volatile, with a 78.6% coefficient of variation and swings between 1.41 and 46.23 token/s.
  • No missing-data limitation applies: all eight models report 12 of 12 samples, 96 valid points, and 100.0% coverage, so the four-hour window is complete.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model with an average throughput of 147.1 token/s (peak 189.36 token/s), while nemotron-3-ultra is the weakest at 13.26 token/s average, dipping as low as 1.41 token/s.
  • The most operationally significant volatility is nemotron-3-ultra, with a 91.1% coefficient of variation and a swing from 1.41 to 46.23 token/s; glm-5.3 also shows instability, dropping from 178.87 token/s at 02:20 to a 62.69 token/s floor mid-window, and deepseek-v4-flash trends upward 29.0%.
  • No missing-data limitation applies: all 96 expected points are valid, every model has 12 of 12 samples, and coverage is 100.0% across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model by average throughput at 149.78 token/s (peak 202.96 token/s), while nemotron-3-ultra is the weakest at 10.78 token/s, roughly one-fourteenth of gemma4:31b's average.
  • The most operationally significant movement is nemotron-3-ultra's recovery: it climbed from a 1.41 token/s low at 03:40 to 46.23 token/s at 05:00, with 115.9% coefficient of variation. glm-5.3 also spiked to 178.87 token/s at 02:20 before settling near 63.64 token/s.
  • No missing-data limitation applies: all 96 expected points are valid (100.0% coverage), and every model has 12 of 12 samples, so gaps do not affect these figures.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 147.02 token/s average throughput, peaking at 202.96 token/s, while nemotron-3-ultra is the weakest at 4.83 token/s average, never exceeding 7.89 token/s.
  • The most operationally significant movement is deepseek-v4-pro's 53.2% upward trend, climbing from 22.18 token/s at 00:20 to sustained 109.93-120.16 token/s late in the window; glm-5.3 shows the highest volatility at 54.2% coefficient of variation, swinging between 58.52 and 217.6 token/s.
  • No missing-data limitation exists: all eight models report 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 140.57 token/s average throughput, narrowly ahead of glm-5.3-flash at 139.99 token/s; nemotron-3-ultra is the weakest at 6.52 token/s average, never exceeding 14.47 token/s.
  • deepseek-v4-pro shows the most operationally significant volatility, with a 113.2% upward trend, 66.5% coefficient of variation, and swings between 15.02 and 122.82 token/s; glm-5.3 is also erratic, ranging 58.52 to 217.6 token/s with 55.8% coefficient of variation.
  • No missing-data limitation applies: all eight models report 12 of 12 expected samples, totaling 96 valid points at 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3-flash is the strongest model at 139.8 token/s average throughput, ahead of gemma4:31b at 127.05 token/s; nemotron-3-ultra is the weakest at 8.05 token/s average, never exceeding 15.74 token/s.
  • deepseek-v4-flash shows the sharpest trend, up 87.4% from 22.73 token/s at 22:20 to a 202.33 token/s peak at 00:40, while deepseek-v4-pro is the most volatile at 61.3% CV, dipping to 15.02 token/s at 23:20; glm-5.2's latest reading fell to 10.55 token/s against a 98.58 token/s maximum.
  • No missing-data limitation applies: all 8 models have 12 of 12 samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3-flash is the strongest model with an average throughput of 133.81 token/s (p95 168.69 token/s), while nemotron-3-ultra is the weakest at 9.52 token/s average, peaking at only 15.74 token/s across the window.
  • deepseek-v4-pro shows the most operationally significant volatility: coefficient of variation 73.8%, swinging between 121.16 token/s and 15.02 token/s, with a -44.9% trend ending at 91.5 token/s; deepseek-v4-flash conversely rose 60.4% to a 202.33 token/s maximum.
  • No missing-data limitation exists: all eight models report 12 of 12 expected samples, with 96 valid points and 100.0% coverage over the four-hour period, so the figures reflect complete observations.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3-flash is the strongest model, averaging 126.73 token/s with a p95 of 168.69 token/s, while nemotron-3-ultra is the weakest at 9.75 token/s average, never exceeding 15.74 token/s.
  • glm-5.3 shows the widest operational swings, ranging from 58.26 to 198.77 token/s with a coefficient of variation of 52.5%; deepseek-v4-pro also swung between 15.02 and 121.16 token/s within the four-hour window.
  • No missing-data limitation applies: all eight models delivered 12 of 12 expected samples, yielding 96 valid points and 100.0% coverage across the 20:03 to 00:03 UTC period.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b delivered the highest average throughput at 126.86 token/s, while nemotron-3-ultra was weakest at 8.04 token/s; glm-5.3-flash ranked second at 113.13 token/s.
  • glm-5.2 was the most volatile model, with a 68.2% coefficient of variation and swings from 5.17 to 198.55 token/s; deepseek-v4-pro fell from roughly 105 token/s to a 21.63 token/s low near 22:00 before recovering to 117.36 token/s, and nemotron-3-ultra rose 163% to 15.74 token/s.
  • No missing-data limitation applies: all 8 models recorded 12 of 12 expected samples, yielding 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3-flash is the strongest model at 117.72 token/s average throughput, ahead of gemma4:31b at 114.96 token/s; nemotron-3-ultra is the weakest at 5.56 token/s average, peaking at only 14.75 token/s.
  • The most operationally significant movement is deepseek-v4-pro's decline of 36.9 percent, falling from a 124.57 token/s peak at 20:00 to 21.63 token/s at 22:00. glm-5.2 also swung sharply, dropping to 5.17 token/s at 21:00 before spiking to 198.55 token/s at 21:20.
  • No missing-data limitation exists: all 8 models have 12 of 12 expected samples, with 96 valid points and 100.0 percent coverage across the four-hour window.
4-hour window · 95 points · 99.0% coverage
  • Highest average throughput was gemma4:31b at 122.12 token/s, with a p95 of 157.83 token/s; lowest was nemotron-3-ultra at 3.33 token/s average, peaking at only 7.34 token/s.
  • deepseek-v4-flash showed the widest volatility, with a coefficient of variation of 65.6% and swings between 21.75 and 165.31 token/s, including drops to 25.22 and 21.75 token/s at 20:00 and 20:20; glm-5.2 also fell to 5.17 token/s at 21:00.
  • deepseek-v4-pro is missing one observation (11 of 12 samples, 91.7% coverage, no 17:40 point), leaving 95 of 96 valid points overall (99.0% coverage).
4-hour window · 71 points · 74.0% coverage
  • Strongest average throughput is gemma4:31b at 118.93 token/s (peak 160.04 token/s); weakest is nemotron-3-ultra at 2.58 token/s, never exceeding 4.21 token/s across its nine samples.
  • deepseek-v4-flash shows the greatest volatility (CV 68.2%), swinging between 24.33 and 165.31 token/s, and falling to 25.22 token/s at 20:00 right after reaching 165.31 token/s at 19:00; glm-5.3 also spiked to 160.9 token/s at 20:00 from roughly 60 token/s.
  • The dataset holds 71 of 96 expected points (74.0% coverage); deepseek-v4-pro has only 8 of 12 samples (66.7%), and no model reports before 17:20, leaving the early window unobserved.
4-hour window · 47 points · 49.0% coverage
  • glm-5.3-flash is the strongest model at 116.21 token/s average throughput, ahead of gemma4:31b at 100.61 token/s; nemotron-3-ultra is the weakest at 2.22 token/s average, with a maximum of only 3.84 token/s.
  • deepseek-v4-flash shows the most operationally significant volatility, ranging from 24.33 to 165.31 token/s with a 70.1% coefficient of variation and a +253.6% trend, closing at its 165.31 token/s peak; gemma4:31b also fell from 154.57 to 35.49 token/s (-49.5%).
  • Only 47 of 72 expected samples (49.0% coverage) are present; each model has 5 or 6 of 12 expected observations, with no data before 17:20 UTC, so averages may not represent the full four-hour window.
4-hour window · 23 points · 24.0% coverage
  • gemma4:31b is the strongest model at 133.74 token/s average throughput, peaking at 154.57 token/s; nemotron-3-ultra is the weakest at 2.21 token/s average, never exceeding 3.02 token/s.
  • glm-5.2 shows the sharpest decline, falling from 79.82 to 22.79 token/s (-58.7%), while glm-5.3-flash nearly doubled from 55.55 to 124.82 token/s (+95.5%); deepseek-v4-pro rose 30.4% to 106.11 token/s.
  • Coverage is only 24.0%: 23 valid samples out of 96 expected across eight models, with all observations clustered between 17:20 and 18:00 UTC and deepseek-v4-pro contributing just 2 samples (16.7% coverage), leaving the earlier hours unmeasured.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model with a 138.66 token/s average throughput (p95 192.93 token/s, maximum 219.57 token/s), while nemotron-3-ultra is the weakest at 2.43 token/s average (p95 4.42 token/s, maximum 6.63 token/s).
  • The most operationally significant pattern is nemotron-3-ultra's sustained decline, trending down 43.1% across the window to a 1.17 token/s minimum at 15:10 UTC with 46.3% coefficient of variation; glm-5.3 also showed sharp swings between 16.24 and 219.57 token/s.
  • No missing-data limitation applies: all eight models report 48 of 48 expected samples, 384 valid points total, and 100.0% coverage across the 12:03 to 16:03 UTC window.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model with an average of 136.51 token/s (p95 192.32 token/s, max 222.35 token/s), while nemotron-3-ultra is the weakest at an average of 2.66 token/s, peaking at only 6.63 token/s.
  • glm-5.3 shows the most operationally significant volatility: despite the highest average, it dropped to 16.24 token/s at 14:05 and 45.69 token/s at 11:40, with a 33.1% coefficient of variation and a -12.9% trend; nemotron-3-ultra also declined 22.6% to 1.3 token/s by 15:00.
  • No missing-data limitation exists: all eight models have 48 of 48 expected samples, 384 valid points, and 100.0% coverage over the four-hour window.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model with an average of 144.30 token/s (peak 222.35 token/s at 11:25 UTC), while nemotron-3-ultra is the weakest at 3.28 token/s average, roughly 44 times slower.
  • All eight models show negative trends, steepest for glm-5.3 (-18.3%) and glm-5.2 (-17.8%); a synchronized dip at 13:05 UTC dropped glm-5.3 to 34.43 token/s and glm-5.2 to 13.32 token/s, and nemotron-3-ultra shows the highest volatility with a CV of 42.6%.
  • No missing-data limitation: every model has 48 of 48 samples with 100.0% coverage, and the dataset totals 384 valid points across the four-hour window.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model at 149.79 token/s average throughput (peak 222.35 token/s, p95 198.57 token/s), while nemotron-3-ultra is the weakest at 3.85 token/s average (peak 7.37 token/s), roughly 39x slower than the leader.
  • The most significant movement is nemotron-3-ultra's decline of 36.3% over the window, with the highest volatility (CV 44.0%) and a drop from a 7.37 token/s peak to 1.46 token/s at 11:20; glm-5.2 also fell 20.6% to a 37.75 token/s latest reading.
  • No missing data: all eight models report 48 of 48 samples (384 valid points, 100.0% coverage), so the only limitation is that results reflect this single four-hour window.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model at 147.21 token/s average throughput, peaking at 222.35 token/s at 11:25, while nemotron-3-ultra is the weakest at 4.62 token/s average and never exceeded 12.79 token/s.
  • The most operationally significant movement is nemotron-3-ultra's 40.2% decline, dropping from roughly 6-7 token/s before 10:00 to between 1.46 and 3.3 token/s in the final two hours; glm-5.3-flash conversely trended up 35.2% to a 97.12 token/s average.
  • No missing-data limitation applies: all eight models report 48 of 48 expected samples, 384 valid points in total, and 100.0% coverage across the 08:03 to 12:03 UTC window.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model at 142.36 token/s average throughput (p95 191.5 token/s, maximum 233.42 token/s), while nemotron-3-ultra is the weakest at 5.64 token/s average, never exceeding 12.79 token/s.
  • The most operationally significant volatility is glm-5.3's repeated sharp collapses: 233.42 token/s at 07:05 dropped to 60.9 token/s at 07:10, and it fell to a 23.83 token/s minimum at 07:40, driving a 33.7% coefficient of variation; glm-5.2 shows the highest relative volatility at 43.8%.
  • No missing-data limitation applies: all 8 models recorded 48 of 48 expected samples, totaling 384 valid points at 100.0% coverage across the four-hour window.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model by average throughput at 147.17 token/s (p95 179.18 token/s), while nemotron-3-ultra is the weakest at 6.07 token/s average, peaking at only 12.79 token/s.
  • The most operationally significant volatility is glm-5.2, with the highest coefficient of variation at 46.7% and swings from 11.21 token/s at 06:45 to 168.54 token/s at 07:05; glm-5.3 also dropped from 233.42 token/s at 07:05 to 60.9 token/s at 07:10 and hit 23.83 token/s at 07:40.
  • No missing-data limitation applies: all eight models report 48 of 48 expected samples, with 384 valid points and 100.0% coverage over the four-hour window.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model with an average throughput of 155.99 token/s (p95 of 181.82 token/s), while nemotron-3-ultra is the weakest at 5.72 token/s average, roughly 27 times slower.
  • glm-5.3 showed the sharpest operational volatility: it spiked to 233.42 token/s at 07:05 UTC, fell to 60.9 token/s at 07:10, and hit a period minimum of 23.83 token/s at 07:40; glm-5.2 had the highest relative variability at 48.3% CV.
  • No missing-data limitation applies: all eight models reported 48 of 48 expected samples, with 384 valid points and 100.0% coverage across the four-hour window.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model at 162.71 token/s average throughput, peaking at 187.42 token/s, while nemotron-3-ultra is the weakest at 5.40 token/s average, roughly 30 times slower.
  • glm-5.3 is also the most stable (CV 12.3%), whereas gemma4:31b swings between 37.67 and 186.13 token/s (CV 29.6%); minimax-m3 shows a -16.3% trend with a 112.27 token/s spike at 05:00 UTC.
  • No missing data: all eight models report 48 of 48 samples with 100.0% coverage (384 valid points), so the only limitation is the four-hour window ending 07:03 UTC.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model, averaging 163.33 token/s (p95 181.72 token/s, max 187.42 token/s), while nemotron-3-ultra is the weakest at 4.79 token/s average, peaking at only 11.34 token/s.
  • The most operationally significant volatility is gemma4:31b, which swings from a 187.82 token/s maximum down to a 37.67 token/s minimum (cv 28.6%) with a -17.3% trend; deepseek-v4-pro also dipped to 21.01 token/s at 03:50 before recovering.
  • No missing-data limitation exists: all eight models report 48 of 48 expected samples, 384 valid points, and 100.0% coverage, though the window covers only four hours.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 posted the strongest average throughput at 160.71 token/s, peaking at 187.42 token/s, while nemotron-3-ultra was weakest at 3.85 token/s average and never exceeded 8.57 token/s.
  • The most operationally significant volatility was gemma4:31b's late-window collapse: between 04:00 and 04:25 it fell from 65.36 token/s to a 37.67 token/s minimum against its 144.11 token/s average, driving a -16.6% trend; deepseek-v4-pro also dipped to 21.01 token/s at 03:50.
  • No missing-data limitation applies: all eight models report 48 of 48 expected samples, 384 valid points total, and 100.0% coverage, so every five-minute interval from 01:05 to 05:00 UTC is represented.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model at 163.46 token/s average throughput, ahead of gemma4:31b at 146.94 token/s, while nemotron-3-ultra is the weakest at 3.31 token/s average and never exceeded 6.32 token/s.
  • minimax-m3 shows the most operationally significant pattern: a -25.1% trend with 40.4% coefficient of variation, falling from 130.0 token/s at 00:05 to a 17.22 token/s minimum at 03:25; deepseek-v4-pro also degraded late, reaching 21.01 token/s at 03:50 before recovering to 96.24 token/s.
  • No missing-data limitation applies: all eight models reported 48 of 48 expected samples with 100.0% coverage across the four-hour window.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model at 163.5 token/s average throughput, ahead of gemma4:31b at 145.74 token/s; nemotron-3-ultra is the weakest at 4.41 token/s average, with a maximum of only 15.71 token/s.
  • The most significant volatility is glm-5.2's drop to 6.37 token/s at 00:25 UTC, recovering to 81.8 token/s by 00:45; minimax-m3 shows the steepest sustained decline, trending -23.9% and falling from a 130.0 token/s spike at 00:05 to 17.98 token/s at 02:30.
  • No missing-data limitation applies: all eight models report 48 of 48 expected samples, 384 valid points, and 100.0% coverage, so the four-hour window is fully represented.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model at 158.82 token/s average throughput (p95 181.59 token/s, maximum 185.08 token/s), while nemotron-3-ultra is the weakest at 5.73 token/s average, never exceeding 15.74 token/s.
  • The most operationally significant pattern is nemotron-3-ultra's sustained decline: from roughly 7-15 token/s before 23:00 it fell to 1.67-4.64 token/s after 00:15, a 64.8% trend drop, and it shows the highest volatility at 62.2% cv.
  • No missing-data limitation applies: all eight models recorded 48 of 48 expected samples, yielding 384 valid points and 100.0% coverage across the four-hour window.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model at 161.20 token/s average throughput (p95 183.11 token/s), while nemotron-3-ultra is the weakest at 7.38 token/s average, peaking at only 17.73 token/s.
  • The most significant movement is glm-5.3-flash climbing 37.6% to an 86.93 token/s average, ending near 110 token/s, while nemotron-3-ultra fell 38.8% to 2.13 token/s; glm-5.2 shows the sharpest volatility (cv 44.1%), dipping to 4.74 token/s at 22:55, and minimax-m3 spiked to 130.0 token/s at 00:05 against a 50.88 token/s average.
  • No missing-data limitation applies: all eight models report 48 of 48 samples with 100.0% coverage and 384 valid points.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model with an average throughput of 156.22 token/s (peaking at 194.9 token/s), ahead of gemma4:31b at 143.42 token/s; nemotron-3-ultra is the weakest at 7.86 token/s average, never exceeding 17.73 token/s.
  • glm-5.3-flash shows the most operationally significant trend, climbing 34.5% over the window from roughly 50-60 token/s early on to 120.64-125.55 token/s in the final three observations; glm-5.2 is the most volatile, swinging between 4.74 and 95.24 token/s with a 42.3% coefficient of variation.
  • No missing-data limitation applies: all eight models report 48 of 48 expected samples, 384 valid points total, and 100.0% coverage across the four-hour window.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model at 155.26 token/s average throughput (maximum 194.9 token/s), while nemotron-3-ultra is the weakest at 6.57 token/s average, never exceeding 17.73 token/s.
  • glm-5.2 shows the most operationally significant volatility, with a 39.2% coefficient of variation and repeated collapses from roughly 70 token/s down to 4.74 token/s at 22:55 and 8.22 token/s at 20:05; nemotron-3-ultra also rose 129.3% over the window but remained far below other models.
  • No missing-data limitation is present: all eight models report 48 of 48 expected samples with 100.0% coverage and 384 valid points, though the dataset spans only this four-hour window.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 delivered the highest average throughput at 148.95 token/s (peaking at 194.90 token/s), while nemotron-3-ultra was weakest at 5.27 token/s average, roughly 28 times slower than the leader.
  • glm-5.3-flash showed the steepest decline, trending down 20.9% from an early 120.40 token/s peak to a 42.11 token/s low, while nemotron-3-ultra was highly volatile (71.0% CV), swinging between 1.54 and 17.73 token/s with a late spike to 17.54 token/s at 21:55.
  • No missing-data limitation applies: all eight models recorded 48 of 48 expected samples, 100.0% coverage, and 384 valid points across the four-hour window.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model with an average throughput of 147.54 token/s (p95 180.18 token/s), while nemotron-3-ultra is the weakest at 3.77 token/s average, roughly 39 times slower than the leader.
  • The most operationally significant volatility is in glm-5.2, which repeatedly collapses from steady 60-80 token/s levels to sharp drops such as 8.22 token/s at 20:05 and 13.8 token/s at 20:40, with a 32.0% coefficient of variation. glm-5.3-flash also shows a -13.6% trend, ending at its 42.11 token/s minimum.
  • No missing-data limitation applies: all eight models report 48 of 48 expected samples with 100.0% coverage, and the dataset contains 384 valid points, so the four-hour window is fully populated.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 posted the strongest average throughput at 143.32 token/s (p95 177.22 token/s), while nemotron-3-ultra was weakest at 3.18 token/s average and never exceeded 4.55 token/s.
  • glm-5.2 was the most volatile model (cv 46.9%), opening the window between 1.55 and 27.57 token/s from 16:05 to 16:40 before recovering to mostly 50-87 token/s; glm-5.3 also logged abrupt single-interval drops to 31.06, 46.25, and 55.12 token/s.
  • No missing-data limitation applies: all eight models reported 48 of 48 expected samples, 384 valid points total, and 100.0% coverage across the four-hour window.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model at 136.18 token/s average throughput, peaking at 181.93 token/s; nemotron-3-ultra is the weakest at 3.02 token/s average, never exceeding 4.55 token/s.
  • glm-5.2 shows the most operationally significant volatility, swinging from a 1.55 token/s low near 16:35 to an 86.94 token/s high at 18:55, with a 57.3% coefficient of variation and a 109.0% trend over the window.
  • No missing-data limitation applies: all eight models report 48 of 48 expected samples, 100.0% coverage, and 384 valid points, so the four-hour window is fully represented.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model with an average throughput of 132.86 token/s (peak 181.93 token/s, p95 176.54 token/s), while nemotron-3-ultra is the weakest at 2.62 token/s average, never exceeding 4.55 token/s.
  • glm-5.2 shows the most operationally significant volatility: between 15:50 and 16:40 it dropped from 14.73 token/s to a low of 1.55 token/s before recovering to 75.42 token/s by 17:05, driving its 57.7% coefficient of variation; glm-5.3 also rose 23.1% across the window.
  • No missing-data limitation applies: all eight models report 48 of 48 expected samples with 100.0% coverage, so the four-hour window is fully represented.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model at 125.51 token/s average throughput, peaking at 174.03 token/s, while nemotron-3-ultra is the weakest at 2.18 token/s average, roughly 58 times slower.
  • The most operationally significant volatility is glm-5.2's collapse between 15:50 and 16:40 UTC, when throughput fell from 14.73 token/s to a low of 1.55 token/s before recovering to 67.78 token/s; its 58.3% coefficient of variation and -42.7% trend far exceed every other model.
  • No missing-data limitation applies: all eight models report 48 of 48 expected samples, 100.0% coverage, and 384 valid points across the four-hour window.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model at 130.49 token/s average throughput (peak 173.13 token/s), while nemotron-3-ultra is the weakest at 2.11 token/s average, never exceeding 3.77 token/s in the window.
  • glm-5.2 shows the most operationally significant decline, trending -28.9% and ending at 5.76 token/s at 16:00 versus its 80.8 token/s maximum; glm-5.3 also dipped sharply to 28.11 token/s at 14:05 before recovering to 124.28 token/s.
  • No missing-data limitation applies: all eight models report 48 of 48 expected samples, 384 valid points total, and 100.0% coverage across the four-hour period.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model with an average throughput of 133.95 token/s (peak 183.58 token/s, p95 172.51 token/s), while nemotron-3-ultra is the weakest at 2.21 token/s average, never exceeding 3.43 token/s.
  • The most operationally significant volatility is glm-5.3's sharp collapse between 14:00 and 14:10 UTC, dropping from 161.34 token/s at 13:55 to 28.11 token/s at 14:05 before recovering to 131.53 token/s at 14:15; its overall trend is -20.9%. gemma4:31b also swings widely, from 37.51 to 160.23 token/s with a 28.5% coefficient of variation.
  • There is no missing-data limitation in this window: all eight models report 48 of 48 expected samples at 100.0% coverage (384 valid points), though the four-hour span limits longer-term assessment.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model, averaging 142.1 token/s with a peak of 183.58 token/s, while nemotron-3-ultra is the weakest, averaging 2.51 token/s and never exceeding 3.76 token/s.
  • The most operationally significant movement is nemotron-3-ultra's steady decline, down 32.0% across the window to a latest 2.02 token/s and a low of 1.23 token/s; glm-5.3 also dropped sharply at 14:00 to 62.57 token/s, its minimum for the period.
  • No missing-data limitation applies: all eight models report 48 of 48 expected samples, 384 valid points, and 100.0% coverage, so the four-hour window is fully observed.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model by average throughput at 141.47 token/s, with a peak of 183.58 token/s and p95 of 173.07 token/s; nemotron-3-ultra is the weakest at 2.84 token/s average, never exceeding 3.92 token/s across the window.
  • glm-5.2 shows the highest volatility at 33.9% CV (stddev 17.12 token/s), swinging between 12.18 and 78.4 token/s, while nemotron-3-ultra declined 18.2% over the four hours, ending at 2.9 token/s.
  • No missing-data limitation applies: all eight models report 48 of 48 expected samples with 100% coverage, totaling 384 valid points, so gaps do not affect these figures.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model at 136.42 token/s average throughput (peak 183.58 token/s), while nemotron-3-ultra is the weakest at 3.23 token/s average, never exceeding 4.91 token/s.
  • glm-5.3-flash shows the steepest improvement, trending up 39.8% and climbing from a 23.02 token/s low to a 129.83 token/s high; glm-5.2 is the most volatile at 35.9% coefficient of variation, swinging between 12.18 and 79.83 token/s.
  • No missing-data limitation applies: all eight models report 48 of 48 expected samples, 384 valid points total, and 100.0% coverage across the four-hour window.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model, averaging 134.73 token/s with a p95 of 169.32 token/s, while nemotron-3-ultra is the weakest at 3.42 token/s average, roughly 39 times slower.
  • The most operationally significant volatility is minimax-m3's collapse to 6.02 token/s at 09:35 followed by recovery to 58.58 token/s at 09:45; glm-5.2 shows the highest relative variability (cv 34.0%), swinging between 18.14 and 79.83 token/s.
  • No missing-data limitation applies: all eight models report 48 of 48 expected samples with 100.0% coverage (384 valid points), so the four-hour window is fully represented.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model with an average throughput of 133.86 token/s (peaking at 183.75 token/s), while nemotron-3-ultra is the weakest at 3.69 token/s average, never exceeding 7.01 token/s.
  • The most operationally significant volatility is minimax-m3's collapse to 6.02 token/s at 09:35 followed by a rebound to 58.58 token/s at 09:45; glm-5.3 also dropped to 42.32 token/s at 09:50 before recovering to 145.28 token/s by 10:00.
  • No missing-data limitation applies: all eight models report 48 of 48 samples with 100.0% coverage (384 valid points), so the four-hour window is fully observed.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model with an average throughput of 143.06 token/s (p95 174.33 token/s), while nemotron-3-ultra is the weakest at 4.32 token/s average, never exceeding 8.74 token/s.
  • Throughput declined over the window for seven of eight models, with nemotron-3-ultra down 24.6% and minimax-m3 down 16.3%; deepseek-v4-flash was the only gainer at +10.5%. Volatility was highest for nemotron-3-ultra (CV 35.8%) and glm-5.2 (CV 29.4%), which dipped to 18.65 token/s at 07:30.
  • No missing-data limitation applies: all eight models recorded 48 of 48 expected samples, coverage is 100.0%, and valid_point_count is 384, so the four-hour window is fully observed.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model with an average throughput of 142.28 token/s (p95 of 179.37 token/s), while nemotron-3-ultra is the weakest at 4.84 token/s average, never exceeding 9.57 token/s.
  • The most operationally significant movement is nemotron-3-ultra's sustained decline, down 32.5% over the window to a latest reading of 4.76 token/s with dips to 2.03 token/s; glm-5.3 also swings sharply, ranging from 36.92 to 183.75 token/s.
  • No missing-data limitation applies: all eight models report 48 of 48 expected samples, 384 valid points overall, and 100.0% coverage across the four-hour period.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 delivered the strongest average throughput at 146.77 token/s (peaking at 183.75 token/s), while nemotron-3-ultra was weakest at 5.44 token/s, roughly 27 times lower.
  • The most significant movement is deepseek-v4-flash's downward trend of -13.7%, ending at 63.75 token/s against a 90.64 token/s average; glm-5.2 showed the highest volatility with a coefficient of variation of 31.9% and a range of 7.02 to 84.5 token/s.
  • No missing-data limitation exists in this window: all eight models report 48 of 48 samples at 100.0% coverage, totaling 384 valid points, though the four-hour span limits broader conclusions.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model at 152.93 token/s average throughput (p95 180.97 token/s), while nemotron-3-ultra is the weakest at 5.45 token/s average, peaking at only 9.57 token/s.
  • glm-5.2 shows the highest volatility (CV 34.8%), dropping to 7.02 token/s at 03:35 and 10.8 token/s at 04:20; glm-5.3 also dipped sharply to 36.92 token/s at 04:20. deepseek-v4-pro was the steadiest performer (CV 16.1%, stddev 16.85 token/s).
  • No missing-data limitation applies: all eight models recorded 48 of 48 expected samples, 100.0% coverage, and 384 valid points across the four-hour window.