← Performance dashboard

Hourly performance insights

Summaries of rolling four-hour performance data

1193 retained summaries
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 124.79 token/s average throughput; nemotron-3-ultra is the weakest at 37.77 token/s average, with a minimum of 3.19 token/s.
  • glm-5.3 shows the steepest decline, trending -32.4% from 142.52 token/s at 21:20 to 92.24 token/s at 01:00, while nemotron-3-ultra is the most volatile (cv 81.9%, range 3.19–112.61 token/s); deepseek-v4-flash rose 27.5% to 120.54 token/s.
  • No missing-data limitation: all 8 models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0%.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput: gemma4:31b at 105.73 token/s (peak 167.44 token/s); weakest: nemotron-3-ultra at 31.70 token/s, roughly a third of the leader.
  • Most operationally significant volatility: nemotron-3-ultra swings between 3.19 and 112.61 token/s with 99.1% coefficient of variation, while deepseek-v4-flash shows the steepest climb, a 77.8% trend from 45.08 to 120.38 token/s.
  • No missing-data limitation: all eight models report 12 of 12 samples with 100.0% coverage, and the dataset's 96 valid points match the expected count.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 delivered the highest average throughput at 99.77 token/s, ahead of gemma4:31b at 90.56 token/s, while nemotron-3-ultra was weakest at 27.28 token/s, far below deepseek-v4-flash at 65.72 token/s.
  • nemotron-3-ultra showed extreme volatility, ranging from 3.19 to 112.61 token/s with a coefficient of variation of 116.2%, and glm-5.3-flash swung from a peak of 188.17 token/s at 21:40 to 11.77 token/s at 22:00.
  • No missing-data limitation applies: all 96 expected points are valid, coverage is 100.0%, and every model has 12 of 12 samples, so the four-hour window is fully populated.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3-flash is the strongest model on average throughput at 103.79 token/s, with a peak of 203.49 token/s, while nemotron-3-ultra is the weakest at 16.53 token/s average, never exceeding 40.67 token/s.
  • The most operationally significant volatility is glm-5.3-flash, which swings between 11.77 and 203.49 token/s with a 59.6% coefficient of variation, and its final 22:00 reading collapsed to 11.77 token/s; nemotron-3-ultra is even less stable at 87.6% CV, oscillating between 3.19 and 40.67 token/s.
  • No missing-data limitation exists in this window: all 8 models report 12 of 12 expected samples, giving 96 valid points and 100.0% coverage over the four-hour period.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput is minimax-m3 at 102.6 token/s across 12 samples (range 65.34 to 166.2 token/s); weakest is nemotron-3-ultra at 15.69 token/s (range 3.05 to 56.32 token/s).
  • glm-5.3-flash shows the most operationally significant volatility: after holding roughly 45 to 71 token/s through 18:20, it spiked to 187.86 token/s at 18:40 and 203.49 token/s at 19:00, ending at 150.42 token/s, a 93% trend with 58.8% coefficient of variation; nemotron-3-ultra is similarly erratic at 100.1%.
  • No missing-data limitation applies: all eight models report 12 of 12 expected samples, 96 valid points total, and 100% coverage over the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • Highest average throughput was minimax-m3 at 110.17 token/s; lowest was nemotron-3-ultra at 20.18 token/s, roughly one-fifth of the leader.
  • glm-5.3-flash showed the sharpest swing, falling to 45.78 token/s at 17:20 before climbing to 203.49 token/s at 19:00; nemotron-3-ultra was the most volatile overall, with a coefficient of variation of 85.9% and readings ranging from 3.05 to 56.32 token/s.
  • No missing-data limitation applies: all eight models delivered 12 of 12 expected samples, giving 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • minimax-m3 is the strongest model by average throughput at 109.03 token/s, ahead of gemma4:31b at 95.7 token/s; nemotron-3-ultra is the weakest at 16.36 token/s, with a minimum of 1.36 token/s.
  • The most operationally significant volatility is nemotron-3-ultra's coefficient of variation of 104.4%, ranging from 1.36 to 56.32 token/s; deepseek-v4-pro shows the steepest trend at +74.2%, recovering from 10.59 token/s at 16:20 to 126.46 token/s at 17:40.
  • No missing-data limitation exists: all eight models have 12 of 12 expected samples, 96 valid points, and 100.0% coverage over the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput was minimax-m3 at 105.91 token/s; weakest was nemotron-3-ultra at 13.12 token/s, with glm-5.2 also low at 40.04 token/s.
  • Volatility is the main operational concern: nemotron-3-ultra swung between 1.36 and 39.27 token/s (105.0% coefficient of variation), and deepseek-v4-pro ranged 10.59 to 102.45 token/s (69.1% CV), while glm-5.3-flash was steadist at 32.3% CV.
  • No missing-data limitation exists: all 8 models delivered 12 of 12 expected samples, giving 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • minimax-m3 is the strongest model by average throughput at 96.04 token/s, slightly ahead of gemma4:31b at 92.81 token/s, while nemotron-3-ultra is the weakest at 15.65 token/s average, including a low of 1.36 token/s.
  • glm-5.3 shows the most operationally significant volatility, with a 53.5% coefficient of variation, falling from 144.61 token/s at 12:20 UTC to 5.17 token/s at 14:00 UTC before recovering to 142.01 token/s at 16:00 UTC; deepseek-v4-pro is similarly erratic at 65.8% CV.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, 96 valid points, and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 104.78 token/s average throughput, peaking at 157.31 token/s; nemotron-3-ultra is the weakest at 16.89 token/s average, never exceeding 54.66 token/s.
  • glm-5.3 shows the most operationally significant volatility, falling from 162.0 token/s at 11:20 to 5.17 token/s at 14:00, a -57.5% trend with 55.2% coefficient of variation; deepseek-v4-pro is similarly erratic at 71.8% CV, ranging from 3.86 to 83.41 token/s.
  • No missing-data limitation applies: all eight models recorded 12 of 12 expected samples, yielding 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput was glm-5.3 at 106.17 token/s, narrowly ahead of gemma4:31b at 104.77 token/s; weakest was nemotron-3-ultra at 16.82 token/s, with a maximum of only 54.66 token/s.
  • The most significant volatility is glm-5.3's late collapse: from a 162.0 token/s peak at 11:20 to 5.17 token/s at 14:00, a -40.3% trend. nemotron-3-ultra (cv 88.0%) and deepseek-v4-pro (cv 59.7%, dipping to 3.86 token/s at 12:00) also swung erratically.
  • No missing-data limitation exists: all 8 models delivered 12 of 12 samples, giving 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model by average throughput at 113.08 token/s, ahead of glm-5.3 at 108.38 token/s; nemotron-3-ultra is the weakest at 16.58 token/s, with a maximum of only 54.66 token/s.
  • The most operationally significant volatility is nemotron-3-ultra's coefficient of variation of 85.2%, swinging between 2.55 and 54.66 token/s, while deepseek-v4-pro shows the steepest decline at -33.9% trend, dipping to 3.86 token/s at 12:00 UTC.
  • No missing-data limitation exists: all eight models have 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 123.55 token/s average throughput, peaking at 171.75 token/s, while nemotron-3-ultra is the weakest at 18.44 token/s average, never exceeding 54.66 token/s.
  • Volatility is the dominant operational signal: nemotron-3-ultra swings between 2.55 and 54.66 token/s (89.7% CV), and deepseek-v4-pro ends at 3.86 token/s against a 112.86 token/s maximum, while glm-5.3 shows the strongest upward trend at +46.3%, closing at 151.78 token/s.
  • No missing-data limitation exists: all eight models report 12 of 12 samples, 96 valid points, and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model by average throughput at 116.16 token/s, well above glm-5.3 at 97.27 token/s, while nemotron-3-ultra is the weakest at 20.79 token/s, far below the next-lowest deepseek-v4-flash at 53.09 token/s.
  • The most operationally significant movement is nemotron-3-ultra's sustained decline, falling 68.4% from 68.29 token/s at 07:20 to 2.55 token/s at 11:00, with volatility of 102.9% coefficient of variation; deepseek-v4-flash also dropped 41.4%, ending at 9.18 token/s.
  • No missing-data limitation exists: all eight models report 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 128.18 token/s average throughput, well ahead of glm-5.3 at 92.89 token/s; nemotron-3-ultra is weakest at 21.22 token/s, with a minimum of 0.64 token/s.
  • deepseek-v4-flash shows the steepest decline, trending -50.5% and falling from a 113.78 token/s peak at 07:20 to 59.57 token/s at 10:00; nemotron-3-ultra is the most volatile, with a 100.0% coefficient of variation and swings between 0.64 and 68.29 token/s.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0%, though the window covers only four hours.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 131.01 token/s average throughput, peaking at 171.75 token/s; nemotron-3-ultra is the weakest at 26.09 token/s average, with a low of 0.64 token/s.
  • deepseek-v4-pro shows the sharpest operational volatility, collapsing from 118.32 token/s at 07:20 to 10.99 token/s at 08:00 before recovering to 112.86 token/s at 08:20; nemotron-3-ultra is also erratic, swinging between 0.64 and 68.29 token/s with a 90.5% coefficient of variation.
  • No missing-data limitation applies: all eight models report 12 of 12 samples, 96 valid points, and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 134.51 token/s average throughput, with a p95 of 167.04 token/s; nemotron-3-ultra is weakest at 26.77 token/s average, never exceeding 58.45 token/s.
  • nemotron-3-ultra shows the most operationally significant volatility: 80.6% coefficient of variation, a -37.6% trend, and swings from 0.64 to 58.45 token/s, ending at 12.96 token/s; glm-5.2 is the only clear gainer, rising 34.4% to 107.88 token/s.
  • No missing-data limitation applies: all eight models report 12 of 12 samples, overall coverage is 100.0% with 96 valid points, so the four-hour window is complete.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 127.58 token/s average throughput (peak 169.76 token/s), while nemotron-3-ultra is the weakest at 38.83 token/s average, with a low of 3.05 token/s.
  • The most operationally significant volatility is the synchronized drop at 06:00 UTC: minimax-m3 fell from 102.64 to 40.59 token/s, glm-5.2 from 91.31 to 20.99 token/s, and nemotron-3-ultra from 50.56 to 3.05 token/s; nemotron-3-ultra also shows the highest variability at 54.4% CV.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, with 96 valid points and 100.0% coverage over the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model by average throughput at 125.41 token/s, ahead of glm-5.3 at 106.70 token/s; nemotron-3-ultra is the weakest at 42.71 token/s, below glm-5.2 at 54.51 token/s.
  • The most significant volatility is deepseek-v4-flash collapsing to 15.26 token/s at 03:40 before recovering to 127.27 token/s at 04:40, a -32.3% trend overall; glm-5.3-flash fell -37.3% from a 191.77 token/s peak, and nemotron-3-ultra shows the highest variability at 50.0% CV.
  • No missing-data limitation exists: all eight models have 12 of 12 samples, with 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 120.43 token/s average throughput, peaking at 163.48 token/s, while nemotron-3-ultra is the weakest at 39.47 token/s average, dipping to 6.36 token/s.
  • The most operationally significant pattern is late-window degradation: glm-5.2 fell 26.5% to 26.81 token/s at 04:00, and deepseek-v4-flash dropped from 152.97 token/s at 00:40 to 43.87 token/s at 04:00; nemotron-3-ultra also shows the highest volatility at 54.7% coefficient of variation.
  • No missing-data limitation applies: all eight models report 12 of 12 expected samples, 96 valid points, and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput was gemma4:31b at 121.4 token/s (p95 175.26 token/s), while nemotron-3-ultra was weakest at 42.8 token/s, never exceeding 79.54 token/s in any observation.
  • The most significant volatility is nemotron-3-ultra with a 50.2% coefficient of variation, swinging between 6.36 and 79.54 token/s; glm-5.3-flash shows the sharpest trend, rising 66.1% and spiking to 191.77 token/s at 01:20, and deepseek-v4-pro dipped to 17.26 token/s at 00:00.
  • No missing-data limitation exists in this window: all eight models delivered 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour period.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model on average throughput at 136.71 token/s (peak 181.94 token/s), while nemotron-3-ultra is the weakest at 36.20 token/s average, roughly a quarter of gemma4:31b's rate.
  • Seven of eight models show negative trends, with glm-5.2 down 29.1% and minimax-m3 down 25.9%; glm-5.3-flash is the exception, up 63.1% after spiking to 191.77 token/s at 01:20, and nemotron-3-ultra shows the highest volatility with a 60.3% coefficient of variation.
  • No missing-data limitation exists: all eight models have 12 of 12 samples, 96 valid points, and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput was gemma4:31b at 131.79 token/s; weakest was nemotron-3-ultra at 33.95 token/s, roughly a quarter of the leader's rate.
  • The most significant shift was glm-5.3-flash, which fell from 190.66 token/s at 21:40 to 61.95 token/s at 01:00 (trend -30.4%), while glm-5.3 trended up 24.6% to a 103.86 token/s average; nemotron-3-ultra was the most volatile, with a 76.5% coefficient of variation and dips to 4.06 token/s.
  • No missing-data limitation applies: all 8 models reported 12 of 12 expected samples, giving 96 valid points and 100.0% coverage over the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b posted the highest average throughput at 132.5 token/s (peak 181.94 token/s), while nemotron-3-ultra was weakest at 35.23 token/s average and never exceeded 93.69 token/s.
  • glm-5.3 showed the sharpest operational shift, jumping from 18.13 token/s at 22:20 to 176.53 token/s at 22:40 and staying above 142 token/s through 00:00 (trend +113.3%); in contrast, glm-5.3-flash collapsed from 190.66 token/s at 21:40 to 43.76 token/s at 22:00.
  • No missing-data limitation applies: all 8 models recorded 12 of 12 expected samples, totaling 96 valid points at 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b delivered the highest average throughput at 130.75 token/s, while nemotron-3-ultra was the weakest at 40.07 token/s, less than a third of the leader and closing the window at just 4.06 token/s.
  • The most operationally significant volatility came from glm-5.3-flash (CV 57.9%), which swung from 190.66 token/s at 21:40 to 43.76 token/s at 22:00; glm-3.3 similarly dropped to 18.13 token/s at 22:20 before recovering to 176.53 token/s.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, 96 valid points, and 100.0% coverage, so the four-hour window is complete.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model by average throughput at 129.26 token/s, ahead of deepseek-v4-flash at 124.50 token/s; nemotron-3-ultra is the weakest at 35.09 token/s, including a low of 0.98 token/s.
  • glm-5.3-flash shows the most extreme volatility (CV 51.7%), ranging from 43.76 to 190.66 token/s, with a late run near 184-191 token/s followed by a drop to 43.76 token/s; nemotron-3-ultra is also unstable (CV 102.0%), swinging between 0.98 and 95.88 token/s.
  • No missing-data limitation applies: all 8 models report 12 of 12 samples, 96 valid points overall, and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model by average throughput at 128.84 token/s, narrowly ahead of gemma4:31b at 126.83 token/s; nemotron-3-ultra is the weakest at 26.72 token/s.
  • The most operationally significant volatility is nemotron-3-ultra, with a coefficient of variation of 120.5% and readings ranging from 0.98 to 95.88 token/s; glm-5.3-flash is also unstable at 52.2%, spiking to 183.8 token/s at 21:00 from a 35.53 token/s low.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, and the dataset reports 96 valid points with 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model at 134.54 token/s average throughput, slightly ahead of gemma4:31b at 130.09 token/s; nemotron-3-ultra is the weakest at 19.78 token/s average, with a low of 0.98 token/s.
  • The most operationally significant volatility is nemotron-3-ultra, swinging from 0.98 to 95.88 token/s (CV 130.5%); glm-5.3-flash and deepseek-v4-pro also vary widely (CV 46.2% and 45.7%), while deepseek-v4-flash is steadiest at CV 13.0%.
  • No missing-data limitation exists: all eight models report 12 of 12 samples, 96 valid points, and 100.0% coverage over the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput was deepseek-v4-flash at 130.97 token/s (peaking at 166.55 token/s at 17:20), while the weakest was nemotron-3-ultra at 10.26 token/s, never exceeding 30.88 token/s and dipping to 0.98 token/s.
  • The most significant volatility came from nemotron-3-ultra (coefficient of variation 80.1%) and glm-5.3, which fell 31.8% from a 173.24 token/s peak at 16:40 to 60.67 token/s at 18:40; deepseek-v4-pro also dropped to 15.65 token/s at 16:40.
  • No missing-data limitation applies: all 8 models delivered 12 of 12 expected samples, yielding 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash delivered the highest average throughput at 126.11 token/s, peaking at 166.55 token/s, while nemotron-3-ultra was the weakest model at 14.39 token/s average and never exceeded 48.43 token/s.
  • deepseek-v4-pro showed the sharpest operational volatility, with a 47.5% coefficient of variation and a -42.2% trend, collapsing from 113.05 token/s at 14:40 to 15.65 token/s at 16:40; nemotron-3-ultra was even less stable at 86.1% CV.
  • No missing-data limitation applies: all eight models report 12 of 12 samples, and the dataset records 96 valid points with 100.0% coverage.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b delivered the highest average throughput at 137.57 token/s (peaking at 179.5 token/s), while nemotron-3-ultra was the weakest at 21.15 token/s average, ending at 6.22 token/s.
  • glm-5.3 showed the sharpest positive trend, up 65.9% from a 59.44 token/s low to a sustained 165.88–173.24 token/s range between 15:40 and 16:40 UTC; nemotron-3-ultra was the most volatile (cv 81.2%), falling 58.1% to 6.22 token/s, and deepseek-v4-pro dipped to 15.65 token/s at 16:40.
  • No missing-data limitation applies: all eight models recorded 12 of 12 expected samples, yielding 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b delivered the highest average throughput at 130.78 token/s, peaking at 179.5 token/s, while nemotron-3-ultra was weakest at 21.36 token/s average and never exceeded 55.92 token/s.
  • glm-5.3 showed the most operationally significant volatility, with a coefficient of variation of 45.4% and swings between 56.14 and 184.53 token/s; nemotron-3-ultra was even less stable at 79.7% CV, oscillating between 6.1 and 55.92 token/s across the window.
  • No missing-data limitation applies: all eight models reported 12 of 12 expected samples, giving 96 valid points and 100.0% coverage, so every twenty-minute interval from 12:20 to 16:00 UTC is represented.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b is the strongest model at 127.18 token/s average throughput, ahead of deepseek-v4-flash at 117.65 token/s; nemotron-3-ultra is the weakest at 19.45 token/s average, dipping as low as 3.48 token/s.
  • glm-5.3 shows the sharpest operational swing: it peaked at 184.53 token/s at 12:40, then dropped to 56.14 token/s at 13:00 and stayed volatile, ending at 80.39 token/s with a -40.9% trend; nemotron-3-ultra is the least stable, with an 88.9% coefficient of variation.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, and the dataset reports 96 valid points with 100.0% coverage.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model at 128.77 token/s average throughput, peaking at 210.67 token/s, while nemotron-3-ultra is the weakest at 17.01 token/s average, never exceeding 55.92 token/s.
  • The most operationally significant movement is glm-5.3's decline of 20.4%, falling from a 184.53 token/s peak at 12:40 to 59.44 token/s at 14:00; nemotron-3-ultra also shows extreme volatility with a 93.4% coefficient of variation.
  • No missing-data limitation applies: all eight models delivered 12 of 12 expected samples, and the dataset reports 96 valid points with 100.0% coverage.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model at 135.83 token/s average throughput (peak 210.67 token/s), while nemotron-3-ultra is the weakest at 19.39 token/s average, peaking at only 73.62 token/s.
  • The most operationally significant volatility is nemotron-3-ultra, with a 107.0% coefficient of variation, swings between 3.48 and 73.62 token/s, and a -69.4% trend ending at 6.83 token/s; glm-5.3 also swung from 184.53 to 56.14 token/s in the final interval.
  • No missing-data limitation applies: all eight models report 12 of 12 expected samples, with 100.0% coverage across the 96 valid points.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput was deepseek-v4-flash at 142.26 token/s (peaking at 210.67 token/s), while nemotron-3-ultra was weakest at 27.04 token/s, roughly one-fifth of the leader's average.
  • The most significant volatility is nemotron-3-ultra: an 83.0% coefficient of variation, a -68.3% trend, and swings from 73.62 token/s down to 3.48 token/s; glm-5.3 moved oppositely, rising 31.7% to close at 134.24 token/s.
  • No missing-data limitation exists in this window: all eight models delivered 12 of 12 expected samples, giving 96 valid points and 100.0% coverage, though the four-hour span limits longer-term conclusions.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model at 146.24 token/s average throughput (peak 210.67 token/s), while nemotron-3-ultra is the weakest at 26.85 token/s average, peaking at only 73.62 token/s.
  • Six of eight models declined over the window; glm-5.2 fell 22.5% to a 33.34 token/s close, and deepseek-v4-pro dipped to 29.39 token/s at 10:40. nemotron-3-ultra was the most volatile, ranging 3.89 to 73.62 token/s with an 84.8% coefficient of variation.
  • No missing-data limitation applies: all 96 expected points are valid, each model has 12 of 12 samples, and coverage is 100.0%.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model at 139.41 token/s average throughput (peak 196.85 token/s), while nemotron-3-ultra is the weakest at 31.51 token/s average, peaking at only 73.62 token/s.
  • nemotron-3-ultra shows the most operationally significant volatility, with a 74.9% coefficient of variation and swings between 3.89 and 73.62 token/s; glm-5.3 is next most volatile at 38.0% CV, ranging 55.08 to 187.24 token/s.
  • No missing-data limitation exists: all 96 expected points are valid (100.0% coverage), and every model has 12 of 12 samples across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput was gemma4:31b at 130.93 token/s; weakest was nemotron-3-ultra at 26.87 token/s, which also recorded the window's lowest single reading of 2.23 token/s at 06:00 UTC.
  • The most operationally significant volatility came from glm-5.3-flash, whose coefficient of variation was 43.5%, swinging between 172.45 and 20.84 token/s; deepseek-v4-flash was the only model with a clear upward trend at +11.2%, peaking at 196.85 token/s at 07:20 UTC.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, the dataset contains 96 valid points, and coverage is 100.0% for the full four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b leads average output throughput at 130.04 token/s, narrowly ahead of deepseek-v4-flash at 126.59 token/s, while nemotron-3-ultra is the weakest model at 23.80 token/s.
  • glm-5.3-flash shows the steepest deterioration, trending down 50.2% from a 185.04 token/s peak at 04:40 to 54.27 token/s at 08:00; nemotron-3-ultra is the most volatile, with a 78.4% coefficient of variation and swings between 2.23 and 60.25 token/s.
  • No missing-data limitation applies: all eight models recorded 12 of 12 expected samples, yielding 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash is the strongest model by average throughput at 131.33 token/s, narrowly ahead of gemma4:31b at 130.9 token/s; nemotron-3-ultra is the weakest at 31.36 token/s average, with a maximum of only 73.68 token/s.
  • The most operationally significant volatility is nemotron-3-ultra, whose coefficient of variation is 71.4%, including drops to 2.23 token/s at 06:00 and 3.01 token/s at 03:40; glm-5.3-flash is also unstable, ranging from 20.84 to 185.04 token/s with 53.4% CV.
  • No missing-data limitation exists: all 96 expected points are valid, coverage is 100.0%, and every model has 12 of 12 samples, so the four-hour window is complete.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput is gemma4:31b at 141.18 token/s (peak 190.48 token/s); weakest is nemotron-3-ultra at 35.60 token/s, never exceeding 73.68 token/s.
  • glm-5.3 shows the sharpest upward trend at +39.5% but with high volatility (stddev 47.99 token/s, cv 40.8%); nemotron-3-ultra is the most volatile (cv 64.7%), collapsing to 2.23 token/s at 06:00, and deepseek-v4-flash fell to 34.47 token/s at the same observation.
  • No missing-data limitation exists: all 8 models report 12 of 12 samples, 96 valid points, and 100.0% coverage over the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b delivered the strongest average throughput at 146.77 token/s, edging deepseek-v4-flash at 139.32 token/s, while nemotron-3-ultra was weakest at 36.32 token/s, far below glm-5.2's 74.4 token/s.
  • glm-5.3-flash showed the most operationally significant volatility, swinging between 22.93 and 196.26 token/s with a 44.0% coefficient of variation and a -23.5% trend; nemotron-3-ultra was similarly erratic at 75.4% CV, dipping to 1.65 token/s.
  • No missing-data limitation applies: all eight models recorded 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window, so the averages rest on complete data.
4-hour window · 96 points · 100.0% coverage
  • Strongest average throughput was gemma4:31b at 151.78 token/s (p95 186.59 token/s); weakest was nemotron-3-ultra at 37.0 token/s, with a minimum of 1.65 token/s.
  • glm-5.3-flash showed the most operationally significant volatility, ranging from 196.26 token/s at 01:20 down to 22.93 token/s at 04:00, with a coefficient of variation of 46.1% and a -22.5% trend; glm-5.3 similarly swung between 172.31 and 62.75 token/s.
  • No missing-data limitation applies: all eight models recorded 12 of 12 expected samples, giving 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • gemma4:31b delivered the strongest average throughput at 133.75 token/s (peaking at 183.41 token/s), while nemotron-3-ultra was weakest at 42.27 token/s, never exceeding 95.76 token/s across the window.
  • nemotron-3-ultra showed the most operationally significant volatility, with a 73.8% coefficient of variation and swings from 1.65 token/s at 01:40 to 95.76 token/s at 23:40; glm-5.3 also oscillated sharply between 62.75 and 172.31 token/s.
  • No missing-data limitation applies: all eight models recorded 12 of 12 expected samples, with 96 valid points and 100.0% coverage over the four-hour period, so the summary reflects complete observations.
4-hour window · 96 points · 100.0% coverage
  • Highest average throughput was deepseek-v4-flash at 139.49 token/s (peak 195.83 token/s); lowest was nemotron-3-ultra at 37.11 token/s, well below the next-lowest model, glm-5.2 at 81.56 token/s.
  • nemotron-3-ultra showed the most operationally significant volatility: coefficient of variation 86.4%, swings from 1.65 token/s at 01:40 to 95.76 token/s at 23:40, a -37.9% trend, and repeated drops below 20 token/s.
  • No missing-data limitation applies: all eight models recorded 12 of 12 samples, valid_point_count is 96, and coverage is 100.0% across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model at 141.90 token/s average throughput, while nemotron-3-ultra is the weakest at 50.87 token/s average, roughly a third of the leader's rate.
  • The most operationally significant volatility is nemotron-3-ultra, which swings from 3.85 to 95.76 token/s with a 63.5% coefficient of variation; glm-5.3 also shows a sharp -31.3% trend, falling from a 186.30 token/s peak to 65.00 token/s at 23:40.
  • No missing-data limitation exists: all eight models have 12 of 12 samples, 96 valid points, and 100.0% coverage, though the four-hour window limits longer-term conclusions.
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4-flash led average throughput at 140.33 token/s, narrowly ahead of glm-5.3 at 140.21 token/s; nemotron-3-ultra was weakest at 48.46 token/s, never exceeding 95.76 token/s.
  • glm-5.3 held 168–186 token/s from 21:40 to 23:00, then dropped to 75.45 and 65.0 token/s at 23:20 and 23:40 before recovering to 136.34 token/s; nemotron-3-ultra was the most volatile, with a 69.6% coefficient of variation and swings between 3.85 and 95.76 token/s.
  • No missing-data limitation applies: all eight models recorded 12 of 12 samples, giving 96 valid points and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model with average throughput of 153.72 token/s (peak 186.30 token/s); nemotron-3-ultra is the weakest at 38.10 token/s average, dipping to 3.85 token/s at 23:00 UTC.
  • The most operationally significant volatility comes from nemotron-3-ultra, swinging between 3.85 and 94.35 token/s (CV 80.2%), and glm-5.3-flash, ranging 62.32 to 193.35 token/s (CV 43.5%); meanwhile glm-5.3 trended up 20.8%, from a 130.10 token/s floor to a 186.30 token/s peak.
  • No missing-data limitation applies: all eight models report 12 of 12 samples, 96 valid points, and 100.0% coverage across the four-hour window.
4-hour window · 96 points · 100.0% coverage
  • glm-5.3 is the strongest model with an average throughput of 146.23 token/s, peaking at 186.3 token/s at 22:00 UTC, while nemotron-3-ultra is the weakest at 35.57 token/s average, never exceeding 94.35 token/s.
  • The most operationally significant volatility comes from nemotron-3-ultra, whose coefficient of variation is 86.1%, swinging between 6.13 and 94.35 token/s and trending up 155.9% over the window; gemma4:31b also fell from a 173.46 token/s peak to 55.4 token/s at 22:00, a 23.4% decline.
  • No missing-data limitation applies: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0%, so the four-hour window is fully populated.