← Performance dashboard

Hourly performance insights

Summaries of rolling four-hour performance data

1203 retained summaries
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model by average throughput at 159.81 token/s (p95 of 182.91 token/s), while nemotron-3-ultra is the weakest at 26.08 token/s average, far below the next-lowest model, minimax-m3, at 68.89 token/s.
  • The most operationally significant volatility is nemotron-3-ultra's step change around 22:25 UTC: it ran between 2.07 and 6.78 token/s for roughly the first two and a half hours, then jumped to peaks of 101.92 token/s, producing a 118.3% coefficient of variation.
  • No missing-data limitation applies: all eight models recorded 48 of 48 expected samples, totaling 384 valid points at 100.0% coverage across the four-hour window.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model at 149.26 token/s average throughput, ahead of glm-5.2 at 135.11 token/s; nemotron-3-ultra is the weakest at 13.92 token/s average, with a minimum of 1.45 token/s.
  • The most operationally significant volatility is nemotron-3-ultra, which ran near 2 to 6 token/s for most of the window, then spiked to 101.92 token/s at 22:35, driving its coefficient of variation to 185.1 percent.
  • No missing-data limitation applies: all eight models report 48 of 48 samples at 100.0 percent coverage, and the dataset's 384 valid points match the expected total.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model with average throughput of 149.09 token/s (p95 183.28 token/s), while nemotron-3-ultra is the weakest at 11.71 token/s average, peaking only once at 92.73 token/s.
  • The most operationally significant event is nemotron-3-ultra's collapse: after 18:55 it drops from 92.73 token/s to 1.45 token/s by 19:40 and stays below 7 token/s through 22:00, a -78.9% trend with 160.7% coefficient of variation.
  • No missing-data limitation applies: all eight models report 48 of 48 expected samples, 384 valid points total, and 100.0% coverage, so the four-hour window is fully represented.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model by average throughput at 151.64 token/s (peak 186.34 token/s), while nemotron-3-ultra is the weakest at 16.46 token/s average (minimum 1.39 token/s).
  • The most operationally significant volatility is nemotron-3-ultra's collapse: after 19:00 it stays between roughly 1.45 and 5.84 token/s through 21:00, versus an earlier peak of 92.73 token/s, producing a 135.2% coefficient of variation and a -89.4% trend.
  • No missing-data limitation applies: all eight models report 48 of 48 expected samples with 100.0% coverage, and the dataset contains 384 valid points across the four-hour window.
4-hour window · 384 points · 100.0% coverage
  • glm-5.3 is the strongest model with average throughput of 137.91 token/s (p95 183.4 token/s, max 186.34 token/s), while nemotron-3-ultra is the weakest at 19.77 token/s average, peaking at only 92.73 token/s.
  • The most operationally significant volatility is nemotron-3-ultra: coefficient of variation 119.9%, with a sustained collapse from 19:05 to 20:00 UTC where readings stay between 1.45 and 5.84 token/s, versus its 92.73 token/s peak at 18:55.
  • No missing-data limitation exists: all eight models report 48 of 48 samples with 100.0% coverage, and 384 valid points match the expected total, so the only constraint is the four-hour window itself.
4-hour window · 384 points · 100.0% coverage
  • glm-5.2 is the strongest model with an average throughput of 124.14 token/s, while nemotron-3-ultra is the weakest at 31.39 token/s, roughly a quarter of the leader's rate.
  • The most significant movement is glm-5.3's recovery: it started near 12.08 token/s at 15:05 UTC and climbed to a 186.34 token/s peak at 17:55, ending at 161.23 token/s. nemotron-3-ultra shows extreme volatility, with a coefficient of variation of 91.0% and swings between 1.35 and 109.36 token/s.
  • No missing-data limitation exists: all eight models report 48 of 48 expected samples, 100.0% coverage, and 384 valid points across the four-hour window.
4-hour window · 384 points · 100.0% coverage
  • glm-5.2 is the strongest model with an average throughput of 133.27 token/s (p95 176.08 token/s), while nemotron-3-ultra is the weakest at 30.40 token/s average, peaking at only 109.36 token/s.
  • The most significant shift is glm-5.3, which climbed from single-digit rates near 6.57 token/s early in the window to a maximum of 186.34 token/s, a 255.5% trend with 74.1% coefficient of variation; nemotron-3-ultra was similarly unstable at 97.8% CV.
  • No missing-data limitation exists: all eight models report 48 of 48 expected samples, 384 valid points total, and 100.0% coverage across the four-hour window.
4-hour window · 384 points · 100.0% coverage
  • glm-5.2 delivered the highest average throughput at 150.74 token/s (p95 204.97 token/s), while nemotron-3-ultra was the weakest at 29.60 token/s average with a minimum of 0.92 token/s.
  • glm-5.3 shows the most operationally significant volatility: throughput collapsed from 81.27 token/s at 14:00 to between 6.57 and 15.52 token/s from roughly 14:10 to 15:10, with a coefficient of variation of 69.3%, before recovering to 150.82 token/s by 17:00.
  • No missing-data limitation applies: all eight models recorded 48 of 48 expected samples, giving 384 valid points and 100.0% coverage across the four-hour window.
4-hour window · 375 points · 97.7% coverage
  • glm-5.2 is the strongest model at 160.44 token/s average throughput (peak 220.37 token/s), while nemotron-3-ultra is the weakest at 30.36 token/s average, ranging from 0.92 to 109.36 token/s across the window.
  • glm-5.3 shows the most operationally significant volatility: it fell from 130.82 token/s at 12:50 to between 6.57 and 15.52 token/s from 14:10 to 15:15, producing a 74.5% coefficient of variation and a -53.3% trend, before partially recovering to 126.23 token/s at 15:45.
  • glm-5.3 is missing 9 of 48 expected samples (81.2% coverage), with no observations from 12:05 to 12:45 UTC, so its early-window behavior is unknown; overall dataset coverage is 97.7% with 375 valid points.
4-hour window · 363 points · 94.5% coverage
  • glm-5.2 is the strongest model at 166.2 token/s average throughput (p95 215.87 token/s); nemotron-3-ultra is the weakest at 24.42 token/s average, peaking at only 93.45 token/s and dropping to 0.92 token/s.
  • glm-5.3 shows the most significant degradation: near 130 token/s at 12:50 UTC, it fell to 11.04 token/s at 14:10 and stayed between 6.57 and 14.67 token/s through 15:00, a 77.4% decline; nemotron-3-ultra is also highly volatile with 101.0% cv.
  • glm-5.3 has only 27 of 48 expected samples (56.2% coverage) with no observations before 12:50 UTC, so its 64.68 token/s average is unreliable; overall coverage is 94.5% with 363 valid points.
4-hour window · 343 points · 89.3% coverage
  • glm-5.2 is the strongest model at 162.28 token/s average throughput (p95 215.91 token/s), while nemotron-3-ultra is the weakest at 21.52 token/s average, never exceeding 63.67 token/s.
  • nemotron-3-ultra shows the most operationally significant volatility, with a 97.6% coefficient of variation and swings between 1.15 and 63.67 token/s, including a 12:00–12:15 stretch stuck near 1–2.5 token/s; gemma4:31b also dipped to 24.09 token/s at 10:20.
  • glm-5.3 has only 14 of 48 expected samples (29.2% coverage), all from 12:50 onward, so its 108.31 token/s average is not comparable to other models; overall dataset coverage is 89.3% across 343 valid points.
4-hour window · 339 points · 88.3% coverage
  • Between 09:03 and 13:03 UTC, glm-5.2 was the strongest model at 159.75 token/s average output throughput, peaking at 223.09 token/s, while nemotron-3-ultra was weakest at 17.29 token/s average.
  • The most operationally significant issue is nemotron-3-ultra's extreme volatility: a 113.1% coefficient of variation with swings between 1.15 and 63.01 token/s, including prolonged stretches near 2 token/s. By contrast, deepseek-v4-pro was the steadiest performer (10.9% CV, 109.68 token/s average).
  • Overall coverage was 88.3% of 339 valid points; glm-5.3 reported only 3 of 48 expected samples (6.2% coverage), so its 96.95 token/s average is unreliable.
4-hour window · 338 points · 88.0% coverage
  • Over the 08:56–12:56 UTC window, glm-5.2 was the strongest model, averaging 160.48 token/s (p95 210.27 token/s), while nemotron-3-ultra was the weakest at 16.45 token/s average.
  • The most operationally significant issue is nemotron-3-ultra's extreme volatility: a 113.2% coefficient of variation, with five-minute readings swinging between 1.15 and 63.01 token/s, including prolonged stretches near 2 token/s. gemma4:31b also showed sharp swings (24.09–181.92 token/s) but trended upward 17.1%.
  • Overall dataset coverage was 88.0% (338 valid points); glm-5.3 reported only 2 of 48 expected samples (4.2% coverage), so its 128.37 token/s average is unreliable and limits window-wide conclusions.
4-hour window · 336 points · 100.0% coverage
  • Across the four-hour window, glm-5.2 delivered the strongest average throughput at 156.92 token/s, while nemotron-3-ultra was the weakest at 18.89 token/s.
  • The most operationally significant volatility occurred in glm-5.2, which frequently swung from over 200 token/s down to brief severe drops near 16.66 token/s. In contrast, deepseek-v4-pro maintained the most stable performance, varying only between 52.41 and 131.48 token/s.
  • Although dataset coverage is 100.0 percent across all 336 valid samples, the five-minute observation interval limits the ability to capture sub-minute latency spikes or exact drop durations.
4-hour window · 336 points · 100.0% coverage

Across the four-hour window, glm-5.2 delivered the strongest average throughput at 157.94 token/s, while nemotron-3-ultra was weakest at 17.67 token/s. Operationally, nemotron-3-ultra showed extreme volatility with a 109.6% coefficient of variation, dropping to a minimum of 1.48 token/s. Additionally, glm-5.2 and gemma4:31b exhibited sharp final-interval declines, falling to 16.66 token/s and 141.34 token/s, respectively. The dataset contains 336 valid points across seven models, achieving 100.0% coverage with no missing data limitations.

4-hour window · 336 points · 100.0% coverage

Across the four-hour window, glm-5.2 delivered the highest average throughput at 164.98 token/s, while nemotron-3-ultra was the weakest at 25.72 token/s. The most operationally significant volatility occurred in nemotron-3-ultra, which experienced a severe throughput collapse after 07:35 UTC, dropping from 40.20 token/s to 4.08 token/s and eventually hitting a minimum of 1.48 token/s. This model also exhibited extreme variability with a coefficient of variation of 95.7 percent. The dataset contains complete coverage with 336 valid points across all models, so there are no missing-data limitations affecting this analysis.

4-hour window · 336 points · 100.0% coverage

Across the four-hour window, glm-5.2 delivered the strongest average throughput at 169.86 token/s, while nemotron-3-ultra was weakest at 35.49 token/s. Operationally, nemotron-3-ultra exhibited severe volatility with a 79.0% coefficient of variation and a 50.7% downward trend, dropping from a 105.57 token/s peak to just 1.55 token/s by 08:15. In contrast, minimax-m3 remained the most stable at 47.95 token/s average with a 15.3% coefficient of variation. The dataset is complete with 100.0% coverage and no missing-data limitations across all 336 valid observations.

4-hour window · 336 points · 100.0% coverage

Across the four-hour window, glm-5.2 delivered the strongest average throughput at 166.89 token/s, while nemotron-3-ultra was weakest at 39.12 token/s. Operationally, nemotron-3-ultra showed extreme volatility with a 71.9% coefficient of variation, dropping from a peak of 105.57 token/s to a low of 2.72 token/s. In contrast, minimax-m3 was the most stable model with a 17.1% coefficient of variation and an average of 49.09 token/s. The dataset contains 336 valid observations across seven models with 100.0% coverage, so there are no missing-data limitations affecting this analysis.

4-hour window · 336 points · 100.0% coverage

Across the four-hour window, glm-5.2 delivered the strongest average throughput at 169.82 token/s, while nemotron-3-ultra was weakest at 47.52 token/s. Operationally, nemotron-3-ultra exhibited extreme volatility, swinging from a low of 4.31 to a high of 116.48 token/s, yielding a coefficient of variation of 59.8 percent. This instability contrasts sharply with the steady performance of glm-5.3-flash, which maintained a much tighter range between 50.51 and 102.02 token/s. The dataset is complete with 100.0 percent coverage across all 336 valid points, meaning there are no missing-data limitations to account for in this analysis.

4-hour window · 336 points · 100.0% coverage

Across the four-hour window, glm-5.2 delivered the strongest average throughput at 172.07 token/s, while nemotron-3-ultra was weakest at 48.92 token/s. The most operationally significant volatility is nemotron-3-ultra's extreme instability, swinging from a high of 120.56 token/s down to 4.31 token/s, yielding a 64.9 percent coefficient of variation and a 23.1 percent downward trend. In contrast, glm-5.3-flash remained the most stable, averaging 87.03 token/s with a 9.7 percent coefficient of variation. Although dataset coverage is 100.0 percent with 336 valid samples, the four-hour duration limits broader operational conclusions.

4-hour window · 336 points · 100.0% coverage

Across the four-hour window, glm-5.2 delivered the highest average throughput at 163.79 token/s, while nemotron-3-ultra was the weakest at 48.87 token/s. Operationally, nemotron-3-ultra exhibited extreme volatility with a 59.7% coefficient of variation, swinging from a maximum of 120.56 token/s down to 4.31 token/s. Additionally, deepseek-v4-pro showed a notable downward trend, dropping 11.5% over the period to a low of 32.43 token/s. The dataset contains 336 valid observations across seven models with 100.0% coverage, so there are no missing-data limitations affecting this analysis.

4-hour window · 335 points · 99.7% coverage

Across the four-hour window, glm-5.2 delivered the strongest average throughput at 156.45 token/s, while nemotron-3-ultra was weakest at 40.18 token/s. The most operationally significant volatility belongs to nemotron-3-ultra, which dropped to 0.9 token/s and showed a 82.2% coefficient of variation, indicating highly unstable performance. In contrast, glm-5.3-flash was the most stable model, averaging 85.48 token/s with a low 9.8% coefficient of variation. A missing-data limitation affects this analysis: nemotron-3-ultra has 47 valid samples instead of the expected 48, resulting in 97.9% coverage. All other models maintained complete 100% coverage across the 335 total valid observations.

4-hour window · 335 points · 99.7% coverage

Across the four-hour window, glm-5.2 delivered the strongest average throughput at 156.43 token/s, while nemotron-3-ultra was weakest at 40.25 token/s. Operationally, nemotron-3-ultra exhibited extreme volatility, dropping to 0.9 token/s around 00:50 UTC before recovering to 145.81 token/s earlier at 23:55 UTC. This volatility reflects a coefficient of variation of 86.7 percent. A missing-data limitation affects nemotron-3-ultra, which has 47 valid samples instead of the expected 48, resulting in 97.9 percent coverage. In contrast, glm-5.3-flash maintained the most stable performance with a low 11.0 percent coefficient of variation and an average of 84.07 token/s.

4-hour window · 335 points · 99.7% coverage

Across the four-hour window, glm-5.2 delivered the highest average throughput at 156.28 token/s, while nemotron-3-ultra was the weakest at 34.85 token/s. The most operationally significant volatility occurred in nemotron-3-ultra, which exhibited extreme instability with a coefficient of variation of 91.2% and a sharp downward trend of -45.9%, plummeting from a peak of 145.81 token/s to near-zero levels around 00:50. This dataset contains a missing-data limitation: nemotron-3-ultra has only 47 valid samples out of an expected 48, missing one five-minute observation entirely. All other models maintained complete 100.0% coverage.

4-hour window · 335 points · 99.7% coverage

Across the four-hour window, glm-5.2 delivered the strongest average throughput at 157.82 token/s, while nemotron-3-ultra was weakest at 35.96 token/s. Operationally, nemotron-3-ultra showed extreme volatility with a 94.2% coefficient of variation and a severe downward trend, dropping from 145.81 token/s to near-zero levels around 00:10. In contrast, glm-5.3-flash maintained the most stable performance with a 12.5% coefficient of variation. The dataset has a minor missing-data limitation: nemotron-3-ultra recorded 47 samples instead of the expected 48, missing one observation, though overall dataset coverage remained high at 99.7%.

4-hour window · 336 points · 100.0% coverage

Over the four-hour window, glm-5.2 delivered the highest average throughput at 165.6 token/s, while minimax-m3 was the weakest at 47.71 token/s. The most operationally significant volatility occurred in nemotron-3-ultra, which exhibited extreme swings from a low of 4.89 token/s to a high of 145.81 token/s, resulting in a coefficient of variation of 70.0 percent. Several other models, including glm-5.2 and deepseek-v4-flash, also experienced periodic sharp throughput drops below 25 token/s. The dataset contains complete coverage with 336 valid points across all seven models, so there are no missing-data limitations affecting this analysis.

4-hour window · 336 points · 100.0% coverage

Across the four-hour window, glm-5.2 delivered the strongest average throughput at 159.66 token/s, while nemotron-3-ultra was weakest at 38.3 token/s. Operationally, glm-5.2 showed high volatility with a 30.7% coefficient of variation, including severe drops to 25.73 token/s. Nemotron-3-ultra exhibited even greater instability with a 68.0% coefficient of variation, ranging from 4.89 to 135.27 token/s. The dataset contains 336 valid points across seven models with 100.0% coverage, so there are no missing-data limitations affecting this analysis.

4-hour window · 336 points · 100.0% coverage

Across the four-hour window, glm-5.2 delivered the strongest average throughput at 143.96 token/s, while nemotron-3-ultra was weakest at 31.39 token/s. The most operationally significant volatility occurred in glm-5.2, which fluctuated heavily between 10.73 and 217.13 token/s despite its high average. Nemotron-3-ultra also showed extreme instability, starting near 1.8 token/s before spiking to 135.27 token/s at the end. The dataset contains 336 valid points across seven models with 100.0 percent coverage, meaning there are no missing-data limitations affecting this analysis.

4-hour window · 336 points · 100.0% coverage

Across the four-hour window, glm-5.2 delivered the strongest average throughput at 144.6 token/s, while nemotron-3-ultra was weakest at 20.66 token/s. The most operationally significant volatility occurred in glm-5.2, which fluctuated from a low of 10.73 token/s to a high of 217.13 token/s, and in nemotron-3-ultra, which showed extreme instability with a coefficient of variation of 104.2 percent. Deepseek-v4-flash exhibited a notable throughput drop late in the window, falling to 46.28 token/s at 20:25. The dataset contains 336 valid points across seven models, achieving 100.0 percent coverage with no missing-data limitation.

4-hour window · 333 points · 99.1% coverage

Across the four-hour window, glm-5.2 had the strongest average throughput at 152.29 token/s, while nemotron-3-ultra was weakest at 13.5 token/s. The most operationally significant volatility occurred in glm-5.2, which dropped sharply from 203.35 token/s at 17:50 to 10.73 token/s at 18:10 before recovering. Nemotron-3-ultra also exhibited extreme instability, fluctuating between 0.6 and 75.07 token/s with a coefficient of variation of 134.8 percent. A missing-data limitation affects nemotron-3-ultra, which recorded only 45 valid samples out of 48 expected, resulting in 93.8 percent coverage and leaving three five-minute intervals unobserved.

4-hour window · 332 points · 98.8% coverage

Across the four-hour window, glm-5.2 delivered the highest average throughput at 158.21 token/s, while nemotron-3-ultra was the weakest at 4.7 token/s. The most operationally significant volatility occurred in glm-5.2, which experienced severe latency spikes dropping throughput to 10.73 token/s at 18:10 before recovering. Overall dataset coverage is 98.8%, but this masks a missing-data limitation for nemotron-3-ultra, which has only 44 valid samples out of 48 expected, restricting visibility into its true performance baseline.

4-hour window · 332 points · 98.8% coverage

Over the four-hour window, glm-5.2 delivered the strongest average throughput at 164.14 token/s, while nemotron-3-ultra was the weakest at 2.85 token/s. The most operationally significant volatility occurred in glm-5.2, which dropped sharply from 213.81 token/s at 17:10 to 25.55 token/s at 17:15 before recovering. Deepseek-v4-flash showed a notable upward trend, increasing by 10.0 percent to an average of 101.5 token/s. A key limitation is missing data: nemotron-3-ultra has only 44 valid samples out of 48 expected, resulting in 91.7 percent coverage, which leaves gaps in its throughput profile.

4-hour window · 331 points · 98.5% coverage

Across the four-hour window, glm-5.2 delivered the highest average throughput at 163.88 token/s, while nemotron-3-ultra was the weakest at 1.64 token/s. The most operationally significant volatility occurred in nemotron-3-ultra, which spiked from sub-1 token/s baselines to 6.35 token/s after 16:30 UTC. Gemma4:31b also showed notable instability, swinging between 24.51 and 136.67 token/s. A missing-data limitation affects this analysis: nemotron-3-ultra recorded only 43 of 48 expected samples, dropping five observations between 16:10 and 16:30 UTC, meaning its recent throughput spike lacks continuous context.

4-hour window · 326 points · 97.0% coverage

Across the four-hour window, glm-5.2 delivered the strongest average throughput at 141.69 token/s, while nemotron-3-ultra was weakest at 0.91 token/s. The most operationally significant volatility was a severe throughput drop around 13:20 UTC, where gemma4:31b fell to 24.51 token/s, glm-5.2 dropped to 42.09 token/s, and glm-5.3-flash reached 24.24 token/s. Despite this drop, glm-5.2 showed a strong upward trend of 39.8 percent, ending at 183.02 token/s. The dataset has a missing-data limitation: nemotron-3-ultra recorded only 44 samples instead of the expected 48, resulting in 91.7 percent coverage, which limits visibility into its exact performance dips.

4-hour window · 326 points · 97.0% coverage

Across the four-hour window, glm-5.2 delivered the strongest average throughput at 121.03 token/s, while nemotron-3-ultra was the weakest at 2.83 token/s. The most operationally significant volatility is glm-5.2's sharp 64.4 percent upward trend, which included a peak of 212.94 token/s despite a high coefficient of variation of 37.1 percent. In contrast, deepseek-v4-pro offered the most stable performance, averaging 102.65 token/s with a low coefficient of variation of 12.9 percent. A missing-data limitation affects this analysis: nemotron-3-ultra captured only 44 of 48 expected samples, dropping to 91.7 percent coverage, which limits the reliability of its calculated averages.

4-hour window · 333 points · 99.1% coverage

Across the four-hour window, glm-5.2 delivered the strongest average throughput at 107.11 token/s, while nemotron-3-ultra was weakest at 8.72 token/s. Nemotron-3-ultra also showed the most operationally significant volatility, with a 161.5% coefficient of variation and a 94.5% downward trend, dropping from a 64.81 token/s peak to 1.17 token/s. In contrast, deepseek-v4-pro offered the most stable performance, averaging 104.33 token/s with a 14.1% coefficient of variation. A missing-data limitation affects nemotron-3-ultra, which recorded only 45 valid samples instead of the expected 48, resulting in 93.8% coverage.

4-hour window · 334 points · 99.4% coverage

Across the four-hour window, deepseek-v4-pro delivered the strongest average throughput at 106.16 token/s, while nemotron-3-ultra was weakest at 11.18 token/s. The most operationally significant volatility is nemotron-3-ultra's severe degradation, trending down 69.9 percent to a latest throughput of 1.05 token/s. Conversely, deepseek-v4-flash showed a positive 15.5 percent trend, peaking at 154.18 token/s. A missing-data limitation affects this analysis: nemotron-3-ultra recorded only 46 valid samples instead of 48, dropping its coverage to 95.8 percent, whereas the other six models maintained complete 48-sample coverage.

4-hour window · 335 points · 99.7% coverage

Across the four-hour window, deepseek-v4-pro had the strongest average output-token throughput at 105.69 token/s, while nemotron-3-ultra was the weakest at 13.14 token/s. The most operationally significant volatility came from nemotron-3-ultra, which exhibited a 95.0% coefficient of variation and dropped to a minimum of 0.77 token/s. In contrast, deepseek-v4-pro maintained the most stable performance with a 14.0% coefficient of variation. Overall dataset coverage was 99.7% across 335 valid points, but a missing-data limitation exists for nemotron-3-ultra, which recorded only 47 samples instead of the expected 48, leaving a gap in its observations.

4-hour window · 336 points · 100.0% coverage

Across the four-hour observation period, deepseek-v4-pro achieved the strongest average throughput at 106.99 token/s, while nemotron-3-ultra was the weakest at 13.72 token/s. The most operationally significant volatility occurred in nemotron-3-ultra, which exhibited a coefficient of variation of 78.4 percent and a sharp throughput spike to 64.81 token/s at 10:55 UTC after hovering near 10 token/s for most of the window. In contrast, deepseek-v4-pro maintained the most stable performance with a coefficient of variation of just 15.4 percent. The dataset contains complete coverage with 336 valid points across all seven models, so there are no missing-data limitations affecting this specific interval.

4-hour window · 336 points · 100.0% coverage

Across the four-hour observation window, deepseek-v4-pro achieved the highest average throughput at 105.82 token/s, while nemotron-3-ultra was the weakest at 11.38 token/s. The most operationally significant volatility occurred within nemotron-3-ultra, which exhibited extreme instability with a coefficient of variation of 69.5 percent and throughput frequently dropping below 10 token/s, despite starting at 45.46 token/s. In contrast, deepseek-v4-pro maintained the most stable performance with a coefficient of variation of just 16.1 percent. The dataset contains 336 valid observations across seven models with 100.0 percent coverage, meaning there are no missing-data limitations affecting this analysis.

4-hour window · 336 points · 100.0% coverage

Across the four-hour window, deepseek-v4-pro had the strongest average throughput at 106.06 token/s, while nemotron-3-ultra was weakest at 16.65 token/s. The most operationally significant volatility appeared in nemotron-3-ultra, which dropped from a 78.58 token/s peak to a 2.46 token/s minimum with a 102.1 percent coefficient of variation, alongside a severe downward trend of 54.8 percent. In contrast, deepseek-v4-pro remained stable with a 16.4 percent coefficient of variation. glm-5.2 also showed notable instability, falling to 21.35 token/s at 06:10. The dataset contains 336 valid observations with 100.0 percent coverage, so no missing-data limitation affects this specific analysis.

4-hour window · 336 points · 100.0% coverage

Across the four-hour window, gemma4:31b delivered the strongest average throughput at 115.32 token/s, while nemotron-3-ultra was the weakest at 22.98 token/s. The most operationally significant volatility came from nemotron-3-ultra, which dropped 62.4 percent over the period to a latest throughput of 8.19 token/s, alongside a coefficient of variation of 87.6 percent. deepseek-v4-flash showed a contrasting upward trend of 13.4 percent but remained highly volatile, swinging between 13.7 and 146.06 token/s. The dataset includes 336 valid observations across seven models with 100.0 percent coverage, so there are no missing-data limitations affecting this analysis.

4-hour window · 336 points · 100.0% coverage

Across the four-hour window, gemma4:31b had the strongest average throughput at 115.75 token/s, while nemotron-3-ultra was weakest at 31.28 token/s. The most operationally significant volatility appeared in deepseek-v4-flash, which dropped from 149.62 token/s to a minimum of 13.7 token/s before recovering, yielding a coefficient of variation of 52.2 percent. Nemotron-3-ultra also showed severe instability, trending down by 42.1 percent and ending at 7.15 token/s. In contrast, deepseek-v4-pro remained the most stable, averaging 105.85 token/s with a coefficient of variation of just 15.8 percent. The dataset includes 336 valid observations across seven models, achieving 100.0 percent coverage with no missing-data limitations.

4-hour window · 336 points · 100.0% coverage

Across the four-hour observation window, gemma4:31b achieved the highest average throughput at 124.69 token/s, while nemotron-3-ultra was the weakest at 34.62 token/s. The most operationally significant volatility occurred in deepseek-v4-flash, which experienced a severe throughput collapse from 149.62 token/s down to 13.7 token/s between 03:20 and 04:05, reflecting a 51.6 coefficient of variation. Nemotron-3-ultra also showed extreme instability, dropping to a minimum of 3.69 token/s. The dataset contains 336 valid observations across seven models with 100.0 percent coverage, so no missing-data limitations affect this specific analysis window.

4-hour window · 336 points · 100.0% coverage

Across the four-hour window, gemma4:31b is the strongest model with an average throughput of 126.65 token/s, while nemotron-3-ultra is the weakest at 32.78 token/s. The most operationally significant volatility appears in deepseek-v4-flash, which dropped sharply from 132.42 token/s at 02:35 to 13.7 token/s at 04:05 before recovering to 140.47 token/s at 04:45. Nemotron-3-ultra also showed extreme instability, spiking to 94.97 token/s at 03:05 after a low of 3.69 token/s at 02:45. The dataset includes 336 valid observations across seven models. Although coverage is complete at 100 percent, the five-minute sampling interval limits the ability to detect sub-minute latency spikes or brief micro-stalls.

4-hour window · 336 points · 100.0% coverage

Across the four-hour observation window, gemma4:31b delivered the strongest average output-token throughput at 118.25 token/s, while nemotron-3-ultra was the weakest at 30.27 token/s. The most operationally significant volatility occurred in deepseek-v4-flash, which maintained an average of 89.54 token/s before suffering a severe throughput collapse after 03:30 UTC, dropping to a minimum of 18.16 token/s. Nemotron-3-ultra also exhibited extreme instability, ranging from 3.59 to 94.97 token/s with a coefficient of variation of 70.3 percent. The dataset contains 336 valid observations across seven models with 100.0 percent coverage, meaning there are no missing-data limitations affecting this analysis.

4-hour window · 336 points · 100.0% coverage

Across the four-hour window, glm-5.2 delivered the highest average throughput at 123.04 token/s, while nemotron-3-ultra was the weakest at 29.26 token/s. The most operationally significant volatility came from glm-5.3-flash, which exhibited extreme swings including a minimum of 17.49 token/s and a maximum of 152.6 token/s, reflecting a 39.3% coefficient of variation. Additionally, nemotron-3-ultra experienced severe periodic drops, plunging to 3.59 token/s. The dataset contains 336 valid observations across seven models, achieving 100.0% coverage. However, a missing-data limitation exists: the dataset lacks any concurrent request volume or concurrency metrics, making it impossible to determine if these throughput fluctuations were caused by varying load or internal system instability.

4-hour window · 336 points · 100.0% coverage

Across the four-hour window, glm-5.2 delivered the strongest average throughput at 129.67 token/s, while nemotron-3-ultra was the weakest at 26.97 token/s. The most operationally significant volatility comes from glm-5.3-flash, which spiked from a 17.49 token/s minimum to a 149.62 token/s maximum, yielding a 42.5% coefficient of variation and a 55.7% upward trend. Nemotron-3-ultra also showed extreme instability, dropping as low as 1.73 token/s with a 90.6% coefficient of variation. The dataset contains 336 valid observations across seven models with 100.0% coverage, so there are no missing-data limitations affecting this analysis.

4-hour window · 336 points · 100.0% coverage

Across the four-hour observation window, glm-5.2 delivered the strongest average output-token throughput at 127.97 token/s, while nemotron-3-ultra was the weakest at 25.45 token/s. The most operationally significant volatility came from glm-5.3-flash, which dropped sharply from 154.48 token/s to 17.49 token/s, reflecting a 44.6% coefficient of variation. Similarly, nemotron-3-ultra exhibited extreme instability, fluctuating between 1.73 token/s and 89.2 token/s. The dataset contains 336 valid observations across seven models, achieving 100.0% coverage with no missing-data limitations. This complete dataset allows for reliable operational assessments of model throughput without gaps.

4-hour window · 336 points · 100.0% coverage

Across the four-hour observation window, gemma4:31b delivered the strongest average output-token throughput at 124.01 token/s, while nemotron-3-ultra was the weakest at 27.11 token/s. The most operationally significant volatility came from glm-5.3-flash, which experienced a 51.7 percent downward trend, dropping from early peaks above 150 token/s to a low of 17.49 token/s. Nemotron-3-ultra also showed extreme instability, oscillating between 0.99 and 89.2 token/s. The dataset contains 336 valid observations across seven models with 100.0 percent coverage, so there are no missing-data limitations affecting this specific period.