Ollama Cloud Performance

Output-token throughput benchmarks every 20 minutes

8/8 models online
gemma4:31b
145.7
output tokens / second
123 tokens · 1.2s request
minimax-m3
62.5
output tokens / second
160 tokens · 2.7s request
glm-5.2
74.9
output tokens / second
160 tokens · 2.3s request
glm-5.3
60.8
output tokens / second
160 tokens · 2.8s request
glm-5.3-flash
137.7
output tokens / second
160 tokens · 1.3s request
deepseek-v4-pro
107.4
output tokens / second
160 tokens · 1.6s request
deepseek-v4.1-flash
74.7
output tokens / second
160 tokens · 2.3s request
nemotron-3-ultra
16.3
output tokens / second
160 tokens · 10.0s request

Hourly performance insights

Analysis of the rolling four-hour, twenty-minute dataset

AI-generated · hourly
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model at an average 184.46 token/s (peaking at 235.93 token/s), while nemotron-3-ultra is the weakest at an average 25.54 token/s, never exceeding 52.47 token/s.
  • glm-5.2 shows the most operationally significant volatility, with a 64.2% coefficient of variation and swings from 10.78 to 96.62 token/s within the four hours; nemotron-3-ultra is similarly erratic at 57.9%.
  • No missing-data limitation exists: all eight models report 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the period.
View full insight feed →

Last 24 hours

Token throughput for Last 24 hours