Ollama Cloud Performance

Output-token throughput benchmarks every 5 minutes

6/6 models online
gemma4:31b
62.7
output tokens / second
123 tokens · 2.1s request
minimax-m3
58.5
output tokens / second
160 tokens · 2.9s request
glm-5.2
83.8
output tokens / second
160 tokens · 2.1s request
deepseek-v4-pro
95.7
output tokens / second
159 tokens · 1.8s request
deepseek-v4-flash
133.6
output tokens / second
160 tokens · 1.4s request
nemotron-3-ultra
39.3
output tokens / second
160 tokens · 4.3s request

Hourly performance insights

GLM-5.2 analysis of the rolling four-hour, five-minute dataset

AI-generated · hourly
4-hour window · 288 points · 100.0% coverage

Across the four-hour window, glm-5.2 had the highest average output-token throughput at 122.38 token/s, while nemotron-3-ultra was weakest at 46.77 token/s. The most operationally significant volatility occurred in glm-5.2, which dropped sharply from a peak of 223.93 token/s at 02:25 to a low of 13.59 token/s at 04:30, reflecting a 60.7 percent downward trend. In contrast, deepseek-v4-flash maintained the most stable performance, averaging 106.89 token/s with a coefficient of variation of 20.7 percent. The dataset contains 288 valid observations across six models, achieving 100.0 percent coverage with no missing-data limitations.

View full insight feed →

Last 24 hours

Token throughput for Last 24 hours