{"summaries":[{"coverage_pct":100.0,"eval_count":155,"generated_at":"2026-08-16T08:03:06.161429+00:00","generated_label":"Aug 16, 2026 \u00b7 08:03 UTC","id":40,"model":"glm-5.2","period_end":"2026-08-16T08:03:01+00:00","period_label":"Aug 16, 2026 \u00b7 04:03 UTC to Aug 16, 2026 \u00b7 08:03 UTC","period_start":"2026-08-16T04:03:01+00:00","prompt_eval_count":4514,"summary":"Across the four-hour window, gemma4:31b delivered the highest average throughput at 106.92 token/s, while nemotron-3-ultra was the weakest at 41.95 token/s. The most operationally significant volatility occurred in deepseek-v4-flash, which dropped sharply from 140.19 token/s at 07:50 to 15.7 token/s at 07:55 before recovering to 123.5 token/s at 08:00. glm-5.2 showed the strongest upward trend, increasing by 33.3 percent to a latest throughput of 116.56 token/s. The dataset includes 288 valid points across six models, achieving 100.0 percent coverage with no missing-data limitations.","summary_en":"Across the four-hour window, gemma4:31b delivered the highest average throughput at 106.92 token/s, while nemotron-3-ultra was the weakest at 41.95 token/s. The most operationally significant volatility occurred in deepseek-v4-flash, which dropped sharply from 140.19 token/s at 07:50 to 15.7 token/s at 07:55 before recovering to 123.5 token/s at 08:00. glm-5.2 showed the strongest upward trend, increasing by 33.3 percent to a latest throughput of 116.56 token/s. The dataset includes 288 valid points across six models, achieving 100.0 percent coverage with no missing-data limitations.","summary_zh":"\u5728\u56db\u5c0f\u65f6\u7684\u65f6\u95f4\u7a97\u53e3\u5185\uff0cgemma4:31b \u7684\u5e73\u5747\u541e\u5410\u91cf\u6700\u9ad8\uff0c\u8fbe\u5230 106.92 \u8bcd\u5143/\u79d2\uff0c\u800c nemotron-3-ultra \u8868\u73b0\u6700\u5f31\uff0c\u4e3a 41.95 \u8bcd\u5143/\u79d2\u3002\u6700\u5177\u8fd0\u7ef4\u610f\u4e49\u7684\u6ce2\u52a8\u51fa\u73b0\u5728 deepseek-v4-flash \u4e2d\uff0c\u5176\u541e\u5410\u91cf\u4ece 07:50 \u7684 140.19 \u8bcd\u5143/\u79d2\u6025\u5267\u4e0b\u964d\u81f3 07:55 \u7684 15.7 \u8bcd\u5143/\u79d2\uff0c\u968f\u540e\u5728 08:00 \u6062\u590d\u81f3 123.5 \u8bcd\u5143/\u79d2\u3002glm-5.2 \u5448\u73b0\u51fa\u6700\u5f3a\u7684\u4e0a\u5347\u8d8b\u52bf\uff0c\u589e\u957f\u4e86 33.3%\uff0c\u6700\u65b0\u541e\u5410\u91cf\u8fbe\u5230 116.56 \u8bcd\u5143/\u79d2\u3002\u8be5\u6570\u636e\u96c6\u5305\u542b\u516d\u4e2a\u6a21\u578b\u7684 288 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u5b9e\u73b0\u4e86 100.0% \u7684\u8986\u76d6\u7387\uff0c\u4e0d\u5b58\u5728\u7f3a\u5931\u6570\u636e\u9650\u5236\u3002","total_duration_ns":1674930438,"translated_at":"2026-08-16T08:03:06.161429+00:00","translation_eval_count":178,"translation_model":"glm-5.2","translation_prompt_eval_count":307,"translation_total_duration_ns":2549595209,"translation_wall_seconds":2.744,"valid_point_count":288,"wall_seconds":1.879},{"coverage_pct":100.0,"eval_count":136,"generated_at":"2026-08-16T07:03:05.439833+00:00","generated_label":"Aug 16, 2026 \u00b7 07:03 UTC","id":39,"model":"glm-5.2","period_end":"2026-08-16T07:03:02+00:00","period_label":"Aug 16, 2026 \u00b7 03:03 UTC to Aug 16, 2026 \u00b7 07:03 UTC","period_start":"2026-08-16T03:03:02+00:00","prompt_eval_count":4514,"summary":"Across the four-hour window, deepseek-v4-flash achieved the strongest average throughput at 103.84 token/s, while nemotron-3-ultra was the weakest at 45.14 token/s. The most operationally significant volatility occurred in glm-5.2, which dropped sharply from a peak of 216.98 token/s down to 13.59 token/s at 04:30, reflecting a coefficient of variation of 53.4 percent. This dataset contains 288 valid observations across six models, corresponding to exactly 100.0 percent coverage of the 48 expected samples per model, meaning there are no missing-data limitations to report.","summary_en":"Across the four-hour window, deepseek-v4-flash achieved the strongest average throughput at 103.84 token/s, while nemotron-3-ultra was the weakest at 45.14 token/s. The most operationally significant volatility occurred in glm-5.2, which dropped sharply from a peak of 216.98 token/s down to 13.59 token/s at 04:30, reflecting a coefficient of variation of 53.4 percent. This dataset contains 288 valid observations across six models, corresponding to exactly 100.0 percent coverage of the 48 expected samples per model, meaning there are no missing-data limitations to report.","summary_zh":"\u5728\u56db\u5c0f\u65f6\u7684\u65f6\u95f4\u7a97\u53e3\u5185\uff0cdeepseek-v4-flash \u5b9e\u73b0\u4e86\u6700\u5f3a\u7684\u5e73\u5747\u541e\u5410\u91cf\uff0c\u8fbe\u5230 103.84 \u8bcd\u5143/\u79d2\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u4e3a 45.14 \u8bcd\u5143/\u79d2\u3002\u6700\u5177\u8fd0\u8425\u610f\u4e49\u7684\u6ce2\u52a8\u51fa\u73b0\u5728 glm-5.2 \u4e2d\uff0c\u8be5\u6a21\u578b\u4ece 216.98 \u8bcd\u5143/\u79d2\u7684\u5cf0\u503c\u6025\u5267\u4e0b\u964d\u81f3 04:30 \u65f6\u7684 13.59 \u8bcd\u5143/\u79d2\uff0c\u53cd\u6620\u51fa 53.4% \u7684\u53d8\u5f02\u7cfb\u6570\u3002\u8be5\u6570\u636e\u96c6\u5305\u542b\u8de8\u516d\u4e2a\u6a21\u578b\u7684 288 \u4e2a\u6709\u6548\u89c2\u6d4b\u503c\uff0c\u76f8\u5f53\u4e8e\u6bcf\u4e2a\u6a21\u578b 48 \u4e2a\u9884\u671f\u6837\u672c\u7684\u6070\u597d 100.0% \u8986\u76d6\u7387\uff0c\u8fd9\u610f\u5473\u7740\u6ca1\u6709\u9700\u8981\u62a5\u544a\u7684\u7f3a\u5931\u6570\u636e\u9650\u5236\u3002","total_duration_ns":1551318628,"translated_at":"2026-08-16T07:03:05.439833+00:00","translation_eval_count":150,"translation_model":"glm-5.2","translation_prompt_eval_count":288,"translation_total_duration_ns":1449773742,"translation_wall_seconds":1.591,"valid_point_count":288,"wall_seconds":1.749},{"coverage_pct":100.0,"eval_count":161,"generated_at":"2026-08-16T06:03:06.995174+00:00","generated_label":"Aug 16, 2026 \u00b7 06:03 UTC","id":38,"model":"glm-5.2","period_end":"2026-08-16T06:03:01+00:00","period_label":"Aug 16, 2026 \u00b7 02:03 UTC to Aug 16, 2026 \u00b7 06:03 UTC","period_start":"2026-08-16T02:03:01+00:00","prompt_eval_count":4516,"summary":"Across the four-hour window, glm-5.2 had the highest average output-token throughput at 122.38 token/s, while nemotron-3-ultra was weakest at 46.77 token/s. The most operationally significant volatility occurred in glm-5.2, which dropped sharply from a peak of 223.93 token/s at 02:25 to a low of 13.59 token/s at 04:30, reflecting a 60.7 percent downward trend. In contrast, deepseek-v4-flash maintained the most stable performance, averaging 106.89 token/s with a coefficient of variation of 20.7 percent. The dataset contains 288 valid observations across six models, achieving 100.0 percent coverage with no missing-data limitations.","summary_en":"Across the four-hour window, glm-5.2 had the highest average output-token throughput at 122.38 token/s, while nemotron-3-ultra was weakest at 46.77 token/s. The most operationally significant volatility occurred in glm-5.2, which dropped sharply from a peak of 223.93 token/s at 02:25 to a low of 13.59 token/s at 04:30, reflecting a 60.7 percent downward trend. In contrast, deepseek-v4-flash maintained the most stable performance, averaging 106.89 token/s with a coefficient of variation of 20.7 percent. The dataset contains 288 valid observations across six models, achieving 100.0 percent coverage with no missing-data limitations.","summary_zh":"\u5728\u56db\u5c0f\u65f6\u7684\u65f6\u95f4\u7a97\u53e3\u5185\uff0cglm-5.2 \u7684\u5e73\u5747\u8f93\u51fa\u8bcd\u5143\u541e\u5410\u91cf\u6700\u9ad8\uff0c\u4e3a 122.38 \u8bcd\u5143/\u79d2\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u4e3a 46.77 \u8bcd\u5143/\u79d2\u3002\u6700\u5177\u8fd0\u8425\u610f\u4e49\u7684\u6ce2\u52a8\u51fa\u73b0\u5728 glm-5.2 \u4e2d\uff0c\u5176\u4ece 02:25 \u7684 223.93 \u8bcd\u5143/\u79d2\u5cf0\u503c\u6025\u5267\u4e0b\u964d\u81f3 04:30 \u7684 13.59 \u8bcd\u5143/\u79d2\u4f4e\u70b9\uff0c\u53cd\u6620\u51fa 60.7% \u7684\u4e0b\u964d\u8d8b\u52bf\u3002\u76f8\u6bd4\u4e4b\u4e0b\uff0cdeepseek-v4-flash \u4fdd\u6301\u4e86\u6700\u7a33\u5b9a\u7684\u6027\u80fd\uff0c\u5e73\u5747\u4e3a 106.89 \u8bcd\u5143/\u79d2\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 20.7%\u3002\u8be5\u6570\u636e\u96c6\u5305\u542b\u8de8\u516d\u4e2a\u6a21\u578b\u7684 288 \u4e2a\u6709\u6548\u89c2\u6d4b\u503c\uff0c\u5b9e\u73b0\u4e86 100.0% \u7684\u8986\u76d6\u7387\uff0c\u6ca1\u6709\u7f3a\u5931\u6570\u636e\u7684\u9650\u5236\u3002","total_duration_ns":2543428450,"translated_at":"2026-08-16T06:03:06.995174+00:00","translation_eval_count":171,"translation_model":"glm-5.2","translation_prompt_eval_count":313,"translation_total_duration_ns":2268334085,"translation_wall_seconds":2.426,"valid_point_count":288,"wall_seconds":2.699},{"coverage_pct":100.0,"eval_count":151,"generated_at":"2026-08-16T05:03:05.618583+00:00","generated_label":"Aug 16, 2026 \u00b7 05:03 UTC","id":37,"model":"glm-5.2","period_end":"2026-08-16T05:03:01+00:00","period_label":"Aug 16, 2026 \u00b7 01:03 UTC to Aug 16, 2026 \u00b7 05:03 UTC","period_start":"2026-08-16T01:03:01+00:00","prompt_eval_count":4518,"summary":"Across the four-hour window, glm-5.2 had the strongest average throughput at 153.49 token/s, while nemotron-3-ultra was weakest at 47.92 token/s. The most operationally significant volatility occurred in glm-5.2, which dropped sharply from over 200 token/s early in the window to a minimum of 13.59 token/s at 04:30, reflecting a 43.8 percent downward trend. In contrast, minimax-m3 remained stable with a coefficient of variation of 18.1 percent and an average of 63.38 token/s. The dataset contains 288 valid points across six models, achieving 100.0 percent coverage with no missing-data limitation.","summary_en":"Across the four-hour window, glm-5.2 had the strongest average throughput at 153.49 token/s, while nemotron-3-ultra was weakest at 47.92 token/s. The most operationally significant volatility occurred in glm-5.2, which dropped sharply from over 200 token/s early in the window to a minimum of 13.59 token/s at 04:30, reflecting a 43.8 percent downward trend. In contrast, minimax-m3 remained stable with a coefficient of variation of 18.1 percent and an average of 63.38 token/s. The dataset contains 288 valid points across six models, achieving 100.0 percent coverage with no missing-data limitation.","summary_zh":"\u5728\u56db\u5c0f\u65f6\u7684\u65f6\u95f4\u7a97\u53e3\u5185\uff0cglm-5.2 \u7684\u5e73\u5747\u541e\u5410\u91cf\u6700\u9ad8\uff0c\u8fbe\u5230 153.49 \u8bcd\u5143/\u79d2\uff0c\u800c nemotron-3-ultra \u6700\u4f4e\uff0c\u4e3a 47.92 \u8bcd\u5143/\u79d2\u3002\u6700\u5177\u8fd0\u8425\u610f\u4e49\u7684\u6ce2\u52a8\u51fa\u73b0\u5728 glm-5.2 \u4e2d\uff0c\u5176\u4ece\u7a97\u53e3\u521d\u671f\u7684 200 \u8bcd\u5143/\u79d2\u4ee5\u4e0a\u6025\u5267\u4e0b\u964d\u81f3 04:30 \u7684\u6700\u4f4e\u70b9 13.59 \u8bcd\u5143/\u79d2\uff0c\u53cd\u6620\u51fa 43.8% \u7684\u4e0b\u964d\u8d8b\u52bf\u3002\u76f8\u6bd4\u4e4b\u4e0b\uff0cminimax-m3 \u4fdd\u6301\u7a33\u5b9a\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 18.1%\uff0c\u5e73\u5747\u503c\u4e3a 63.38 \u8bcd\u5143/\u79d2\u3002\u8be5\u6570\u636e\u96c6\u5305\u542b\u8de8\u516d\u4e2a\u6a21\u578b\u7684 288 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u5b9e\u73b0\u4e86 100.0% \u7684\u8986\u76d6\u7387\uff0c\u6ca1\u6709\u7f3a\u5931\u6570\u636e\u7684\u9650\u5236\u3002","total_duration_ns":1977100077,"translated_at":"2026-08-16T05:03:05.618583+00:00","translation_eval_count":159,"translation_model":"glm-5.2","translation_prompt_eval_count":303,"translation_total_duration_ns":1771674741,"translation_wall_seconds":1.926,"valid_point_count":288,"wall_seconds":2.158},{"coverage_pct":100.0,"eval_count":158,"generated_at":"2026-08-16T04:03:06.536012+00:00","generated_label":"Aug 16, 2026 \u00b7 04:03 UTC","id":36,"model":"glm-5.2","period_end":"2026-08-16T04:03:01+00:00","period_label":"Aug 16, 2026 \u00b7 00:03 UTC to Aug 16, 2026 \u00b7 04:03 UTC","period_start":"2026-08-16T00:03:01+00:00","prompt_eval_count":4518,"summary":"Across the 288 valid observations, glm-5.2 is the strongest model with an average throughput of 186.51 token/s, while nemotron-3-ultra is the weakest at 46.66 token/s. The most operationally significant trend is the sharp late-window throughput degradation in glm-5.2, which dropped from 216.98 token/s at 03:15 to a minimum of 58.75 token/s by 04:00, yielding a -10.9% period trend. Deepseek-v4-pro also exhibited high volatility, spiking between 123.03 token/s and 26.57 token/s. The dataset has 100.0% coverage with no missing-data limitation across the expected 48 samples per model.","summary_en":"Across the 288 valid observations, glm-5.2 is the strongest model with an average throughput of 186.51 token/s, while nemotron-3-ultra is the weakest at 46.66 token/s. The most operationally significant trend is the sharp late-window throughput degradation in glm-5.2, which dropped from 216.98 token/s at 03:15 to a minimum of 58.75 token/s by 04:00, yielding a -10.9% period trend. Deepseek-v4-pro also exhibited high volatility, spiking between 123.03 token/s and 26.57 token/s. The dataset has 100.0% coverage with no missing-data limitation across the expected 48 samples per model.","summary_zh":"\u5728288\u4e2a\u6709\u6548\u89c2\u6d4b\u6570\u636e\u4e2d\uff0cglm-5.2\u662f\u6700\u5f3a\u7684\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a186.51\u8bcd\u5143/\u79d2\uff0c\u800cnemotron-3-ultra\u6700\u5f31\uff0c\u4e3a46.66\u8bcd\u5143/\u79d2\u3002\u6700\u5177\u8fd0\u7ef4\u610f\u4e49\u7684\u8d8b\u52bf\u662fglm-5.2\u5728\u65f6\u95f4\u7a97\u53e3\u540e\u671f\u7684\u541e\u5410\u91cf\u6025\u5267\u4e0b\u964d\uff0c\u4ece03:15\u7684216.98\u8bcd\u5143/\u79d2\u964d\u81f304:00\u7684\u6700\u4f4e58.75\u8bcd\u5143/\u79d2\uff0c\u4ea7\u751f-10.9%\u7684\u5468\u671f\u8d8b\u52bf\u3002Deepseek-v4-pro\u4e5f\u8868\u73b0\u51fa\u9ad8\u6ce2\u52a8\u6027\uff0c\u5728123.03\u8bcd\u5143/\u79d2\u548c26.57\u8bcd\u5143/\u79d2\u4e4b\u95f4\u5267\u70c8\u6ce2\u52a8\u3002\u8be5\u6570\u636e\u96c6\u7684\u8986\u76d6\u7387\u4e3a100.0%\uff0c\u5728\u6bcf\u6a21\u578b\u9884\u671f\u768448\u4e2a\u6837\u672c\u4e2d\u4e0d\u5b58\u5728\u7f3a\u5931\u6570\u636e\u9650\u5236\u3002","total_duration_ns":2605990182,"translated_at":"2026-08-16T04:03:06.536012+00:00","translation_eval_count":159,"translation_model":"glm-5.2","translation_prompt_eval_count":310,"translation_total_duration_ns":1776490343,"translation_wall_seconds":1.925,"valid_point_count":288,"wall_seconds":2.78},{"coverage_pct":100.0,"eval_count":155,"generated_at":"2026-08-16T03:03:03.521819+00:00","generated_label":"Aug 16, 2026 \u00b7 03:03 UTC","id":35,"model":"glm-5.2","period_end":"2026-08-16T03:03:01+00:00","period_label":"Aug 15, 2026 \u00b7 23:03 UTC to Aug 16, 2026 \u00b7 03:03 UTC","period_start":"2026-08-15T23:03:01+00:00","prompt_eval_count":4521,"summary":"Across the four-hour window, glm-5.2 is the strongest model with an average throughput of 196.91 token/s, while nemotron-3-ultra is the weakest at 41.71 token/s. The most operationally significant volatility comes from nemotron-3-ultra, which exhibits a coefficient of variation of 42.4% and drops to a minimum of 5.25 token/s, despite a 34.5% upward trend. Conversely, glm-5.2 maintains the highest stability with a coefficient of variation of just 10.2%. The dataset includes 288 valid observations, achieving 100.0% coverage across all 48 expected samples per model, meaning there are no missing-data limitations affecting this analysis.","summary_en":"Across the four-hour window, glm-5.2 is the strongest model with an average throughput of 196.91 token/s, while nemotron-3-ultra is the weakest at 41.71 token/s. The most operationally significant volatility comes from nemotron-3-ultra, which exhibits a coefficient of variation of 42.4% and drops to a minimum of 5.25 token/s, despite a 34.5% upward trend. Conversely, glm-5.2 maintains the highest stability with a coefficient of variation of just 10.2%. The dataset includes 288 valid observations, achieving 100.0% coverage across all 48 expected samples per model, meaning there are no missing-data limitations affecting this analysis.","summary_zh":"\u5728\u56db\u5c0f\u65f6\u7684\u65f6\u95f4\u7a97\u53e3\u5185\uff0cglm-5.2 \u662f\u8868\u73b0\u6700\u5f3a\u7684\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 196.91 \u8bcd\u5143/\u79d2\uff0c\u800c nemotron-3-ultra \u8868\u73b0\u6700\u5f31\uff0c\u4e3a 41.71 \u8bcd\u5143/\u79d2\u3002\u6700\u5177\u64cd\u4f5c\u610f\u4e49\u7684\u6ce2\u52a8\u6027\u6765\u81ea nemotron-3-ultra\uff0c\u5176\u53d8\u5f02\u7cfb\u6570\u4e3a 42.4%\uff0c\u5c3d\u7ba1\u5448 34.5% \u7684\u4e0a\u5347\u8d8b\u52bf\uff0c\u4f46\u6700\u4f4e\u964d\u81f3 5.25 \u8bcd\u5143/\u79d2\u3002\u76f8\u53cd\uff0cglm-5.2 \u4fdd\u6301\u4e86\u6700\u9ad8\u7684\u7a33\u5b9a\u6027\uff0c\u53d8\u5f02\u7cfb\u6570\u4ec5\u4e3a 10.2%\u3002\u6570\u636e\u96c6\u5305\u542b 288 \u4e2a\u6709\u6548\u89c2\u6d4b\u503c\uff0c\u5728\u6240\u6709\u6bcf\u4e2a\u6a21\u578b\u9884\u671f\u7684 48 \u4e2a\u6837\u672c\u4e2d\u5b9e\u73b0\u4e86 100.0% \u7684\u8986\u76d6\u7387\uff0c\u8fd9\u610f\u5473\u7740\u4e0d\u5b58\u5728\u5f71\u54cd\u6b64\u5206\u6790\u7684\u7f3a\u5931\u6570\u636e\u9650\u5236\u3002","total_duration_ns":1047788031,"translated_at":"2026-08-16T03:03:03.521819+00:00","translation_eval_count":155,"translation_model":"glm-5.2","translation_prompt_eval_count":307,"translation_total_duration_ns":640359691,"translation_wall_seconds":0.815,"valid_point_count":288,"wall_seconds":1.243},{"coverage_pct":100.0,"eval_count":141,"generated_at":"2026-08-16T02:03:02.870934+00:00","generated_label":"Aug 16, 2026 \u00b7 02:03 UTC","id":34,"model":"glm-5.2","period_end":"2026-08-16T02:03:01+00:00","period_label":"Aug 15, 2026 \u00b7 22:03 UTC to Aug 16, 2026 \u00b7 02:03 UTC","period_start":"2026-08-15T22:03:01+00:00","prompt_eval_count":4522,"summary":"Across the four-hour window, glm-5.2 delivered the strongest average throughput at 195.03 token/s, while nemotron-3-ultra was weakest at 39.86 token/s. The most operationally significant volatility came from nemotron-3-ultra, which dropped to 5.25 token/s at 23:30 and exhibited a coefficient of variation of 43.9 percent. Deepseek-v4-pro and deepseek-v4-flash also experienced severe transient drops to 8.78 and 14.69 token/s, respectively. Dataset coverage is 100.0 percent with 288 valid points, meaning no missing-data limitation affects this specific analysis window.","summary_en":"Across the four-hour window, glm-5.2 delivered the strongest average throughput at 195.03 token/s, while nemotron-3-ultra was weakest at 39.86 token/s. The most operationally significant volatility came from nemotron-3-ultra, which dropped to 5.25 token/s at 23:30 and exhibited a coefficient of variation of 43.9 percent. Deepseek-v4-pro and deepseek-v4-flash also experienced severe transient drops to 8.78 and 14.69 token/s, respectively. Dataset coverage is 100.0 percent with 288 valid points, meaning no missing-data limitation affects this specific analysis window.","summary_zh":"\u5728\u56db\u5c0f\u65f6\u7684\u65f6\u95f4\u7a97\u53e3\u5185\uff0cglm-5.2 \u4ee5 195.03 \u8bcd\u5143/\u79d2\u7684\u5e73\u5747\u541e\u5410\u91cf\u8868\u73b0\u6700\u5f3a\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u4e3a 39.86 \u8bcd\u5143/\u79d2\u3002\u6700\u5177\u64cd\u4f5c\u610f\u4e49\u7684\u6ce2\u52a8\u6765\u81ea nemotron-3-ultra\uff0c\u5b83\u5728 23:30 \u964d\u81f3 5.25 \u8bcd\u5143/\u79d2\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 43.9%\u3002Deepseek-v4-pro \u548c deepseek-v4-flash \u4e5f\u7ecf\u5386\u4e86\u4e25\u91cd\u7684\u77ac\u65f6\u4e0b\u964d\uff0c\u5206\u522b\u964d\u81f3 8.78 \u548c 14.69 \u8bcd\u5143/\u79d2\u3002\u6570\u636e\u96c6\u8986\u76d6\u7387\u4e3a 100.0%\uff0c\u5305\u542b 288 \u4e2a\u6709\u6548\u70b9\uff0c\u8fd9\u610f\u5473\u7740\u6ca1\u6709\u7f3a\u5931\u6570\u636e\u9650\u5236\u5f71\u54cd\u8fd9\u4e00\u7279\u5b9a\u5206\u6790\u7a97\u53e3\u3002","total_duration_ns":1088032711,"translated_at":"2026-08-16T02:21:37.510584+00:00","translation_eval_count":140,"translation_model":"glm-5.2","translation_prompt_eval_count":246,"translation_total_duration_ns":763033059,"translation_wall_seconds":0.906,"valid_point_count":288,"wall_seconds":1.256},{"coverage_pct":100.0,"eval_count":138,"generated_at":"2026-08-16T01:03:02.516550+00:00","generated_label":"Aug 16, 2026 \u00b7 01:03 UTC","id":33,"model":"glm-5.2","period_end":"2026-08-16T01:03:01+00:00","period_label":"Aug 15, 2026 \u00b7 21:03 UTC to Aug 16, 2026 \u00b7 01:03 UTC","period_start":"2026-08-15T21:03:01+00:00","prompt_eval_count":4521,"summary":"Across the four-hour window, glm-5.2 delivered the strongest average throughput at 188.70 token/s, while nemotron-3-ultra was weakest at 32.04 token/s. Operationally, nemotron-3-ultra exhibited extreme volatility with a 59.3% coefficient of variation, swinging between 4.42 and 67.56 token/s. Additionally, deepseek-v4-pro and deepseek-v4-flash experienced severe transient drops, plummeting to 8.78 and 14.69 token/s respectively. The dataset includes 288 valid observations across six models, achieving 100.0% coverage with no missing data limitations.","summary_en":"Across the four-hour window, glm-5.2 delivered the strongest average throughput at 188.70 token/s, while nemotron-3-ultra was weakest at 32.04 token/s. Operationally, nemotron-3-ultra exhibited extreme volatility with a 59.3% coefficient of variation, swinging between 4.42 and 67.56 token/s. Additionally, deepseek-v4-pro and deepseek-v4-flash experienced severe transient drops, plummeting to 8.78 and 14.69 token/s respectively. The dataset includes 288 valid observations across six models, achieving 100.0% coverage with no missing data limitations.","summary_zh":"\u5728\u56db\u5c0f\u65f6\u7684\u65f6\u95f4\u7a97\u53e3\u5185\uff0cglm-5.2 \u4ee5 188.70 \u8bcd\u5143/\u79d2\u7684\u5e73\u5747\u541e\u5410\u91cf\u8868\u73b0\u6700\u5f3a\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u4e3a 32.04 \u8bcd\u5143/\u79d2\u3002\u5728\u8fd0\u884c\u65b9\u9762\uff0cnemotron-3-ultra \u8868\u73b0\u51fa\u6781\u5927\u7684\u6ce2\u52a8\u6027\uff0c\u53d8\u5f02\u7cfb\u6570\u8fbe 59.3%\uff0c\u5728 4.42 \u548c 67.56 \u8bcd\u5143/\u79d2\u4e4b\u95f4\u6ce2\u52a8\u3002\u6b64\u5916\uff0cdeepseek-v4-pro \u548c deepseek-v4-flash \u7ecf\u5386\u4e86\u4e25\u91cd\u7684\u77ac\u65f6\u4e0b\u964d\uff0c\u5206\u522b\u9aa4\u964d\u81f3 8.78 \u548c 14.69 \u8bcd\u5143/\u79d2\u3002\u8be5\u6570\u636e\u96c6\u5305\u542b\u8de8\u516d\u4e2a\u6a21\u578b\u7684 288 \u4e2a\u6709\u6548\u89c2\u6d4b\u503c\uff0c\u5b9e\u73b0\u4e86 100.0% \u7684\u8986\u76d6\u7387\uff0c\u6ca1\u6709\u7f3a\u5931\u6570\u636e\u9650\u5236\u3002","total_duration_ns":910178712,"translated_at":"2026-08-16T02:21:36.603666+00:00","translation_eval_count":150,"translation_model":"glm-5.2","translation_prompt_eval_count":243,"translation_total_duration_ns":836527145,"translation_wall_seconds":0.987,"valid_point_count":288,"wall_seconds":1.071},{"coverage_pct":100.0,"eval_count":155,"generated_at":"2026-08-16T00:03:02.438631+00:00","generated_label":"Aug 16, 2026 \u00b7 00:03 UTC","id":32,"model":"glm-5.2","period_end":"2026-08-16T00:03:01+00:00","period_label":"Aug 15, 2026 \u00b7 20:03 UTC to Aug 16, 2026 \u00b7 00:03 UTC","period_start":"2026-08-15T20:03:01+00:00","prompt_eval_count":4519,"summary":"Across the four-hour window, glm-5.2 delivered the strongest average throughput at 186.23 token/s, while nemotron-3-ultra was the weakest at 32.32 token/s. Operationally, nemotron-3-ultra exhibited extreme volatility with a 61.5% coefficient of variation, frequently dropping below 10 token/s. Additionally, deepseek-v4-pro and deepseek-v4-flash experienced severe transient slowdowns, plunging to 8.78 token/s and 14.69 token/s respectively. Although the dataset reports 100.0% coverage across all 288 valid samples, the five-minute granularity limits the ability to detect sub-interval micro-outages or pinpoint the exact duration of these sharp throughput drops.","summary_en":"Across the four-hour window, glm-5.2 delivered the strongest average throughput at 186.23 token/s, while nemotron-3-ultra was the weakest at 32.32 token/s. Operationally, nemotron-3-ultra exhibited extreme volatility with a 61.5% coefficient of variation, frequently dropping below 10 token/s. Additionally, deepseek-v4-pro and deepseek-v4-flash experienced severe transient slowdowns, plunging to 8.78 token/s and 14.69 token/s respectively. Although the dataset reports 100.0% coverage across all 288 valid samples, the five-minute granularity limits the ability to detect sub-interval micro-outages or pinpoint the exact duration of these sharp throughput drops.","summary_zh":"\u5728\u6574\u4e2a\u56db\u5c0f\u65f6\u7a97\u53e3\u671f\u5185\uff0cglm-5.2 \u8868\u73b0\u51fa\u6700\u5f3a\u7684\u5e73\u5747\u541e\u5410\u91cf\uff0c\u8fbe\u5230 186.23 \u8bcd\u5143/\u79d2\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u4e3a 32.32 \u8bcd\u5143/\u79d2\u3002\u5728\u8fd0\u884c\u65b9\u9762\uff0cnemotron-3-ultra \u8868\u73b0\u51fa\u6781\u7aef\u7684\u6ce2\u52a8\u6027\uff0c\u53d8\u5f02\u7cfb\u6570\u8fbe 61.5%\uff0c\u9891\u7e41\u8dcc\u7834 10 \u8bcd\u5143/\u79d2\u3002\u6b64\u5916\uff0cdeepseek-v4-pro \u548c deepseek-v4-flash \u7ecf\u5386\u4e86\u4e25\u91cd\u7684\u77ac\u65f6\u51cf\u901f\uff0c\u5206\u522b\u9aa4\u964d\u81f3 8.78 \u8bcd\u5143/\u79d2\u548c 14.69 \u8bcd\u5143/\u79d2\u3002\u5c3d\u7ba1\u6570\u636e\u96c6\u663e\u793a\u6240\u6709 288 \u4e2a\u6709\u6548\u6837\u672c\u7684\u8986\u76d6\u7387\u8fbe\u5230 100.0%\uff0c\u4f46\u4e94\u5206\u949f\u7684\u7c92\u5ea6\u9650\u5236\u4e86\u68c0\u6d4b\u5b50\u533a\u95f4\u5fae\u4e2d\u65ad\u6216\u7cbe\u786e\u5b9a\u4f4d\u8fd9\u4e9b\u541e\u5410\u91cf\u6025\u5267\u4e0b\u964d\u786e\u5207\u6301\u7eed\u65f6\u95f4\u7684\u80fd\u529b\u3002","total_duration_ns":1023606722,"translated_at":"2026-08-16T02:21:35.614305+00:00","translation_eval_count":159,"translation_model":"glm-5.2","translation_prompt_eval_count":260,"translation_total_duration_ns":1007748151,"translation_wall_seconds":1.166,"valid_point_count":288,"wall_seconds":1.221},{"coverage_pct":100.0,"eval_count":149,"generated_at":"2026-08-15T23:03:02.721216+00:00","generated_label":"Aug 15, 2026 \u00b7 23:03 UTC","id":31,"model":"glm-5.2","period_end":"2026-08-15T23:03:01+00:00","period_label":"Aug 15, 2026 \u00b7 19:03 UTC to Aug 15, 2026 \u00b7 23:03 UTC","period_start":"2026-08-15T19:03:01+00:00","prompt_eval_count":4518,"summary":"Across the four-hour window, glm-5.2 delivered the strongest average throughput at 178.42 token/s, while nemotron-3-ultra was the weakest at 33.78 token/s. Operationally, nemotron-3-ultra showed severe volatility with a 58.4% coefficient of variation and a sharp downward trend, dropping from 45.26 token/s at 19:05 to 10.34 token/s by 23:00. Deepseek-v4-flash also exhibited instability, plunging to 14.69 token/s at 22:25. Although dataset coverage is 100.0% with 288 valid points, the four-hour window limits visibility into diurnal load patterns.","summary_en":"Across the four-hour window, glm-5.2 delivered the strongest average throughput at 178.42 token/s, while nemotron-3-ultra was the weakest at 33.78 token/s. Operationally, nemotron-3-ultra showed severe volatility with a 58.4% coefficient of variation and a sharp downward trend, dropping from 45.26 token/s at 19:05 to 10.34 token/s by 23:00. Deepseek-v4-flash also exhibited instability, plunging to 14.69 token/s at 22:25. Although dataset coverage is 100.0% with 288 valid points, the four-hour window limits visibility into diurnal load patterns.","summary_zh":"\u5728\u56db\u5c0f\u65f6\u7a97\u53e3\u5185\uff0cglm-5.2 \u63d0\u4f9b\u4e86\u6700\u5f3a\u7684\u5e73\u5747\u541e\u5410\u91cf\uff0c\u8fbe\u5230 178.42 \u8bcd\u5143/\u79d2\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u4e3a 33.78 \u8bcd\u5143/\u79d2\u3002\u5728\u8fd0\u884c\u65b9\u9762\uff0cnemotron-3-ultra \u8868\u73b0\u51fa\u4e25\u91cd\u7684\u6ce2\u52a8\u6027\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 58.4%\uff0c\u5e76\u5448\u73b0\u6025\u5267\u4e0b\u964d\u8d8b\u52bf\uff0c\u4ece 19:05 \u7684 45.26 \u8bcd\u5143/\u79d2\u964d\u81f3 23:00 \u7684 10.34 \u8bcd\u5143/\u79d2\u3002Deepseek-v4-flash \u4e5f\u8868\u73b0\u51fa\u4e0d\u7a33\u5b9a\u6027\uff0c\u5728 22:25 \u9aa4\u964d\u81f3 14.69 \u8bcd\u5143/\u79d2\u3002\u5c3d\u7ba1\u6570\u636e\u96c6\u8986\u76d6\u7387\u4e3a 100.0%\uff0c\u5305\u542b 288 \u4e2a\u6709\u6548\u70b9\uff0c\u4f46\u56db\u5c0f\u65f6\u7a97\u53e3\u9650\u5236\u4e86\u5bf9\u663c\u591c\u8d1f\u8f7d\u6a21\u5f0f\u7684\u53ef\u89c1\u6027\u3002","total_duration_ns":991041194,"translated_at":"2026-08-16T02:21:34.446980+00:00","translation_eval_count":160,"translation_model":"glm-5.2","translation_prompt_eval_count":254,"translation_total_duration_ns":834708769,"translation_wall_seconds":0.988,"valid_point_count":288,"wall_seconds":1.168},{"coverage_pct":100.0,"eval_count":149,"generated_at":"2026-08-15T22:03:02.822484+00:00","generated_label":"Aug 15, 2026 \u00b7 22:03 UTC","id":30,"model":"glm-5.2","period_end":"2026-08-15T22:03:01+00:00","period_label":"Aug 15, 2026 \u00b7 18:03 UTC to Aug 15, 2026 \u00b7 22:03 UTC","period_start":"2026-08-15T18:03:01+00:00","prompt_eval_count":4514,"summary":"Across the four-hour window, glm-5.2 delivered the strongest average throughput at 175.08 token/s, while nemotron-3-ultra was weakest at 33.08 token/s. Operationally, nemotron-3-ultra showed extreme volatility, swinging between 4.29 and 68.67 token/s with a coefficient of variation of 57.6%, and its throughput trended downward by 21.2%. In contrast, minimax-m3 maintained the most stable performance, averaging 66.99 token/s with a coefficient of variation of just 17.1%. The dataset includes 288 valid observations, achieving 100.0% coverage across all six models with no missing data limitations.","summary_en":"Across the four-hour window, glm-5.2 delivered the strongest average throughput at 175.08 token/s, while nemotron-3-ultra was weakest at 33.08 token/s. Operationally, nemotron-3-ultra showed extreme volatility, swinging between 4.29 and 68.67 token/s with a coefficient of variation of 57.6%, and its throughput trended downward by 21.2%. In contrast, minimax-m3 maintained the most stable performance, averaging 66.99 token/s with a coefficient of variation of just 17.1%. The dataset includes 288 valid observations, achieving 100.0% coverage across all six models with no missing data limitations.","summary_zh":"\u5728\u56db\u5c0f\u65f6\u7684\u65f6\u95f4\u7a97\u53e3\u5185\uff0cglm-5.2 \u7684\u5e73\u5747\u541e\u5410\u91cf\u6700\u9ad8\uff0c\u8fbe\u5230 175.08 \u8bcd\u5143/\u79d2\uff0c\u800c nemotron-3-ultra \u6700\u4f4e\uff0c\u4e3a 33.08 \u8bcd\u5143/\u79d2\u3002\u5728\u8fd0\u884c\u8868\u73b0\u4e0a\uff0cnemotron-3-ultra \u6ce2\u52a8\u6781\u5927\uff0c\u5728 4.29 \u548c 68.67 \u8bcd\u5143/\u79d2\u4e4b\u95f4\u6ce2\u52a8\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 57.6%\uff0c\u5176\u541e\u5410\u91cf\u5448 21.2% \u7684\u4e0b\u964d\u8d8b\u52bf\u3002\u76f8\u6bd4\u4e4b\u4e0b\uff0cminimax-m3 \u4fdd\u6301\u4e86\u6700\u7a33\u5b9a\u7684\u6027\u80fd\uff0c\u5e73\u5747\u4e3a 66.99 \u8bcd\u5143/\u79d2\uff0c\u53d8\u5f02\u7cfb\u6570\u4ec5\u4e3a 17.1%\u3002\u8be5\u6570\u636e\u96c6\u5305\u542b 288 \u4e2a\u6709\u6548\u89c2\u6d4b\u503c\uff0c\u5728\u6240\u6709\u516d\u4e2a\u6a21\u578b\u4e2d\u5b9e\u73b0\u4e86 100.0% \u7684\u8986\u76d6\u7387\uff0c\u6ca1\u6709\u7f3a\u5931\u6570\u636e\u9650\u5236\u3002","total_duration_ns":1021275304,"translated_at":"2026-08-16T02:21:33.455067+00:00","translation_eval_count":157,"translation_model":"glm-5.2","translation_prompt_eval_count":254,"translation_total_duration_ns":761437068,"translation_wall_seconds":0.903,"valid_point_count":288,"wall_seconds":1.192},{"coverage_pct":100.0,"eval_count":142,"generated_at":"2026-08-15T21:03:03.072004+00:00","generated_label":"Aug 15, 2026 \u00b7 21:03 UTC","id":29,"model":"glm-5.2","period_end":"2026-08-15T21:03:01+00:00","period_label":"Aug 15, 2026 \u00b7 17:03 UTC to Aug 15, 2026 \u00b7 21:03 UTC","period_start":"2026-08-15T17:03:01+00:00","prompt_eval_count":4514,"summary":"Across the four-hour window, glm-5.2 delivered the strongest average throughput at 172.69 token/s, while nemotron-3-ultra was the weakest at 37.07 token/s. Operationally, deepseek-v4-pro shows the most significant degrading trend, dropping 16.1 percent to a latest reading of 64.85 token/s. Meanwhile, gemma4:31b exhibited extreme volatility, swinging between a minimum of 32.12 token/s and a maximum of 178.8 token/s. The dataset contains 288 valid observations across six models with 100.0 percent coverage, meaning there are no missing-data limitations affecting this operational review.","summary_en":"Across the four-hour window, glm-5.2 delivered the strongest average throughput at 172.69 token/s, while nemotron-3-ultra was the weakest at 37.07 token/s. Operationally, deepseek-v4-pro shows the most significant degrading trend, dropping 16.1 percent to a latest reading of 64.85 token/s. Meanwhile, gemma4:31b exhibited extreme volatility, swinging between a minimum of 32.12 token/s and a maximum of 178.8 token/s. The dataset contains 288 valid observations across six models with 100.0 percent coverage, meaning there are no missing-data limitations affecting this operational review.","summary_zh":"\u5728\u56db\u5c0f\u65f6\u7684\u65f6\u95f4\u7a97\u53e3\u5185\uff0cglm-5.2 \u63d0\u4f9b\u4e86\u6700\u5f3a\u7684\u5e73\u5747\u541e\u5410\u91cf\uff0c\u8fbe\u5230 172.69 \u8bcd\u5143/\u79d2\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u4e3a 37.07 \u8bcd\u5143/\u79d2\u3002\u5728\u8fd0\u884c\u65b9\u9762\uff0cdeepseek-v4-pro \u8868\u73b0\u51fa\u6700\u663e\u8457\u7684\u4e0b\u964d\u8d8b\u52bf\uff0c\u4e0b\u964d 16.1%\uff0c\u6700\u65b0\u8bfb\u6570\u4e3a 64.85 \u8bcd\u5143/\u79d2\u3002\u4e0e\u6b64\u540c\u65f6\uff0cgemma4:31b \u8868\u73b0\u51fa\u6781\u7aef\u6ce2\u52a8\u6027\uff0c\u5728 32.12 \u8bcd\u5143/\u79d2\u7684\u6700\u5c0f\u503c\u548c 178.8 \u8bcd\u5143/\u79d2\u7684\u6700\u5927\u503c\u4e4b\u95f4\u6446\u52a8\u3002\u8be5\u6570\u636e\u96c6\u5305\u542b\u8de8\u516d\u4e2a\u6a21\u578b\u7684 288 \u4e2a\u6709\u6548\u89c2\u6d4b\u503c\uff0c\u8986\u76d6\u7387\u4e3a 100.0%\uff0c\u8fd9\u610f\u5473\u7740\u4e0d\u5b58\u5728\u5f71\u54cd\u672c\u6b21\u8fd0\u884c\u5ba1\u67e5\u7684\u7f3a\u5931\u6570\u636e\u9650\u5236\u3002","total_duration_ns":964711452,"translated_at":"2026-08-16T02:21:32.550372+00:00","translation_eval_count":155,"translation_model":"glm-5.2","translation_prompt_eval_count":247,"translation_total_duration_ns":793858382,"translation_wall_seconds":0.94,"valid_point_count":288,"wall_seconds":1.153},{"coverage_pct":100.0,"eval_count":146,"generated_at":"2026-08-15T20:03:02.608834+00:00","generated_label":"Aug 15, 2026 \u00b7 20:03 UTC","id":28,"model":"glm-5.2","period_end":"2026-08-15T20:03:01+00:00","period_label":"Aug 15, 2026 \u00b7 16:03 UTC to Aug 15, 2026 \u00b7 20:03 UTC","period_start":"2026-08-15T16:03:01+00:00","prompt_eval_count":4518,"summary":"Across the four-hour window, glm-5.2 delivered the strongest average throughput at 168.37 token/s, while nemotron-3-ultra was weakest at 31.61 token/s. The most operationally significant volatility occurred in nemotron-3-ultra, which logged a 57.1% coefficient of variation and a severe minimum of 2.26 token/s, indicating highly unstable generation. Conversely, deepseek-v4-pro exhibited a notable downward trend, dropping 10.4% to an average of 78.81 token/s. The dataset contains 288 valid points across all six models with 100.0% coverage, meaning there are no missing-data limitations affecting this analysis.","summary_en":"Across the four-hour window, glm-5.2 delivered the strongest average throughput at 168.37 token/s, while nemotron-3-ultra was weakest at 31.61 token/s. The most operationally significant volatility occurred in nemotron-3-ultra, which logged a 57.1% coefficient of variation and a severe minimum of 2.26 token/s, indicating highly unstable generation. Conversely, deepseek-v4-pro exhibited a notable downward trend, dropping 10.4% to an average of 78.81 token/s. The dataset contains 288 valid points across all six models with 100.0% coverage, meaning there are no missing-data limitations affecting this analysis.","summary_zh":"\u5728\u56db\u5c0f\u65f6\u7684\u65f6\u95f4\u7a97\u53e3\u5185\uff0cglm-5.2 \u63d0\u4f9b\u4e86\u6700\u5f3a\u7684\u5e73\u5747\u541e\u5410\u91cf\uff0c\u8fbe\u5230 168.37 \u8bcd\u5143/\u79d2\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u4e3a 31.61 \u8bcd\u5143/\u79d2\u3002\u6700\u5177\u64cd\u4f5c\u610f\u4e49\u7684\u6ce2\u52a8\u53d1\u751f\u5728 nemotron-3-ultra \u4e2d\uff0c\u5176\u8bb0\u5f55\u4e86 57.1% \u7684\u53d8\u5f02\u7cfb\u6570\u548c 2.26 \u8bcd\u5143/\u79d2\u7684\u4e25\u91cd\u6700\u4f4e\u503c\uff0c\u8868\u660e\u751f\u6210\u6781\u4e0d\u7a33\u5b9a\u3002\u76f8\u53cd\uff0cdeepseek-v4-pro \u8868\u73b0\u51fa\u663e\u8457\u7684\u4e0b\u964d\u8d8b\u52bf\uff0c\u4e0b\u964d\u4e86 10.4%\uff0c\u5e73\u5747\u964d\u81f3 78.81 \u8bcd\u5143/\u79d2\u3002\u8be5\u6570\u636e\u96c6\u5305\u542b\u6240\u6709\u516d\u4e2a\u6a21\u578b\u7684 288 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u8986\u76d6\u7387\u4e3a 100.0%\uff0c\u8fd9\u610f\u5473\u7740\u4e0d\u5b58\u5728\u5f71\u54cd\u6b64\u5206\u6790\u7684\u7f3a\u5931\u6570\u636e\u9650\u5236\u3002","total_duration_ns":1008211979,"translated_at":"2026-08-16T02:21:31.607254+00:00","translation_eval_count":149,"translation_model":"glm-5.2","translation_prompt_eval_count":251,"translation_total_duration_ns":732774414,"translation_wall_seconds":0.884,"valid_point_count":288,"wall_seconds":1.469},{"coverage_pct":100.0,"eval_count":150,"generated_at":"2026-08-15T19:03:02.562184+00:00","generated_label":"Aug 15, 2026 \u00b7 19:03 UTC","id":27,"model":"glm-5.2","period_end":"2026-08-15T19:03:01+00:00","period_label":"Aug 15, 2026 \u00b7 15:03 UTC to Aug 15, 2026 \u00b7 19:03 UTC","period_start":"2026-08-15T15:03:01+00:00","prompt_eval_count":4519,"summary":"Over the four-hour window, glm-5.2 delivered the strongest average throughput at 172.86 token/s, while nemotron-3-ultra was weakest at 25.1 token/s. Operationally, nemotron-3-ultra showed extreme volatility, crashing to 1.67 token/s around 15:35 before recovering to 72.03 token/s at 16:45. Deepseek-v4-pro also exhibited sharp early swings, dropping to 11.73 token/s at 15:05. All six models achieved 100.0 percent coverage across 288 valid observations, meaning no missing-data limitation affects this specific dataset. However, the dataset lacks per-request concurrency details, limiting deeper capacity analysis.","summary_en":"Over the four-hour window, glm-5.2 delivered the strongest average throughput at 172.86 token/s, while nemotron-3-ultra was weakest at 25.1 token/s. Operationally, nemotron-3-ultra showed extreme volatility, crashing to 1.67 token/s around 15:35 before recovering to 72.03 token/s at 16:45. Deepseek-v4-pro also exhibited sharp early swings, dropping to 11.73 token/s at 15:05. All six models achieved 100.0 percent coverage across 288 valid observations, meaning no missing-data limitation affects this specific dataset. However, the dataset lacks per-request concurrency details, limiting deeper capacity analysis.","summary_zh":"\u5728\u56db\u5c0f\u65f6\u7684\u65f6\u95f4\u7a97\u53e3\u5185\uff0cglm-5.2 \u63d0\u4f9b\u4e86\u6700\u5f3a\u7684\u5e73\u5747\u541e\u5410\u91cf\uff0c\u8fbe\u5230 172.86 \u8bcd\u5143/\u79d2\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u4e3a 25.1 \u8bcd\u5143/\u79d2\u3002\u5728\u8fd0\u884c\u65b9\u9762\uff0cnemotron-3-ultra \u8868\u73b0\u51fa\u6781\u7aef\u7684\u6ce2\u52a8\u6027\uff0c\u5728 15:35 \u5de6\u53f3\u66b4\u8dcc\u81f3 1.67 \u8bcd\u5143/\u79d2\uff0c\u968f\u540e\u5728 16:45 \u6062\u590d\u81f3 72.03 \u8bcd\u5143/\u79d2\u3002Deepseek-v4-pro \u4e5f\u8868\u73b0\u51fa\u65e9\u671f\u7684\u5267\u70c8\u6ce2\u52a8\uff0c\u5728 15:05 \u964d\u81f3 11.73 \u8bcd\u5143/\u79d2\u3002\u6240\u6709\u516d\u4e2a\u6a21\u578b\u5728 288 \u6b21\u6709\u6548\u89c2\u6d4b\u4e2d\u5747\u5b9e\u73b0\u4e86 100.0% \u7684\u8986\u76d6\u7387\uff0c\u8fd9\u610f\u5473\u7740\u6ca1\u6709\u7f3a\u5931\u6570\u636e\u9650\u5236\u5f71\u54cd\u8fd9\u4e00\u7279\u5b9a\u6570\u636e\u96c6\u3002\u7136\u800c\uff0c\u8be5\u6570\u636e\u96c6\u7f3a\u4e4f\u6bcf\u4e2a\u8bf7\u6c42\u7684\u5e76\u53d1\u7ec6\u8282\uff0c\u9650\u5236\u4e86\u66f4\u6df1\u5c42\u6b21\u7684\u5bb9\u91cf\u5206\u6790\u3002","total_duration_ns":1110461040,"translated_at":"2026-08-16T02:21:30.721150+00:00","translation_eval_count":177,"translation_model":"glm-5.2","translation_prompt_eval_count":255,"translation_total_duration_ns":1242583670,"translation_wall_seconds":1.393,"valid_point_count":288,"wall_seconds":1.276},{"coverage_pct":100.0,"eval_count":147,"generated_at":"2026-08-15T18:03:02.774330+00:00","generated_label":"Aug 15, 2026 \u00b7 18:03 UTC","id":26,"model":"glm-5.2","period_end":"2026-08-15T18:03:01+00:00","period_label":"Aug 15, 2026 \u00b7 14:03 UTC to Aug 15, 2026 \u00b7 18:03 UTC","period_start":"2026-08-15T14:03:01+00:00","prompt_eval_count":4520,"summary":"Across the four-hour window, glm-5.2 delivered the strongest average throughput at 177.87 token/s, while nemotron-3-ultra was weakest at 21.25 token/s. The most operationally significant volatility occurred in nemotron-3-ultra, which exhibited an 84.7% coefficient of variation and dropped to a minimum of 1.67 token/s. Deepseek-v4-pro also showed notable instability, falling to 11.73 token/s. The dataset contains 288 valid observations across six models, achieving 100.0% coverage. Because the dataset spans only four hours, it lacks longer-duration context, limiting the ability to assess diurnal patterns or sustained capacity degradation.","summary_en":"Across the four-hour window, glm-5.2 delivered the strongest average throughput at 177.87 token/s, while nemotron-3-ultra was weakest at 21.25 token/s. The most operationally significant volatility occurred in nemotron-3-ultra, which exhibited an 84.7% coefficient of variation and dropped to a minimum of 1.67 token/s. Deepseek-v4-pro also showed notable instability, falling to 11.73 token/s. The dataset contains 288 valid observations across six models, achieving 100.0% coverage. Because the dataset spans only four hours, it lacks longer-duration context, limiting the ability to assess diurnal patterns or sustained capacity degradation.","summary_zh":"\u5728\u56db\u5c0f\u65f6\u7684\u65f6\u95f4\u7a97\u53e3\u5185\uff0cglm-5.2 \u7684\u5e73\u5747\u541e\u5410\u91cf\u6700\u9ad8\uff0c\u8fbe\u5230 177.87 \u8bcd\u5143/\u79d2\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u4e3a 21.25 \u8bcd\u5143/\u79d2\u3002\u6700\u5177\u64cd\u4f5c\u610f\u4e49\u7684\u6ce2\u52a8\u51fa\u73b0\u5728 nemotron-3-ultra \u4e2d\uff0c\u5176\u53d8\u5f02\u7cfb\u6570\u9ad8\u8fbe 84.7%\uff0c\u5e76\u8dcc\u81f3 1.67 \u8bcd\u5143/\u79d2\u7684\u6700\u4f4e\u503c\u3002Deepseek-v4-pro \u4e5f\u8868\u73b0\u51fa\u660e\u663e\u7684\u4e0d\u7a33\u5b9a\u6027\uff0c\u964d\u81f3 11.73 \u8bcd\u5143/\u79d2\u3002\u8be5\u6570\u636e\u96c6\u5305\u542b\u8de8\u516d\u4e2a\u6a21\u578b\u7684 288 \u4e2a\u6709\u6548\u89c2\u6d4b\u503c\uff0c\u5b9e\u73b0\u4e86 100.0% \u7684\u8986\u76d6\u7387\u3002\u7531\u4e8e\u6570\u636e\u96c6\u4ec5\u8de8\u8d8a\u56db\u5c0f\u65f6\uff0c\u7f3a\u4e4f\u66f4\u957f\u65f6\u95f4\u8de8\u5ea6\u7684\u4e0a\u4e0b\u6587\uff0c\u9650\u5236\u4e86\u5bf9\u663c\u591c\u6a21\u5f0f\u6216\u6301\u7eed\u6027\u80fd\u8870\u9000\u8fdb\u884c\u8bc4\u4f30\u7684\u80fd\u529b\u3002","total_duration_ns":1074188290,"translated_at":"2026-08-16T02:21:29.325454+00:00","translation_eval_count":151,"translation_model":"glm-5.2","translation_prompt_eval_count":252,"translation_total_duration_ns":943652467,"translation_wall_seconds":1.097,"valid_point_count":288,"wall_seconds":1.238},{"coverage_pct":100.0,"eval_count":133,"generated_at":"2026-08-15T17:03:02.824351+00:00","generated_label":"Aug 15, 2026 \u00b7 17:03 UTC","id":25,"model":"glm-5.2","period_end":"2026-08-15T17:03:01+00:00","period_label":"Aug 15, 2026 \u00b7 13:03 UTC to Aug 15, 2026 \u00b7 17:03 UTC","period_start":"2026-08-15T13:03:01+00:00","prompt_eval_count":4520,"summary":"Across the four-hour window, glm-5.2 delivered the strongest average throughput at 183.76 token/s, while nemotron-3-ultra was weakest at 20.1 token/s. Operationally, nemotron-3-ultra exhibited the most significant volatility and degradation, plunging from a 72.07 token/s peak to sustained sub-3 token/s lows with a -39.8 percent trend. Although the dataset reports 100.0 percent coverage across all 288 valid point counts, the five-minute granularity limits the ability to detect sub-interval latency spikes or brief throttling events that would impact real-time user experience.","summary_en":"Across the four-hour window, glm-5.2 delivered the strongest average throughput at 183.76 token/s, while nemotron-3-ultra was weakest at 20.1 token/s. Operationally, nemotron-3-ultra exhibited the most significant volatility and degradation, plunging from a 72.07 token/s peak to sustained sub-3 token/s lows with a -39.8 percent trend. Although the dataset reports 100.0 percent coverage across all 288 valid point counts, the five-minute granularity limits the ability to detect sub-interval latency spikes or brief throttling events that would impact real-time user experience.","summary_zh":"\u5728\u56db\u5c0f\u65f6\u7684\u65f6\u95f4\u7a97\u53e3\u5185\uff0cglm-5.2 \u63d0\u4f9b\u4e86\u6700\u5f3a\u7684\u5e73\u5747\u541e\u5410\u91cf\uff0c\u8fbe\u5230 183.76 \u8bcd\u5143/\u79d2\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u4e3a 20.1 \u8bcd\u5143/\u79d2\u3002\u5728\u8fd0\u884c\u65b9\u9762\uff0cnemotron-3-ultra \u8868\u73b0\u51fa\u6700\u663e\u8457\u7684\u6ce2\u52a8\u548c\u6027\u80fd\u4e0b\u964d\uff0c\u4ece 72.07 \u8bcd\u5143/\u79d2\u7684\u5cf0\u503c\u9aa4\u964d\u81f3\u6301\u7eed\u4f4e\u4e8e 3 \u8bcd\u5143/\u79d2\u7684\u4f4e\u70b9\uff0c\u8d8b\u52bf\u4e3a -39.8%\u3002\u5c3d\u7ba1\u6570\u636e\u96c6\u62a5\u544a\u5728\u6240\u6709 288 \u4e2a\u6709\u6548\u70b9\u8ba1\u6570\u4e0a\u5b9e\u73b0\u4e86 100.0% \u7684\u8986\u76d6\u7387\uff0c\u4f46\u4e94\u5206\u949f\u7684\u7c92\u5ea6\u9650\u5236\u4e86\u5bf9\u5b50\u533a\u95f4\u5ef6\u8fdf\u5cf0\u503c\u6216\u4f1a\u5f71\u54cd\u5b9e\u65f6\u7528\u6237\u4f53\u9a8c\u7684\u77ed\u6682\u9650\u6d41\u4e8b\u4ef6\u7684\u68c0\u6d4b\u80fd\u529b\u3002","total_duration_ns":1132254935,"translated_at":"2026-08-16T02:21:28.225328+00:00","translation_eval_count":143,"translation_model":"glm-5.2","translation_prompt_eval_count":238,"translation_total_duration_ns":848876121,"translation_wall_seconds":1.005,"valid_point_count":288,"wall_seconds":1.328},{"coverage_pct":100.0,"eval_count":148,"generated_at":"2026-08-15T16:03:03.439563+00:00","generated_label":"Aug 15, 2026 \u00b7 16:03 UTC","id":24,"model":"glm-5.2","period_end":"2026-08-15T16:03:01+00:00","period_label":"Aug 15, 2026 \u00b7 12:03 UTC to Aug 15, 2026 \u00b7 16:03 UTC","period_start":"2026-08-15T12:03:01+00:00","prompt_eval_count":4516,"summary":"Across the four-hour window, glm-5.2 delivered the strongest average throughput at 184.47 token/s, while nemotron-3-ultra was weakest at 23.35 token/s. The most operationally significant volatility is nemotron-3-ultra's severe degradation, dropping from a 72.07 token/s peak to a 1.67 token/s minimum with a 85.4% coefficient of variation and a -46.4% trend. Deepseek-v4-flash also showed instability, plunging to 7.06 token/s at 13:00. Dataset coverage is 100.0% across all 288 valid points, meaning there are no missing-data limitations impacting this analysis.","summary_en":"Across the four-hour window, glm-5.2 delivered the strongest average throughput at 184.47 token/s, while nemotron-3-ultra was weakest at 23.35 token/s. The most operationally significant volatility is nemotron-3-ultra's severe degradation, dropping from a 72.07 token/s peak to a 1.67 token/s minimum with a 85.4% coefficient of variation and a -46.4% trend. Deepseek-v4-flash also showed instability, plunging to 7.06 token/s at 13:00. Dataset coverage is 100.0% across all 288 valid points, meaning there are no missing-data limitations impacting this analysis.","summary_zh":"\u5728\u56db\u5c0f\u65f6\u7684\u65f6\u95f4\u7a97\u53e3\u5185\uff0cglm-5.2 \u63d0\u4f9b\u4e86\u6700\u5f3a\u7684\u5e73\u5747\u541e\u5410\u91cf\uff0c\u8fbe\u5230 184.47 \u8bcd\u5143/\u79d2\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u4e3a 23.35 \u8bcd\u5143/\u79d2\u3002\u6700\u5177\u64cd\u4f5c\u610f\u4e49\u7684\u6ce2\u52a8\u662f nemotron-3-ultra \u7684\u4e25\u91cd\u6027\u80fd\u4e0b\u964d\uff0c\u4ece 72.07 \u8bcd\u5143/\u79d2\u7684\u5cf0\u503c\u8dcc\u81f3 1.67 \u8bcd\u5143/\u79d2\u7684\u6700\u4f4e\u503c\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 85.4%\uff0c\u8d8b\u52bf\u4e3a -46.4%\u3002Deepseek-v4-flash \u4e5f\u8868\u73b0\u51fa\u4e0d\u7a33\u5b9a\u6027\uff0c\u5728 13:00 \u9aa4\u964d\u81f3 7.06 \u8bcd\u5143/\u79d2\u3002\u6240\u6709 288 \u4e2a\u6709\u6548\u70b9\u7684\u6570\u636e\u96c6\u8986\u76d6\u7387\u4e3a 100.0%\uff0c\u8fd9\u610f\u5473\u7740\u4e0d\u5b58\u5728\u5f71\u54cd\u672c\u5206\u6790\u7684\u7f3a\u5931\u6570\u636e\u9650\u5236\u3002","total_duration_ns":1206878282,"translated_at":"2026-08-16T02:21:27.217378+00:00","translation_eval_count":151,"translation_model":"glm-5.2","translation_prompt_eval_count":253,"translation_total_duration_ns":770529996,"translation_wall_seconds":0.91,"valid_point_count":288,"wall_seconds":1.544},{"coverage_pct":100.0,"eval_count":128,"generated_at":"2026-08-15T15:03:02.662979+00:00","generated_label":"Aug 15, 2026 \u00b7 15:03 UTC","id":23,"model":"glm-5.2","period_end":"2026-08-15T15:03:01+00:00","period_label":"Aug 15, 2026 \u00b7 11:03 UTC to Aug 15, 2026 \u00b7 15:03 UTC","period_start":"2026-08-15T11:03:01+00:00","prompt_eval_count":4514,"summary":"Across the four-hour window, glm-5.2 delivered the strongest average throughput at 186.74 token/s, while nemotron-3-ultra was the weakest at 25.67 token/s. The most operationally significant volatility occurred in nemotron-3-ultra, which exhibited extreme instability with a coefficient of variation of 78.0% and throughput dropping to a minimum of 2.3 token/s. This level of volatility indicates highly erratic performance degradation compared to the more stable models. The dataset is complete with 100.0% coverage and no missing data limitations across all 48 expected samples per model.","summary_en":"Across the four-hour window, glm-5.2 delivered the strongest average throughput at 186.74 token/s, while nemotron-3-ultra was the weakest at 25.67 token/s. The most operationally significant volatility occurred in nemotron-3-ultra, which exhibited extreme instability with a coefficient of variation of 78.0% and throughput dropping to a minimum of 2.3 token/s. This level of volatility indicates highly erratic performance degradation compared to the more stable models. The dataset is complete with 100.0% coverage and no missing data limitations across all 48 expected samples per model.","summary_zh":"\u5728\u56db\u5c0f\u65f6\u7684\u65f6\u95f4\u7a97\u53e3\u5185\uff0cglm-5.2 \u63d0\u4f9b\u4e86\u6700\u5f3a\u7684\u5e73\u5747\u541e\u5410\u91cf\uff0c\u8fbe\u5230 186.74 \u8bcd\u5143/\u79d2\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u4e3a 25.67 \u8bcd\u5143/\u79d2\u3002\u6700\u5177\u64cd\u4f5c\u610f\u4e49\u7684\u6ce2\u52a8\u6027\u51fa\u73b0\u5728 nemotron-3-ultra \u4e2d\uff0c\u5176\u8868\u73b0\u51fa\u6781\u5ea6\u7684\u4e0d\u7a33\u5b9a\uff0c\u53d8\u5f02\u7cfb\u6570\u9ad8\u8fbe 78.0%\uff0c\u541e\u5410\u91cf\u964d\u81f3\u6700\u4f4e 2.3 \u8bcd\u5143/\u79d2\u3002\u8fd9\u79cd\u6ce2\u52a8\u6c34\u5e73\u8868\u660e\uff0c\u4e0e\u66f4\u7a33\u5b9a\u7684\u6a21\u578b\u76f8\u6bd4\uff0c\u5176\u6027\u80fd\u4e0b\u964d\u6781\u4e0d\u7a33\u5b9a\u3002\u6570\u636e\u96c6\u5b8c\u6574\uff0c\u8986\u76d6\u7387\u4e3a 100.0%\uff0c\u6bcf\u4e2a\u6a21\u578b\u9884\u671f\u7684\u5168\u90e8 48 \u4e2a\u6837\u672c\u5747\u65e0\u7f3a\u5931\u6570\u636e\u9650\u5236\u3002","total_duration_ns":1139079004,"translated_at":"2026-08-16T02:21:26.304476+00:00","translation_eval_count":129,"translation_model":"glm-5.2","translation_prompt_eval_count":233,"translation_total_duration_ns":755101149,"translation_wall_seconds":0.908,"valid_point_count":288,"wall_seconds":1.322},{"coverage_pct":100.0,"eval_count":125,"generated_at":"2026-08-15T14:03:02.802502+00:00","generated_label":"Aug 15, 2026 \u00b7 14:03 UTC","id":22,"model":"glm-5.2","period_end":"2026-08-15T14:03:01+00:00","period_label":"Aug 15, 2026 \u00b7 10:03 UTC to Aug 15, 2026 \u00b7 14:03 UTC","period_start":"2026-08-15T10:03:01+00:00","prompt_eval_count":4516,"summary":"Across the four-hour window, glm-5.2 delivered the strongest average throughput at 181.76 token/s, while nemotron-3-ultra was weakest at 27.42 token/s. The most operationally significant volatility occurred in deepseek-v4-pro, which trended down 18.1 percent and dropped to a minimum of 13.99 token/s. Deepseek-v4-flash also showed severe instability, plummeting to 7.06 token/s at 13:00. Dataset coverage is complete with 288 valid observations, so there are no missing-data limitations affecting this analysis.","summary_en":"Across the four-hour window, glm-5.2 delivered the strongest average throughput at 181.76 token/s, while nemotron-3-ultra was weakest at 27.42 token/s. The most operationally significant volatility occurred in deepseek-v4-pro, which trended down 18.1 percent and dropped to a minimum of 13.99 token/s. Deepseek-v4-flash also showed severe instability, plummeting to 7.06 token/s at 13:00. Dataset coverage is complete with 288 valid observations, so there are no missing-data limitations affecting this analysis.","summary_zh":"\u5728\u56db\u5c0f\u65f6\u7684\u65f6\u95f4\u7a97\u53e3\u5185\uff0cglm-5.2 \u63d0\u4f9b\u4e86\u6700\u5f3a\u7684\u5e73\u5747\u541e\u5410\u91cf\uff0c\u8fbe\u5230 181.76 \u8bcd\u5143/\u79d2\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u4e3a 27.42 \u8bcd\u5143/\u79d2\u3002\u6700\u5177\u64cd\u4f5c\u610f\u4e49\u7684\u6ce2\u52a8\u51fa\u73b0\u5728 deepseek-v4-pro \u4e2d\uff0c\u5176\u8d8b\u52bf\u4e0b\u964d\u4e86 18.1%\uff0c\u5e76\u8dcc\u81f3 13.99 \u8bcd\u5143/\u79d2\u7684\u6700\u4f4e\u503c\u3002Deepseek-v4-flash \u4e5f\u8868\u73b0\u51fa\u4e25\u91cd\u7684\u4e0d\u7a33\u5b9a\u6027\uff0c\u5728 13:00 \u9aa4\u964d\u81f3 7.06 \u8bcd\u5143/\u79d2\u3002\u6570\u636e\u96c6\u8986\u76d6\u5b8c\u6574\uff0c\u5305\u542b 288 \u4e2a\u6709\u6548\u89c2\u6d4b\u503c\uff0c\u56e0\u6b64\u4e0d\u5b58\u5728\u5f71\u54cd\u6b64\u5206\u6790\u7684\u7f3a\u5931\u6570\u636e\u9650\u5236\u3002","total_duration_ns":873786316,"translated_at":"2026-08-16T02:21:25.393506+00:00","translation_eval_count":132,"translation_model":"glm-5.2","translation_prompt_eval_count":230,"translation_total_duration_ns":761069261,"translation_wall_seconds":0.911,"valid_point_count":288,"wall_seconds":1.059},{"coverage_pct":100.0,"eval_count":165,"generated_at":"2026-08-15T13:03:03.173138+00:00","generated_label":"Aug 15, 2026 \u00b7 13:03 UTC","id":21,"model":"glm-5.2","period_end":"2026-08-15T13:03:01+00:00","period_label":"Aug 15, 2026 \u00b7 09:03 UTC to Aug 15, 2026 \u00b7 13:03 UTC","period_start":"2026-08-15T09:03:01+00:00","prompt_eval_count":4518,"summary":"Across the four-hour window, glm-5.2 delivered the strongest average throughput at 182.1 token/s, while nemotron-3-ultra was weakest at 29.76 token/s. The most operationally significant volatility is nemotron-3-ultra's severe late-morning throughput collapse, dropping from 50.35 token/s at 11:15 to 3.23 token/s at 12:05, driving its 60.5% coefficient of variation. Additionally, deepseek-v4-flash and deepseek-v4-pro experienced sharp final-interval drops to 7.06 token/s and 13.99 token/s, respectively. The dataset is complete with 100.0% coverage and 288 valid points, so there are no missing-data limitations affecting this analysis.","summary_en":"Across the four-hour window, glm-5.2 delivered the strongest average throughput at 182.1 token/s, while nemotron-3-ultra was weakest at 29.76 token/s. The most operationally significant volatility is nemotron-3-ultra's severe late-morning throughput collapse, dropping from 50.35 token/s at 11:15 to 3.23 token/s at 12:05, driving its 60.5% coefficient of variation. Additionally, deepseek-v4-flash and deepseek-v4-pro experienced sharp final-interval drops to 7.06 token/s and 13.99 token/s, respectively. The dataset is complete with 100.0% coverage and 288 valid points, so there are no missing-data limitations affecting this analysis.","summary_zh":"\u5728\u8fd9\u56db\u4e2a\u5c0f\u65f6\u7684\u65f6\u95f4\u7a97\u53e3\u5185\uff0cglm-5.2 \u4ee5 182.1 \u8bcd\u5143/\u79d2\u7684\u5e73\u5747\u541e\u5410\u91cf\u8868\u73b0\u6700\u5f3a\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u4e3a 29.76 \u8bcd\u5143/\u79d2\u3002\u6700\u5177\u64cd\u4f5c\u610f\u4e49\u7684\u6ce2\u52a8\u662f nemotron-3-ultra \u5728\u4e0a\u5348\u4e34\u8fd1\u7ed3\u675f\u65f6\u541e\u5410\u91cf\u7684\u4e25\u91cd\u5d29\u6e83\uff0c\u4ece 11:15 \u7684 50.35 \u8bcd\u5143/\u79d2\u964d\u81f3 12:05 \u7684 3.23 \u8bcd\u5143/\u79d2\uff0c\u5bfc\u81f4\u5176\u53d8\u5f02\u7cfb\u6570\u8fbe\u5230 60.5%\u3002\u6b64\u5916\uff0cdeepseek-v4-flash \u548c deepseek-v4-pro \u5728\u6700\u540e\u4e00\u4e2a\u65f6\u95f4\u95f4\u9694\u5206\u522b\u6025\u5267\u4e0b\u964d\u81f3 7.06 \u8bcd\u5143/\u79d2\u548c 13.99 \u8bcd\u5143/\u79d2\u3002\u8be5\u6570\u636e\u96c6\u5b8c\u6574\uff0c\u8986\u76d6\u7387\u4e3a 100.0%\uff0c\u5305\u542b 288 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u56e0\u6b64\u4e0d\u5b58\u5728\u5f71\u54cd\u672c\u5206\u6790\u7684\u7f3a\u5931\u6570\u636e\u9650\u5236\u3002","total_duration_ns":1171865681,"translated_at":"2026-08-16T02:21:24.481383+00:00","translation_eval_count":166,"translation_model":"glm-5.2","translation_prompt_eval_count":270,"translation_total_duration_ns":1033130799,"translation_wall_seconds":1.193,"valid_point_count":288,"wall_seconds":1.362},{"coverage_pct":100.0,"eval_count":139,"generated_at":"2026-08-15T12:03:02.570017+00:00","generated_label":"Aug 15, 2026 \u00b7 12:03 UTC","id":20,"model":"glm-5.2","period_end":"2026-08-15T12:03:01+00:00","period_label":"Aug 15, 2026 \u00b7 08:03 UTC to Aug 15, 2026 \u00b7 12:03 UTC","period_start":"2026-08-15T08:03:01+00:00","prompt_eval_count":4523,"summary":"Across the four-hour window, glm-5.2 delivered the strongest average throughput at 185.4 token/s, while nemotron-3-ultra was weakest at 31.63 token/s. The most operationally significant volatility is nemotron-3-ultra's 64.9% coefficient of variation, including a drop to 0.4 token/s at 08:35 and a declining trend of -37.1% to 4.92 token/s by 12:00. Although dataset coverage is 100.0% with 288 valid points, the five-minute sampling interval limits the detection of sub-five-minute micro-outages or transient latency spikes.","summary_en":"Across the four-hour window, glm-5.2 delivered the strongest average throughput at 185.4 token/s, while nemotron-3-ultra was weakest at 31.63 token/s. The most operationally significant volatility is nemotron-3-ultra's 64.9% coefficient of variation, including a drop to 0.4 token/s at 08:35 and a declining trend of -37.1% to 4.92 token/s by 12:00. Although dataset coverage is 100.0% with 288 valid points, the five-minute sampling interval limits the detection of sub-five-minute micro-outages or transient latency spikes.","summary_zh":"\u5728\u56db\u5c0f\u65f6\u7684\u65f6\u95f4\u7a97\u53e3\u5185\uff0cglm-5.2 \u63d0\u4f9b\u4e86\u6700\u5f3a\u7684\u5e73\u5747\u541e\u5410\u91cf\uff0c\u8fbe\u5230 185.4 \u8bcd\u5143/\u79d2\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u4e3a 31.63 \u8bcd\u5143/\u79d2\u3002\u6700\u5177\u8fd0\u7ef4\u610f\u4e49\u7684\u6ce2\u52a8\u662f nemotron-3-ultra \u9ad8\u8fbe 64.9% \u7684\u53d8\u5f02\u7cfb\u6570\uff0c\u5305\u62ec\u5728 08:35 \u964d\u81f3 0.4 \u8bcd\u5143/\u79d2\uff0c\u4ee5\u53ca\u5230 12:00 \u65f6\u4e0b\u964d -37.1% \u81f3 4.92 \u8bcd\u5143/\u79d2\u7684\u4e0b\u964d\u8d8b\u52bf\u3002\u5c3d\u7ba1\u6570\u636e\u96c6\u8986\u76d6\u7387\u8fbe\u5230 100.0% \u4e14\u5305\u542b 288 \u4e2a\u6709\u6548\u70b9\uff0c\u4f46\u4e94\u5206\u949f\u7684\u91c7\u6837\u95f4\u9694\u9650\u5236\u4e86\u5bf9\u4e94\u5206\u949f\u4ee5\u5185\u7684\u5fae\u4e2d\u65ad\u6216\u77ac\u65f6\u5ef6\u8fdf\u5cf0\u503c\u7684\u68c0\u6d4b\u3002","total_duration_ns":1028224831,"translated_at":"2026-08-16T02:21:23.285323+00:00","translation_eval_count":148,"translation_model":"glm-5.2","translation_prompt_eval_count":244,"translation_total_duration_ns":1006173272,"translation_wall_seconds":1.15,"valid_point_count":288,"wall_seconds":1.357},{"coverage_pct":100.0,"eval_count":127,"generated_at":"2026-08-15T11:03:02.828481+00:00","generated_label":"Aug 15, 2026 \u00b7 11:03 UTC","id":19,"model":"glm-5.2","period_end":"2026-08-15T11:03:01+00:00","period_label":"Aug 15, 2026 \u00b7 07:03 UTC to Aug 15, 2026 \u00b7 11:03 UTC","period_start":"2026-08-15T07:03:01+00:00","prompt_eval_count":4524,"summary":"Across the four-hour dataset, glm-5.2 delivered the strongest average throughput at 177.71 token/s, while nemotron-3-ultra was weakest at 35.06 token/s. The most operationally significant volatility occurred in nemotron-3-ultra, which dropped to 0.4 token/s at 08:35, reflecting severe latency instability. deepseek-v4-flash showed the largest negative trend at -10.6 percent. Although coverage is 100 percent across all models, the dataset is limited by the absence of request concurrency metadata, preventing isolation of tenant load effects on these throughput drops.","summary_en":"Across the four-hour dataset, glm-5.2 delivered the strongest average throughput at 177.71 token/s, while nemotron-3-ultra was weakest at 35.06 token/s. The most operationally significant volatility occurred in nemotron-3-ultra, which dropped to 0.4 token/s at 08:35, reflecting severe latency instability. deepseek-v4-flash showed the largest negative trend at -10.6 percent. Although coverage is 100 percent across all models, the dataset is limited by the absence of request concurrency metadata, preventing isolation of tenant load effects on these throughput drops.","summary_zh":"\u5728\u6574\u4e2a\u56db\u5c0f\u65f6\u7684\u6570\u636e\u96c6\u4e2d\uff0cglm-5.2 \u63d0\u4f9b\u4e86\u6700\u5f3a\u7684\u5e73\u5747\u541e\u5410\u91cf\uff0c\u8fbe\u5230 177.71 \u8bcd\u5143/\u79d2\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u4e3a 35.06 \u8bcd\u5143/\u79d2\u3002\u6700\u5177\u64cd\u4f5c\u610f\u4e49\u7684\u6ce2\u52a8\u51fa\u73b0\u5728 nemotron-3-ultra \u4e2d\uff0c\u8be5\u6a21\u578b\u5728 08:35 \u964d\u81f3 0.4 \u8bcd\u5143/\u79d2\uff0c\u53cd\u6620\u51fa\u4e25\u91cd\u7684\u5ef6\u8fdf\u4e0d\u7a33\u5b9a\u6027\u3002deepseek-v4-flash \u5448\u73b0\u51fa\u6700\u5927\u7684\u8d1f\u5411\u8d8b\u52bf\uff0c\u4e3a -10.6%\u3002\u5c3d\u7ba1\u6240\u6709\u6a21\u578b\u7684\u8986\u76d6\u7387\u5747\u4e3a 100%\uff0c\u4f46\u8be5\u6570\u636e\u96c6\u56e0\u7f3a\u5c11\u8bf7\u6c42\u5e76\u53d1\u5143\u6570\u636e\u800c\u53d7\u5230\u9650\u5236\uff0c\u65e0\u6cd5\u9694\u79bb\u79df\u6237\u8d1f\u8f7d\u5bf9\u8fd9\u4e9b\u541e\u5410\u91cf\u4e0b\u964d\u7684\u5f71\u54cd\u3002","total_duration_ns":1052513118,"translated_at":"2026-08-16T02:21:22.133997+00:00","translation_eval_count":138,"translation_model":"glm-5.2","translation_prompt_eval_count":232,"translation_total_duration_ns":952910040,"translation_wall_seconds":1.11,"valid_point_count":288,"wall_seconds":1.286},{"coverage_pct":100.0,"eval_count":153,"generated_at":"2026-08-15T10:03:03.355046+00:00","generated_label":"Aug 15, 2026 \u00b7 10:03 UTC","id":18,"model":"glm-5.2","period_end":"2026-08-15T10:03:02+00:00","period_label":"Aug 15, 2026 \u00b7 06:03 UTC to Aug 15, 2026 \u00b7 10:03 UTC","period_start":"2026-08-15T06:03:02+00:00","prompt_eval_count":4526,"summary":"Across the four-hour window, glm-5.2 delivered the strongest average throughput at 182.69 token/s, while nemotron-3-ultra was weakest at 37.46 token/s. Operationally, nemotron-3-ultra exhibited the most severe volatility, plunging to 0.4 token/s at 08:35, which indicates intermittent service disruptions. Conversely, minimax-m3 showed a notable declining trend, dropping 11.1 percent to a latest throughput of 25.93 token/s. Although the dataset reports 100.0 percent coverage across all 288 valid point counts, the five-minute observation interval lacks the granularity needed to determine if these extreme throughput drops represent total outages or momentary throttling.","summary_en":"Across the four-hour window, glm-5.2 delivered the strongest average throughput at 182.69 token/s, while nemotron-3-ultra was weakest at 37.46 token/s. Operationally, nemotron-3-ultra exhibited the most severe volatility, plunging to 0.4 token/s at 08:35, which indicates intermittent service disruptions. Conversely, minimax-m3 showed a notable declining trend, dropping 11.1 percent to a latest throughput of 25.93 token/s. Although the dataset reports 100.0 percent coverage across all 288 valid point counts, the five-minute observation interval lacks the granularity needed to determine if these extreme throughput drops represent total outages or momentary throttling.","summary_zh":"\u5728\u6574\u4e2a\u56db\u5c0f\u65f6\u7a97\u53e3\u671f\u5185\uff0cglm-5.2 \u8868\u73b0\u51fa\u6700\u5f3a\u7684\u5e73\u5747\u541e\u5410\u91cf\uff0c\u8fbe\u5230 182.69 \u8bcd\u5143/\u79d2\uff0c\u800c nemotron-3-ultra \u8868\u73b0\u6700\u5f31\uff0c\u4e3a 37.46 \u8bcd\u5143/\u79d2\u3002\u5728\u8fd0\u884c\u65b9\u9762\uff0cnemotron-3-ultra \u8868\u73b0\u51fa\u6700\u4e25\u91cd\u7684\u6ce2\u52a8\u6027\uff0c\u5728 08:35 \u9aa4\u964d\u81f3 0.4 \u8bcd\u5143/\u79d2\uff0c\u8fd9\u8868\u660e\u5b58\u5728\u95f4\u6b47\u6027\u7684\u670d\u52a1\u4e2d\u65ad\u3002\u76f8\u53cd\uff0cminimax-m3 \u5448\u73b0\u51fa\u660e\u663e\u7684\u4e0b\u964d\u8d8b\u52bf\uff0c\u4e0b\u964d\u4e86 11.1%\uff0c\u6700\u65b0\u541e\u5410\u91cf\u4e3a 25.93 \u8bcd\u5143/\u79d2\u3002\u5c3d\u7ba1\u6570\u636e\u96c6\u62a5\u544a\u5728\u6240\u6709 288 \u4e2a\u6709\u6548\u70b9\u8ba1\u6570\u4e0a\u5747\u5b9e\u73b0\u4e86 100.0% \u7684\u8986\u76d6\u7387\uff0c\u4f46\u4e94\u5206\u949f\u7684\u89c2\u6d4b\u95f4\u9694\u7f3a\u4e4f\u786e\u5b9a\u8fd9\u4e9b\u6781\u7aef\u541e\u5410\u91cf\u4e0b\u964d\u662f\u4ee3\u8868\u5b8c\u5168\u4e2d\u65ad\u8fd8\u662f\u77ac\u65f6\u9650\u6d41\u6240\u9700\u7684\u7c92\u5ea6\u3002","total_duration_ns":1188275835,"translated_at":"2026-08-16T02:21:21.022053+00:00","translation_eval_count":170,"translation_model":"glm-5.2","translation_prompt_eval_count":258,"translation_total_duration_ns":1092887515,"translation_wall_seconds":1.244,"valid_point_count":288,"wall_seconds":1.344},{"coverage_pct":100.0,"eval_count":125,"generated_at":"2026-08-15T09:03:02.986710+00:00","generated_label":"Aug 15, 2026 \u00b7 09:03 UTC","id":17,"model":"glm-5.2","period_end":"2026-08-15T09:03:01+00:00","period_label":"Aug 15, 2026 \u00b7 05:03 UTC to Aug 15, 2026 \u00b7 09:03 UTC","period_start":"2026-08-15T05:03:01+00:00","prompt_eval_count":4522,"summary":"Across the four-hour window, glm-5.2 delivered the strongest average throughput at 181.57 token/s, while nemotron-3-ultra was weakest at 35.11 token/s. The most operationally significant volatility is nemotron-3-ultra's severe throughput collapse between 08:30 and 08:35 UTC, where output plummeted from 1.47 token/s to 0.4 token/s before recovering to 28.18 token/s. This dataset contains no missing-data limitation, as all six models achieved 100.0% coverage across their expected 48 samples.","summary_en":"Across the four-hour window, glm-5.2 delivered the strongest average throughput at 181.57 token/s, while nemotron-3-ultra was weakest at 35.11 token/s. The most operationally significant volatility is nemotron-3-ultra's severe throughput collapse between 08:30 and 08:35 UTC, where output plummeted from 1.47 token/s to 0.4 token/s before recovering to 28.18 token/s. This dataset contains no missing-data limitation, as all six models achieved 100.0% coverage across their expected 48 samples.","summary_zh":"\u5728\u56db\u5c0f\u65f6\u7684\u65f6\u95f4\u7a97\u53e3\u5185\uff0cglm-5.2 \u63d0\u4f9b\u4e86\u6700\u5f3a\u7684\u5e73\u5747\u541e\u5410\u91cf\uff0c\u8fbe\u5230 181.57 \u8bcd\u5143/\u79d2\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u4e3a 35.11 \u8bcd\u5143/\u79d2\u3002\u6700\u5177\u64cd\u4f5c\u610f\u4e49\u7684\u6ce2\u52a8\u662f nemotron-3-ultra \u5728 08:30 \u81f3 08:35 UTC \u4e4b\u95f4\u4e25\u91cd\u7684\u541e\u5410\u91cf\u5d29\u6e83\uff0c\u5176\u8f93\u51fa\u4ece 1.47 \u8bcd\u5143/\u79d2\u66b4\u8dcc\u81f3 0.4 \u8bcd\u5143/\u79d2\uff0c\u968f\u540e\u6062\u590d\u81f3 28.18 \u8bcd\u5143/\u79d2\u3002\u6b64\u6570\u636e\u96c6\u4e0d\u5b58\u5728\u7f3a\u5931\u6570\u636e\u7684\u5c40\u9650\u6027\uff0c\u56e0\u4e3a\u6240\u6709\u516d\u4e2a\u6a21\u578b\u5728\u5176\u9884\u671f\u7684 48 \u4e2a\u6837\u672c\u4e2d\u5747\u5b9e\u73b0\u4e86 100.0% \u7684\u8986\u76d6\u7387\u3002","total_duration_ns":1061083509,"translated_at":"2026-08-16T02:21:19.774103+00:00","translation_eval_count":137,"translation_model":"glm-5.2","translation_prompt_eval_count":230,"translation_total_duration_ns":759814719,"translation_wall_seconds":0.924,"valid_point_count":288,"wall_seconds":1.259},{"coverage_pct":100.0,"eval_count":142,"generated_at":"2026-08-15T08:03:02.661559+00:00","generated_label":"Aug 15, 2026 \u00b7 08:03 UTC","id":16,"model":"glm-5.2","period_end":"2026-08-15T08:03:01+00:00","period_label":"Aug 15, 2026 \u00b7 04:03 UTC to Aug 15, 2026 \u00b7 08:03 UTC","period_start":"2026-08-15T04:03:01+00:00","prompt_eval_count":4519,"summary":"Across the four-hour window, glm-5.2 delivered the strongest average throughput at 180.86 token/s, while nemotron-3-ultra was weakest at 30.78 token/s. The most operationally significant volatility appears in glm-5.2, which despite its high average suffered sharp five-minute drops to 39.01 token/s and 58.29 token/s. Nemotron-3-ultra also showed extreme instability, ranging from 4.16 to 69.35 token/s with a 49.2% coefficient of variation. All six models achieved 100% coverage with 288 valid points, so there is no missing-data limitation in this dataset.","summary_en":"Across the four-hour window, glm-5.2 delivered the strongest average throughput at 180.86 token/s, while nemotron-3-ultra was weakest at 30.78 token/s. The most operationally significant volatility appears in glm-5.2, which despite its high average suffered sharp five-minute drops to 39.01 token/s and 58.29 token/s. Nemotron-3-ultra also showed extreme instability, ranging from 4.16 to 69.35 token/s with a 49.2% coefficient of variation. All six models achieved 100% coverage with 288 valid points, so there is no missing-data limitation in this dataset.","summary_zh":"\u5728\u56db\u5c0f\u65f6\u7684\u65f6\u95f4\u7a97\u53e3\u5185\uff0cglm-5.2 \u8868\u73b0\u51fa\u6700\u5f3a\u7684\u5e73\u5747\u541e\u5410\u91cf\uff0c\u8fbe\u5230 180.86 \u8bcd\u5143/\u79d2\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u4e3a 30.78 \u8bcd\u5143/\u79d2\u3002\u6700\u5177\u64cd\u4f5c\u610f\u4e49\u7684\u6ce2\u52a8\u51fa\u73b0\u5728 glm-5.2 \u4e2d\uff0c\u5c3d\u7ba1\u5176\u5e73\u5747\u503c\u8f83\u9ad8\uff0c\u4f46\u5728\u4e94\u5206\u949f\u5185\u5374\u6025\u5267\u4e0b\u964d\u81f3 39.01 \u8bcd\u5143/\u79d2\u548c 58.29 \u8bcd\u5143/\u79d2\u3002Nemotron-3-ultra \u4e5f\u8868\u73b0\u51fa\u6781\u7aef\u7684\u4e0d\u7a33\u5b9a\u6027\uff0c\u8303\u56f4\u4ece 4.16 \u5230 69.35 \u8bcd\u5143/\u79d2\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 49.2%\u3002\u6240\u6709\u516d\u4e2a\u6a21\u578b\u5747\u5b9e\u73b0\u4e86 100% \u7684\u8986\u76d6\u7387\uff0c\u5305\u542b 288 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u56e0\u6b64\u8be5\u6570\u636e\u96c6\u4e0d\u5b58\u5728\u7f3a\u5931\u6570\u636e\u7684\u9650\u5236\u3002","total_duration_ns":967365450,"translated_at":"2026-08-16T02:21:18.846778+00:00","translation_eval_count":150,"translation_model":"glm-5.2","translation_prompt_eval_count":247,"translation_total_duration_ns":715857294,"translation_wall_seconds":0.878,"valid_point_count":288,"wall_seconds":1.134},{"coverage_pct":100.0,"eval_count":133,"generated_at":"2026-08-15T07:03:02.908341+00:00","generated_label":"Aug 15, 2026 \u00b7 07:03 UTC","id":15,"model":"glm-5.2","period_end":"2026-08-15T07:03:01+00:00","period_label":"Aug 15, 2026 \u00b7 03:03 UTC to Aug 15, 2026 \u00b7 07:03 UTC","period_start":"2026-08-15T03:03:01+00:00","prompt_eval_count":4520,"summary":"Across the four-hour window, glm-5.2 delivered the strongest average throughput at 187.27 token/s, while nemotron-3-ultra was weakest at 33.11 token/s. Operationally, deepseek-v4-flash and nemotron-3-ultra showed severe volatility, with coefficients of variation reaching 35.9% and 52.5% respectively, driven by sharp throughput drops to 22.35 token/s and 4.16 token/s. Although dataset coverage is 100.0%, the analysis is limited by the absence of concurrent request counts, preventing differentiation between model-side latency spikes and user-driven load fluctuations.","summary_en":"Across the four-hour window, glm-5.2 delivered the strongest average throughput at 187.27 token/s, while nemotron-3-ultra was weakest at 33.11 token/s. Operationally, deepseek-v4-flash and nemotron-3-ultra showed severe volatility, with coefficients of variation reaching 35.9% and 52.5% respectively, driven by sharp throughput drops to 22.35 token/s and 4.16 token/s. Although dataset coverage is 100.0%, the analysis is limited by the absence of concurrent request counts, preventing differentiation between model-side latency spikes and user-driven load fluctuations.","summary_zh":"\u5728\u56db\u5c0f\u65f6\u7684\u65f6\u95f4\u7a97\u53e3\u5185\uff0cglm-5.2 \u4ee5 187.27 \u8bcd\u5143/\u79d2\u7684\u5e73\u5747\u541e\u5410\u91cf\u8868\u73b0\u6700\u5f3a\uff0c\u800c nemotron-3-ultra \u4ee5 33.11 \u8bcd\u5143/\u79d2\u8868\u73b0\u6700\u5f31\u3002\u5728\u8fd0\u884c\u65b9\u9762\uff0cdeepseek-v4-flash \u548c nemotron-3-ultra \u8868\u73b0\u51fa\u4e25\u91cd\u7684\u6ce2\u52a8\u6027\uff0c\u53d8\u5f02\u7cfb\u6570\u5206\u522b\u8fbe\u5230 35.9% \u548c 52.5%\uff0c\u8fd9\u662f\u7531\u541e\u5410\u91cf\u9aa4\u964d\u81f3 22.35 \u8bcd\u5143/\u79d2\u548c 4.16 \u8bcd\u5143/\u79d2\u6240\u81f4\u3002\u5c3d\u7ba1\u6570\u636e\u96c6\u8986\u76d6\u7387\u4e3a 100.0%\uff0c\u4f46\u5206\u6790\u56e0\u7f3a\u5c11\u5e76\u53d1\u8bf7\u6c42\u6570\u91cf\u800c\u53d7\u9650\uff0c\u4ece\u800c\u65e0\u6cd5\u533a\u5206\u6a21\u578b\u4fa7\u7684\u5ef6\u8fdf\u6fc0\u589e\u4e0e\u7528\u6237\u9a71\u52a8\u7684\u8d1f\u8f7d\u6ce2\u52a8\u3002","total_duration_ns":946757713,"translated_at":"2026-08-16T02:21:17.966088+00:00","translation_eval_count":144,"translation_model":"glm-5.2","translation_prompt_eval_count":238,"translation_total_duration_ns":701502318,"translation_wall_seconds":0.844,"valid_point_count":288,"wall_seconds":1.119},{"coverage_pct":100.0,"eval_count":141,"generated_at":"2026-08-15T06:03:03.305501+00:00","generated_label":"Aug 15, 2026 \u00b7 06:03 UTC","id":14,"model":"glm-5.2","period_end":"2026-08-15T06:03:02+00:00","period_label":"Aug 15, 2026 \u00b7 02:03 UTC to Aug 15, 2026 \u00b7 06:03 UTC","period_start":"2026-08-15T02:03:02+00:00","prompt_eval_count":4516,"summary":"Across the four-hour window, glm-5.2 delivered the strongest average output-token throughput at 185.69 token/s, while nemotron-3-ultra was the weakest at 34.49 token/s. The most operationally significant volatility occurred in nemotron-3-ultra, which dropped sharply from a 69.73 token/s peak to a 4.16 token/s minimum, reflecting a 41.5 percent downward trend. deepseek-v4-pro also showed notable instability, falling to 20.18 token/s. The dataset contains 288 valid observations across six models with 100.0 percent coverage, meaning there are no missing-data limitations affecting this analysis.","summary_en":"Across the four-hour window, glm-5.2 delivered the strongest average output-token throughput at 185.69 token/s, while nemotron-3-ultra was the weakest at 34.49 token/s. The most operationally significant volatility occurred in nemotron-3-ultra, which dropped sharply from a 69.73 token/s peak to a 4.16 token/s minimum, reflecting a 41.5 percent downward trend. deepseek-v4-pro also showed notable instability, falling to 20.18 token/s. The dataset contains 288 valid observations across six models with 100.0 percent coverage, meaning there are no missing-data limitations affecting this analysis.","summary_zh":"\u5728\u6574\u4e2a\u56db\u5c0f\u65f6\u7684\u65f6\u95f4\u7a97\u53e3\u5185\uff0cglm-5.2 \u63d0\u4f9b\u4e86\u6700\u5f3a\u7684\u5e73\u5747\u8f93\u51fa\u8bcd\u5143\u541e\u5410\u91cf\uff0c\u4e3a 185.69 \u8bcd\u5143/\u79d2\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u4e3a 34.49 \u8bcd\u5143/\u79d2\u3002\u6700\u5177\u64cd\u4f5c\u610f\u4e49\u7684\u6ce2\u52a8\u53d1\u751f\u5728 nemotron-3-ultra \u4e0a\uff0c\u5176\u4ece 69.73 \u8bcd\u5143/\u79d2\u7684\u5cf0\u503c\u6025\u5267\u4e0b\u964d\u81f3 4.16 \u8bcd\u5143/\u79d2\u7684\u6700\u4f4e\u503c\uff0c\u53cd\u6620\u51fa 41.5% \u7684\u4e0b\u964d\u8d8b\u52bf\u3002deepseek-v4-pro \u4e5f\u8868\u73b0\u51fa\u660e\u663e\u7684\u4e0d\u7a33\u5b9a\u6027\uff0c\u964d\u81f3 20.18 \u8bcd\u5143/\u79d2\u3002\u8be5\u6570\u636e\u96c6\u5305\u542b\u8de8\u516d\u4e2a\u6a21\u578b\u7684 288 \u4e2a\u6709\u6548\u89c2\u6d4b\u503c\uff0c\u8986\u76d6\u7387\u4e3a 100.0%\uff0c\u8fd9\u610f\u5473\u7740\u4e0d\u5b58\u5728\u5f71\u54cd\u6b64\u5206\u6790\u7684\u7f3a\u5931\u6570\u636e\u9650\u5236\u3002","total_duration_ns":1059117233,"translated_at":"2026-08-16T02:21:17.119507+00:00","translation_eval_count":149,"translation_model":"glm-5.2","translation_prompt_eval_count":246,"translation_total_duration_ns":715394702,"translation_wall_seconds":0.872,"valid_point_count":288,"wall_seconds":1.222},{"coverage_pct":100.0,"eval_count":158,"generated_at":"2026-08-15T05:03:03.342842+00:00","generated_label":"Aug 15, 2026 \u00b7 05:03 UTC","id":13,"model":"glm-5.2","period_end":"2026-08-15T05:03:02+00:00","period_label":"Aug 15, 2026 \u00b7 01:03 UTC to Aug 15, 2026 \u00b7 05:03 UTC","period_start":"2026-08-15T01:03:02+00:00","prompt_eval_count":4517,"summary":"Across the four-hour window, glm-5.2 delivered the strongest average throughput at 190.22 token/s, while nemotron-3-ultra was weakest at 37.63 token/s. The most operationally significant volatility appears in deepseek-v4-flash, which swung sharply between 18.33 and 152.68 token/s with a coefficient of variation of 46.4 percent, indicating highly unstable response times. Additionally, deepseek-v4-pro exhibited a notable downward trend, dropping 19.0 percent over the period and hitting a low of 20.18 token/s. Although coverage is 100.0 percent, the dataset lacks concurrent request volume data, limiting the ability to determine if throughput drops stem from capacity saturation or internal model inefficiencies.","summary_en":"Across the four-hour window, glm-5.2 delivered the strongest average throughput at 190.22 token/s, while nemotron-3-ultra was weakest at 37.63 token/s. The most operationally significant volatility appears in deepseek-v4-flash, which swung sharply between 18.33 and 152.68 token/s with a coefficient of variation of 46.4 percent, indicating highly unstable response times. Additionally, deepseek-v4-pro exhibited a notable downward trend, dropping 19.0 percent over the period and hitting a low of 20.18 token/s. Although coverage is 100.0 percent, the dataset lacks concurrent request volume data, limiting the ability to determine if throughput drops stem from capacity saturation or internal model inefficiencies.","summary_zh":"\u5728\u6574\u4e2a\u56db\u5c0f\u65f6\u7a97\u53e3\u5185\uff0cglm-5.2 \u7684\u5e73\u5747\u541e\u5410\u91cf\u6700\u9ad8\uff0c\u8fbe\u5230 190.22 \u8bcd\u5143/\u79d2\uff0c\u800c nemotron-3-ultra \u6700\u4f4e\uff0c\u4e3a 37.63 \u8bcd\u5143/\u79d2\u3002\u6700\u5177\u64cd\u4f5c\u610f\u4e49\u7684\u6ce2\u52a8\u51fa\u73b0\u5728 deepseek-v4-flash \u4e2d\uff0c\u5176\u5728 18.33 \u548c 152.68 \u8bcd\u5143/\u79d2\u4e4b\u95f4\u5267\u70c8\u6ce2\u52a8\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 46.4%\uff0c\u8868\u660e\u54cd\u5e94\u65f6\u95f4\u6781\u4e0d\u7a33\u5b9a\u3002\u6b64\u5916\uff0cdeepseek-v4-pro \u5448\u73b0\u51fa\u663e\u8457\u7684\u4e0b\u964d\u8d8b\u52bf\uff0c\u5728\u6b64\u671f\u95f4\u4e0b\u964d\u4e86 19.0%\uff0c\u5e76\u89e6\u53ca 20.18 \u8bcd\u5143/\u79d2\u7684\u4f4e\u70b9\u3002\u5c3d\u7ba1\u8986\u76d6\u7387\u4e3a 100.0%\uff0c\u4f46\u6570\u636e\u96c6\u7f3a\u4e4f\u5e76\u53d1\u8bf7\u6c42\u91cf\u6570\u636e\uff0c\u9650\u5236\u4e86\u5224\u65ad\u541e\u5410\u91cf\u4e0b\u964d\u662f\u6e90\u4e8e\u5bb9\u91cf\u9971\u548c\u8fd8\u662f\u5185\u90e8\u6a21\u578b\u4f4e\u6548\u7684\u80fd\u529b\u3002","total_duration_ns":1068821205,"translated_at":"2026-08-16T02:21:16.245628+00:00","translation_eval_count":159,"translation_model":"glm-5.2","translation_prompt_eval_count":263,"translation_total_duration_ns":783331712,"translation_wall_seconds":0.936,"valid_point_count":288,"wall_seconds":1.229},{"coverage_pct":100.0,"eval_count":131,"generated_at":"2026-08-15T04:03:03.247202+00:00","generated_label":"Aug 15, 2026 \u00b7 04:03 UTC","id":12,"model":"glm-5.2","period_end":"2026-08-15T04:03:02+00:00","period_label":"Aug 15, 2026 \u00b7 00:03 UTC to Aug 15, 2026 \u00b7 04:03 UTC","period_start":"2026-08-15T00:03:02+00:00","prompt_eval_count":4516,"summary":"Over the four-hour window, glm-5.2 had the strongest average throughput at 187.89 token/s, while nemotron-3-ultra was weakest at 41.06 token/s. The most operationally significant volatility occurred in deepseek-v4-flash, which exhibited extreme swings between 155.36 token/s and 18.33 token/s, resulting in a coefficient of variation of 52.5%. This erratic performance complicates capacity planning. Although dataset coverage is 100.0% with no missing samples, the analysis is limited by the brief four-hour duration, which prevents assessing diurnal patterns or long-term degradation.","summary_en":"Over the four-hour window, glm-5.2 had the strongest average throughput at 187.89 token/s, while nemotron-3-ultra was weakest at 41.06 token/s. The most operationally significant volatility occurred in deepseek-v4-flash, which exhibited extreme swings between 155.36 token/s and 18.33 token/s, resulting in a coefficient of variation of 52.5%. This erratic performance complicates capacity planning. Although dataset coverage is 100.0% with no missing samples, the analysis is limited by the brief four-hour duration, which prevents assessing diurnal patterns or long-term degradation.","summary_zh":"\u5728\u56db\u5c0f\u65f6\u7684\u65f6\u95f4\u7a97\u53e3\u5185\uff0cglm-5.2 \u7684\u5e73\u5747\u541e\u5410\u91cf\u6700\u9ad8\uff0c\u8fbe\u5230 187.89 \u8bcd\u5143/\u79d2\uff0c\u800c nemotron-3-ultra \u6700\u4f4e\uff0c\u4e3a 41.06 \u8bcd\u5143/\u79d2\u3002\u6700\u5177\u8fd0\u8425\u610f\u4e49\u7684\u6ce2\u52a8\u51fa\u73b0\u5728 deepseek-v4-flash \u4e2d\uff0c\u5176\u5728 155.36 \u8bcd\u5143/\u79d2\u548c 18.33 \u8bcd\u5143/\u79d2\u4e4b\u95f4\u8868\u73b0\u51fa\u5267\u70c8\u6ce2\u52a8\uff0c\u5bfc\u81f4\u53d8\u5f02\u7cfb\u6570\u4e3a 52.5%\u3002\u8fd9\u79cd\u4e0d\u7a33\u5b9a\u7684\u6027\u80fd\u4f7f\u5bb9\u91cf\u89c4\u5212\u53d8\u5f97\u590d\u6742\u3002\u5c3d\u7ba1\u6570\u636e\u96c6\u8986\u76d6\u7387\u4e3a 100.0% \u4e14\u6ca1\u6709\u7f3a\u5931\u6837\u672c\uff0c\u4f46\u5206\u6790\u53d7\u9650\u4e8e\u77ed\u6682\u7684\u56db\u4e2a\u5c0f\u65f6\u6301\u7eed\u65f6\u95f4\uff0c\u65e0\u6cd5\u8bc4\u4f30\u663c\u591c\u6a21\u5f0f\u6216\u957f\u671f\u9000\u5316\u3002","total_duration_ns":1071111432,"translated_at":"2026-08-16T02:21:15.305223+00:00","translation_eval_count":128,"translation_model":"glm-5.2","translation_prompt_eval_count":236,"translation_total_duration_ns":687621366,"translation_wall_seconds":0.833,"valid_point_count":288,"wall_seconds":1.238},{"coverage_pct":100.0,"eval_count":138,"generated_at":"2026-08-15T03:03:02.688616+00:00","generated_label":"Aug 15, 2026 \u00b7 03:03 UTC","id":11,"model":"glm-5.2","period_end":"2026-08-15T03:03:01+00:00","period_label":"Aug 14, 2026 \u00b7 23:03 UTC to Aug 15, 2026 \u00b7 03:03 UTC","period_start":"2026-08-14T23:03:01+00:00","prompt_eval_count":4515,"summary":"Across the four-hour window, glm-5.2 delivered the strongest average throughput at 187.87 token/s, while nemotron-3-ultra was weakest at 38.29 token/s. The most operationally significant volatility occurred in deepseek-v4-flash, which exhibited extreme swings between 155.36 token/s and 18.33 token/s, yielding a coefficient of variation of 50.6 percent. This erratic performance creates unpredictable latency for end users. Although dataset coverage is marked as 100.0 percent with 288 valid points, the five-minute sampling interval lacks sub-minute resolution, limiting the ability to detect brief micro-spike degradations.","summary_en":"Across the four-hour window, glm-5.2 delivered the strongest average throughput at 187.87 token/s, while nemotron-3-ultra was weakest at 38.29 token/s. The most operationally significant volatility occurred in deepseek-v4-flash, which exhibited extreme swings between 155.36 token/s and 18.33 token/s, yielding a coefficient of variation of 50.6 percent. This erratic performance creates unpredictable latency for end users. Although dataset coverage is marked as 100.0 percent with 288 valid points, the five-minute sampling interval lacks sub-minute resolution, limiting the ability to detect brief micro-spike degradations.","summary_zh":"\u5728\u56db\u5c0f\u65f6\u7684\u65f6\u95f4\u7a97\u53e3\u5185\uff0cglm-5.2 \u4ee5 187.87 \u8bcd\u5143/\u79d2\u63d0\u4f9b\u4e86\u6700\u5f3a\u7684\u5e73\u5747\u541e\u5410\u91cf\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u4e3a 38.29 \u8bcd\u5143/\u79d2\u3002\u6700\u5177\u64cd\u4f5c\u610f\u4e49\u7684\u6ce2\u52a8\u51fa\u73b0\u5728 deepseek-v4-flash \u4e2d\uff0c\u5176\u5728 155.36 \u8bcd\u5143/\u79d2\u548c 18.33 \u8bcd\u5143/\u79d2\u4e4b\u95f4\u8868\u73b0\u51fa\u5267\u70c8\u6ce2\u52a8\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 50.6%\u3002\u8fd9\u79cd\u4e0d\u7a33\u5b9a\u7684\u6027\u80fd\u7ed9\u6700\u7ec8\u7528\u6237\u5e26\u6765\u4e86\u4e0d\u53ef\u9884\u6d4b\u7684\u5ef6\u8fdf\u3002\u5c3d\u7ba1\u6570\u636e\u96c6\u8986\u76d6\u7387\u6807\u8bb0\u4e3a 100.0%\uff0c\u5305\u542b 288 \u4e2a\u6709\u6548\u70b9\uff0c\u4f46\u4e94\u5206\u949f\u7684\u91c7\u6837\u95f4\u9694\u7f3a\u4e4f\u4e9a\u5206\u949f\u7ea7\u7684\u5206\u8fa8\u7387\uff0c\u9650\u5236\u4e86\u5bf9\u77ed\u6682\u5fae\u5cf0\u503c\u6027\u80fd\u4e0b\u964d\u7684\u68c0\u6d4b\u80fd\u529b\u3002","total_duration_ns":939504617,"translated_at":"2026-08-16T02:21:14.468624+00:00","translation_eval_count":143,"translation_model":"glm-5.2","translation_prompt_eval_count":243,"translation_total_duration_ns":837369379,"translation_wall_seconds":0.997,"valid_point_count":288,"wall_seconds":1.096},{"coverage_pct":100.0,"eval_count":143,"generated_at":"2026-08-15T02:03:02.481524+00:00","generated_label":"Aug 15, 2026 \u00b7 02:03 UTC","id":10,"model":"glm-5.2","period_end":"2026-08-15T02:03:01+00:00","period_label":"Aug 14, 2026 \u00b7 22:03 UTC to Aug 15, 2026 \u00b7 02:03 UTC","period_start":"2026-08-14T22:03:01+00:00","prompt_eval_count":4517,"summary":"Across the four-hour window, glm-5.2 delivered the strongest average throughput at 186.96 token/s, while nemotron-3-ultra was weakest at 38.17 token/s. The most operationally significant volatility occurred in deepseek-v4-flash, which exhibited extreme swings between 19.96 and 155.36 token/s, reflecting a 51.2% coefficient of variation. Similarly, nemotron-3-ultra suffered severe micro-drops, hitting a low of 1.9 token/s. Although the dataset reports 100.0% coverage across all models, the five-minute sampling interval lacks the granularity needed to isolate the root causes of these sharp transient drops.","summary_en":"Across the four-hour window, glm-5.2 delivered the strongest average throughput at 186.96 token/s, while nemotron-3-ultra was weakest at 38.17 token/s. The most operationally significant volatility occurred in deepseek-v4-flash, which exhibited extreme swings between 19.96 and 155.36 token/s, reflecting a 51.2% coefficient of variation. Similarly, nemotron-3-ultra suffered severe micro-drops, hitting a low of 1.9 token/s. Although the dataset reports 100.0% coverage across all models, the five-minute sampling interval lacks the granularity needed to isolate the root causes of these sharp transient drops.","summary_zh":"\u5728\u6574\u4e2a\u56db\u5c0f\u65f6\u7a97\u53e3\u5185\uff0cglm-5.2 \u63d0\u4f9b\u4e86\u6700\u5f3a\u7684\u5e73\u5747\u541e\u5410\u91cf\uff0c\u8fbe\u5230 186.96 \u8bcd\u5143/\u79d2\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u4e3a 38.17 \u8bcd\u5143/\u79d2\u3002\u6700\u5177\u64cd\u4f5c\u610f\u4e49\u7684\u6ce2\u52a8\u51fa\u73b0\u5728 deepseek-v4-flash \u4e2d\uff0c\u8be5\u6a21\u578b\u5728 19.96 \u548c 155.36 \u8bcd\u5143/\u79d2\u4e4b\u95f4\u8868\u73b0\u51fa\u5267\u70c8\u6ce2\u52a8\uff0c\u53cd\u6620\u51fa 51.2% \u7684\u53d8\u5f02\u7cfb\u6570\u3002\u540c\u6837\uff0cnemotron-3-ultra \u906d\u9047\u4e86\u4e25\u91cd\u7684\u5fae\u8dcc\uff0c\u6700\u4f4e\u964d\u81f3 1.9 \u8bcd\u5143/\u79d2\u3002\u5c3d\u7ba1\u6570\u636e\u96c6\u62a5\u544a\u6240\u6709\u6a21\u578b\u7684\u8986\u76d6\u7387\u4e3a 100.0%\uff0c\u4f46\u4e94\u5206\u949f\u7684\u91c7\u6837\u95f4\u9694\u7f3a\u4e4f\u9694\u79bb\u8fd9\u4e9b\u6025\u5267\u77ac\u65f6\u4e0b\u964d\u6839\u672c\u539f\u56e0\u6240\u9700\u7684\u7c92\u5ea6\u3002","total_duration_ns":996952877,"translated_at":"2026-08-16T02:21:13.468796+00:00","translation_eval_count":145,"translation_model":"glm-5.2","translation_prompt_eval_count":248,"translation_total_duration_ns":809700640,"translation_wall_seconds":0.966,"valid_point_count":288,"wall_seconds":1.183},{"coverage_pct":100.0,"eval_count":157,"generated_at":"2026-08-15T01:03:02.886684+00:00","generated_label":"Aug 15, 2026 \u00b7 01:03 UTC","id":9,"model":"glm-5.2","period_end":"2026-08-15T01:03:01+00:00","period_label":"Aug 14, 2026 \u00b7 21:03 UTC to Aug 15, 2026 \u00b7 01:03 UTC","period_start":"2026-08-14T21:03:01+00:00","prompt_eval_count":4516,"summary":"Across the four-hour window, glm-5.2 delivered the strongest average throughput at 185.62 token/s, while nemotron-3-ultra was weakest at 35.94 token/s. The most operationally significant volatility is deepseek-v4-flash, which exhibited extreme swings between 19.96 and 155.36 token/s with a coefficient of variation of 55.4 percent, indicating highly unstable performance. Additionally, glm-5.2 experienced severe transient drops, plunging to 46.55 token/s at 22:30 and 88.69 token/s at 00:00. The dataset contains 288 valid points across six models with 100.0 percent coverage, meaning there are no missing-data limitations affecting this analysis.","summary_en":"Across the four-hour window, glm-5.2 delivered the strongest average throughput at 185.62 token/s, while nemotron-3-ultra was weakest at 35.94 token/s. The most operationally significant volatility is deepseek-v4-flash, which exhibited extreme swings between 19.96 and 155.36 token/s with a coefficient of variation of 55.4 percent, indicating highly unstable performance. Additionally, glm-5.2 experienced severe transient drops, plunging to 46.55 token/s at 22:30 and 88.69 token/s at 00:00. The dataset contains 288 valid points across six models with 100.0 percent coverage, meaning there are no missing-data limitations affecting this analysis.","summary_zh":"\u5728\u56db\u5c0f\u65f6\u7684\u65f6\u95f4\u7a97\u53e3\u5185\uff0cglm-5.2 \u63d0\u4f9b\u4e86\u6700\u5f3a\u7684\u5e73\u5747\u541e\u5410\u91cf\uff0c\u8fbe\u5230 185.62 \u8bcd\u5143/\u79d2\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u4e3a 35.94 \u8bcd\u5143/\u79d2\u3002\u6700\u5177\u64cd\u4f5c\u610f\u4e49\u7684\u6ce2\u52a8\u51fa\u73b0\u5728 deepseek-v4-flash\uff0c\u8be5\u6a21\u578b\u5728 19.96 \u548c 155.36 \u8bcd\u5143/\u79d2\u4e4b\u95f4\u8868\u73b0\u51fa\u5267\u70c8\u6ce2\u52a8\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 55.4%\uff0c\u8868\u660e\u5176\u6027\u80fd\u9ad8\u5ea6\u4e0d\u7a33\u5b9a\u3002\u6b64\u5916\uff0cglm-5.2 \u7ecf\u5386\u4e86\u4e25\u91cd\u7684\u77ac\u65f6\u4e0b\u964d\uff0c\u5728 22:30 \u9aa4\u964d\u81f3 46.55 \u8bcd\u5143/\u79d2\uff0c\u5e76\u5728 00:00 \u964d\u81f3 88.69 \u8bcd\u5143/\u79d2\u3002\u8be5\u6570\u636e\u96c6\u5305\u542b\u8de8\u516d\u4e2a\u6a21\u578b\u7684 288 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u8986\u76d6\u7387\u4e3a 100.0%\uff0c\u8fd9\u610f\u5473\u7740\u4e0d\u5b58\u5728\u5f71\u54cd\u672c\u5206\u6790\u7684\u7f3a\u5931\u6570\u636e\u9650\u5236\u3002","total_duration_ns":994860169,"translated_at":"2026-08-16T02:21:12.499846+00:00","translation_eval_count":170,"translation_model":"glm-5.2","translation_prompt_eval_count":262,"translation_total_duration_ns":777966201,"translation_wall_seconds":0.931,"valid_point_count":288,"wall_seconds":1.165},{"coverage_pct":100.0,"eval_count":160,"generated_at":"2026-08-15T00:03:02.760596+00:00","generated_label":"Aug 15, 2026 \u00b7 00:03 UTC","id":8,"model":"glm-5.2","period_end":"2026-08-15T00:03:01+00:00","period_label":"Aug 14, 2026 \u00b7 20:03 UTC to Aug 15, 2026 \u00b7 00:03 UTC","period_start":"2026-08-14T20:03:01+00:00","prompt_eval_count":4517,"summary":"Across the four-hour window, glm-5.2 delivered the strongest average throughput at 191.34 token/s, while nemotron-3-ultra was weakest at 33.9 token/s. The most operationally significant volatility is deepseek-v4-flash, which swung sharply between 21.36 and 142.82 token/s with a 59.9 coefficient of variation, indicating highly unstable performance. Additionally, glm-5.2 experienced a severe throughput drop to 46.55 token/s at 22:30 before recovering. Although dataset coverage is 100.0 percent with 288 valid points, the five-minute observation interval limits the ability to detect sub-five-minute microbursts or brief outages, meaning rapid transient degradations are not captured.","summary_en":"Across the four-hour window, glm-5.2 delivered the strongest average throughput at 191.34 token/s, while nemotron-3-ultra was weakest at 33.9 token/s. The most operationally significant volatility is deepseek-v4-flash, which swung sharply between 21.36 and 142.82 token/s with a 59.9 coefficient of variation, indicating highly unstable performance. Additionally, glm-5.2 experienced a severe throughput drop to 46.55 token/s at 22:30 before recovering. Although dataset coverage is 100.0 percent with 288 valid points, the five-minute observation interval limits the ability to detect sub-five-minute microbursts or brief outages, meaning rapid transient degradations are not captured.","summary_zh":"\u5728\u56db\u5c0f\u65f6\u7684\u65f6\u95f4\u7a97\u53e3\u5185\uff0cglm-5.2 \u63d0\u4f9b\u4e86\u6700\u5f3a\u7684\u5e73\u5747\u541e\u5410\u91cf\uff0c\u8fbe\u5230 191.34 \u8bcd\u5143/\u79d2\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u4e3a 33.9 \u8bcd\u5143/\u79d2\u3002\u6700\u5177\u64cd\u4f5c\u610f\u4e49\u7684\u6ce2\u52a8\u6765\u81ea deepseek-v4-flash\uff0c\u5176\u5728 21.36 \u548c 142.82 \u8bcd\u5143/\u79d2\u4e4b\u95f4\u5267\u70c8\u6ce2\u52a8\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 59.9\uff0c\u8868\u660e\u6027\u80fd\u6781\u4e0d\u7a33\u5b9a\u3002\u6b64\u5916\uff0cglm-5.2 \u5728 22:30 \u7ecf\u5386\u4e86\u4e25\u91cd\u7684\u541e\u5410\u91cf\u4e0b\u964d\uff0c\u964d\u81f3 46.55 \u8bcd\u5143/\u79d2\uff0c\u968f\u540e\u6062\u590d\u3002\u5c3d\u7ba1\u6570\u636e\u96c6\u8986\u76d6\u7387\u4e3a 100.0%\uff0c\u5305\u542b 288 \u4e2a\u6709\u6548\u70b9\uff0c\u4f46\u4e94\u5206\u949f\u7684\u89c2\u6d4b\u95f4\u9694\u9650\u5236\u4e86\u5bf9\u4e94\u5206\u949f\u4ee5\u4e0b\u7684\u5fae\u7a81\u53d1\u6216\u77ed\u6682\u4e2d\u65ad\u7684\u68c0\u6d4b\u80fd\u529b\uff0c\u8fd9\u610f\u5473\u7740\u5feb\u901f\u7684\u77ac\u65f6\u6027\u80fd\u4e0b\u964d\u65e0\u6cd5\u88ab\u6355\u83b7\u3002","total_duration_ns":1059195161,"translated_at":"2026-08-16T02:21:11.566114+00:00","translation_eval_count":168,"translation_model":"glm-5.2","translation_prompt_eval_count":265,"translation_total_duration_ns":847676865,"translation_wall_seconds":1.021,"valid_point_count":288,"wall_seconds":1.232},{"coverage_pct":100.0,"eval_count":148,"generated_at":"2026-08-14T23:03:03.025270+00:00","generated_label":"Aug 14, 2026 \u00b7 23:03 UTC","id":7,"model":"glm-5.2","period_end":"2026-08-14T23:03:01+00:00","period_label":"Aug 14, 2026 \u00b7 19:03 UTC to Aug 14, 2026 \u00b7 23:03 UTC","period_start":"2026-08-14T19:03:01+00:00","prompt_eval_count":4521,"summary":"Over the four-hour window, glm-5.2 delivered the strongest average throughput at 194.24 token/s, while nemotron-3-ultra was weakest at 29.28 token/s. The most operationally significant volatility is deepseek-v4-flash, which swung sharply between 21.36 and 135.06 token/s with a coefficient of variation of 60.1 percent. Nemotron-3-ultra also showed extreme instability, dropping to 3.68 token/s before trending upward by 80.8 percent. The dataset records 100.0 percent coverage across all six models with 288 valid points, but the four-hour duration limits any missing-data assessment of longer-term capacity planning.","summary_en":"Over the four-hour window, glm-5.2 delivered the strongest average throughput at 194.24 token/s, while nemotron-3-ultra was weakest at 29.28 token/s. The most operationally significant volatility is deepseek-v4-flash, which swung sharply between 21.36 and 135.06 token/s with a coefficient of variation of 60.1 percent. Nemotron-3-ultra also showed extreme instability, dropping to 3.68 token/s before trending upward by 80.8 percent. The dataset records 100.0 percent coverage across all six models with 288 valid points, but the four-hour duration limits any missing-data assessment of longer-term capacity planning.","summary_zh":"\u5728\u56db\u5c0f\u65f6\u7684\u65f6\u95f4\u7a97\u53e3\u5185\uff0cglm-5.2 \u4ee5 194.24 \u8bcd\u5143/\u79d2\u7684\u5e73\u5747\u541e\u5410\u91cf\u8868\u73b0\u6700\u5f3a\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u4e3a 29.28 \u8bcd\u5143/\u79d2\u3002\u6700\u5177\u64cd\u4f5c\u610f\u4e49\u7684\u6ce2\u52a8\u51fa\u73b0\u5728 deepseek-v4-flash\uff0c\u5176\u6570\u503c\u5728 21.36 \u548c 135.06 \u8bcd\u5143/\u79d2\u4e4b\u95f4\u5267\u70c8\u6ce2\u52a8\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 60.1%\u3002Nemotron-3-ultra \u4e5f\u8868\u73b0\u51fa\u6781\u7aef\u7684\u4e0d\u7a33\u5b9a\u6027\uff0c\u5728\u56de\u5347 80.8% \u4e4b\u524d\u66fe\u964d\u81f3 3.68 \u8bcd\u5143/\u79d2\u3002\u8be5\u6570\u636e\u96c6\u8bb0\u5f55\u4e86\u6240\u6709\u516d\u4e2a\u6a21\u578b 100.0% \u7684\u8986\u76d6\u7387\uff0c\u5305\u542b 288 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u4f46\u56db\u5c0f\u65f6\u7684\u6301\u7eed\u65f6\u95f4\u9650\u5236\u4e86\u5bf9\u957f\u671f\u5bb9\u91cf\u89c4\u5212\u4e2d\u7f3a\u5931\u6570\u636e\u7684\u4efb\u4f55\u8bc4\u4f30\u3002","total_duration_ns":1089862012,"translated_at":"2026-08-16T02:21:10.542872+00:00","translation_eval_count":156,"translation_model":"glm-5.2","translation_prompt_eval_count":253,"translation_total_duration_ns":844887702,"translation_wall_seconds":0.991,"valid_point_count":288,"wall_seconds":1.257},{"coverage_pct":100.0,"eval_count":155,"generated_at":"2026-08-14T22:03:02.855734+00:00","generated_label":"Aug 14, 2026 \u00b7 22:03 UTC","id":6,"model":"glm-5.2","period_end":"2026-08-14T22:03:01+00:00","period_label":"Aug 14, 2026 \u00b7 18:03 UTC to Aug 14, 2026 \u00b7 22:03 UTC","period_start":"2026-08-14T18:03:01+00:00","prompt_eval_count":4517,"summary":"Across the four-hour window, glm-5.2 is the strongest model with an average throughput of 196.78 token/s, while nemotron-3-ultra is the weakest at 23.55 token/s. The most operationally significant volatility comes from deepseek-v4-flash and nemotron-3-ultra, which exhibit extreme throughput swings; deepseek-v4-flash fluctuates between 12.02 and 135.06 token/s, and nemotron-3-ultra varies from 2.64 to 71.78 token/s. This level of instability creates highly unpredictable latency for affected workloads. The dataset shows 100.0 percent coverage with 288 valid points, meaning there are no missing-data limitations impacting this specific analysis.","summary_en":"Across the four-hour window, glm-5.2 is the strongest model with an average throughput of 196.78 token/s, while nemotron-3-ultra is the weakest at 23.55 token/s. The most operationally significant volatility comes from deepseek-v4-flash and nemotron-3-ultra, which exhibit extreme throughput swings; deepseek-v4-flash fluctuates between 12.02 and 135.06 token/s, and nemotron-3-ultra varies from 2.64 to 71.78 token/s. This level of instability creates highly unpredictable latency for affected workloads. The dataset shows 100.0 percent coverage with 288 valid points, meaning there are no missing-data limitations impacting this specific analysis.","summary_zh":"\u5728\u56db\u5c0f\u65f6\u7684\u65f6\u95f4\u7a97\u53e3\u5185\uff0cglm-5.2 \u662f\u8868\u73b0\u6700\u5f3a\u7684\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 196.78 \u8bcd\u5143/\u79d2\uff0c\u800c nemotron-3-ultra \u662f\u8868\u73b0\u6700\u5f31\u7684\uff0c\u4e3a 23.55 \u8bcd\u5143/\u79d2\u3002\u6700\u5177\u64cd\u4f5c\u610f\u4e49\u7684\u6ce2\u52a8\u6765\u81ea deepseek-v4-flash \u548c nemotron-3-ultra\uff0c\u5b83\u4eec\u8868\u73b0\u51fa\u6781\u7aef\u7684\u541e\u5410\u91cf\u6446\u52a8\uff1bdeepseek-v4-flash \u5728 12.02 \u81f3 135.06 \u8bcd\u5143/\u79d2\u4e4b\u95f4\u6ce2\u52a8\uff0cnemotron-3-ultra \u5728 2.64 \u81f3 71.78 \u8bcd\u5143/\u79d2\u4e4b\u95f4\u53d8\u5316\u3002\u8fd9\u79cd\u7a0b\u5ea6\u7684\u4e0d\u7a33\u5b9a\u6027\u4e3a\u53d7\u5f71\u54cd\u7684\u5de5\u4f5c\u8d1f\u8f7d\u9020\u6210\u4e86\u9ad8\u5ea6\u4e0d\u53ef\u9884\u6d4b\u7684\u5ef6\u8fdf\u3002\u6570\u636e\u96c6\u663e\u793a 100.0% \u7684\u8986\u76d6\u7387\uff0c\u5305\u542b 288 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u8fd9\u610f\u5473\u7740\u6ca1\u6709\u7f3a\u5931\u6570\u636e\u7684\u9650\u5236\u5f71\u54cd\u8fd9\u4e00\u7279\u5b9a\u5206\u6790\u3002","total_duration_ns":1123850093,"translated_at":"2026-08-16T02:21:09.549024+00:00","translation_eval_count":164,"translation_model":"glm-5.2","translation_prompt_eval_count":260,"translation_total_duration_ns":926642221,"translation_wall_seconds":1.076,"valid_point_count":288,"wall_seconds":1.287},{"coverage_pct":100.0,"eval_count":168,"generated_at":"2026-08-14T21:03:02.944484+00:00","generated_label":"Aug 14, 2026 \u00b7 21:03 UTC","id":5,"model":"glm-5.2","period_end":"2026-08-14T21:03:01+00:00","period_label":"Aug 14, 2026 \u00b7 17:03 UTC to Aug 14, 2026 \u00b7 21:03 UTC","period_start":"2026-08-14T17:03:01+00:00","prompt_eval_count":4519,"summary":"Across the four-hour window, glm-5.2 delivered the strongest average throughput at 193.53 token/s, while nemotron-3-ultra was weakest at 24.61 token/s. Operationally, deepseek-v4-flash exhibited severe volatility with a 50.8% coefficient of variation, swinging between 11.88 and 133.32 token/s. Additionally, nemotron-3-ultra experienced a sharp operational degradation, dropping from 44.21 token/s at 17:05 to a minimum of 2.64 token/s at 18:30, reflecting a -26.5% trend. The dataset contains 288 valid observations, achieving 100.0% coverage across all 48 expected samples per model, meaning there are no missing-data limitations to constrain this analysis.","summary_en":"Across the four-hour window, glm-5.2 delivered the strongest average throughput at 193.53 token/s, while nemotron-3-ultra was weakest at 24.61 token/s. Operationally, deepseek-v4-flash exhibited severe volatility with a 50.8% coefficient of variation, swinging between 11.88 and 133.32 token/s. Additionally, nemotron-3-ultra experienced a sharp operational degradation, dropping from 44.21 token/s at 17:05 to a minimum of 2.64 token/s at 18:30, reflecting a -26.5% trend. The dataset contains 288 valid observations, achieving 100.0% coverage across all 48 expected samples per model, meaning there are no missing-data limitations to constrain this analysis.","summary_zh":"\u5728\u56db\u5c0f\u65f6\u7684\u65f6\u95f4\u7a97\u53e3\u5185\uff0cglm-5.2 \u4ee5 193.53 \u8bcd\u5143/\u79d2\u7684\u5e73\u5747\u541e\u5410\u91cf\u8868\u73b0\u6700\u5f3a\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u4e3a 24.61 \u8bcd\u5143/\u79d2\u3002\u5728\u8fd0\u884c\u65b9\u9762\uff0cdeepseek-v4-flash \u8868\u73b0\u51fa\u4e25\u91cd\u7684\u6ce2\u52a8\u6027\uff0c\u53d8\u5f02\u7cfb\u6570\u8fbe 50.8%\uff0c\u5728 11.88 \u548c 133.32 \u8bcd\u5143/\u79d2\u4e4b\u95f4\u6446\u52a8\u3002\u6b64\u5916\uff0cnemotron-3-ultra \u7ecf\u5386\u4e86\u6025\u5267\u7684\u8fd0\u884c\u6027\u80fd\u4e0b\u964d\uff0c\u4ece 17:05 \u7684 44.21 \u8bcd\u5143/\u79d2\u8dcc\u81f3 18:30 \u7684\u6700\u4f4e 2.64 \u8bcd\u5143/\u79d2\uff0c\u53cd\u6620\u51fa -26.5% \u7684\u8d8b\u52bf\u3002\u8be5\u6570\u636e\u96c6\u5305\u542b 288 \u4e2a\u6709\u6548\u89c2\u6d4b\u503c\uff0c\u5728\u6240\u6709 48 \u4e2a\u6bcf\u6a21\u578b\u9884\u671f\u6837\u672c\u4e2d\u5b9e\u73b0\u4e86 100.0% \u7684\u8986\u76d6\u7387\uff0c\u8fd9\u610f\u5473\u7740\u4e0d\u5b58\u5728\u7f3a\u5931\u6570\u636e\u9650\u5236\u6765\u7ea6\u675f\u672c\u5206\u6790\u3002","total_duration_ns":1200739854,"translated_at":"2026-08-16T02:21:08.470410+00:00","translation_eval_count":183,"translation_model":"glm-5.2","translation_prompt_eval_count":273,"translation_total_duration_ns":824763871,"translation_wall_seconds":0.968,"valid_point_count":288,"wall_seconds":1.364},{"coverage_pct":93.8,"eval_count":154,"generated_at":"2026-08-14T20:03:03.065986+00:00","generated_label":"Aug 14, 2026 \u00b7 20:03 UTC","id":4,"model":"glm-5.2","period_end":"2026-08-14T20:03:01+00:00","period_label":"Aug 14, 2026 \u00b7 16:03 UTC to Aug 14, 2026 \u00b7 20:03 UTC","period_start":"2026-08-14T16:03:01+00:00","prompt_eval_count":4283,"summary":"Across the four-hour window, glm-5.2 delivered the strongest average throughput at 186.2 token/s, while nemotron-3-ultra was weakest at 26.77 token/s. Operationally, nemotron-3-ultra exhibited severe volatility and a sharp degradation, dropping from 56.06 token/s at 16:20 to 2.64 token/s at 18:30. deepseek-v4-flash also showed instability, spiking to 159.21 token/s at 16:25 before crashing to 8.22 token/s at 16:50. This analysis is limited by missing data; each model recorded 45 valid samples out of an expected 48, resulting in 93.8 percent coverage.","summary_en":"Across the four-hour window, glm-5.2 delivered the strongest average throughput at 186.2 token/s, while nemotron-3-ultra was weakest at 26.77 token/s. Operationally, nemotron-3-ultra exhibited severe volatility and a sharp degradation, dropping from 56.06 token/s at 16:20 to 2.64 token/s at 18:30. deepseek-v4-flash also showed instability, spiking to 159.21 token/s at 16:25 before crashing to 8.22 token/s at 16:50. This analysis is limited by missing data; each model recorded 45 valid samples out of an expected 48, resulting in 93.8 percent coverage.","summary_zh":"\u5728\u6574\u4e2a\u56db\u5c0f\u65f6\u7a97\u53e3\u671f\u5185\uff0cglm-5.2 \u8868\u73b0\u51fa\u6700\u5f3a\u7684\u5e73\u5747\u541e\u5410\u91cf\uff0c\u8fbe\u5230 186.2 \u8bcd\u5143/\u79d2\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u4e3a 26.77 \u8bcd\u5143/\u79d2\u3002\u5728\u8fd0\u884c\u65b9\u9762\uff0cnemotron-3-ultra \u8868\u73b0\u51fa\u4e25\u91cd\u7684\u6ce2\u52a8\u6027\u548c\u6025\u5267\u7684\u6027\u80fd\u4e0b\u964d\uff0c\u4ece 16:20 \u7684 56.06 \u8bcd\u5143/\u79d2\u8dcc\u843d\u81f3 18:30 \u7684 2.64 \u8bcd\u5143/\u79d2\u3002deepseek-v4-flash \u540c\u6837\u663e\u793a\u51fa\u4e0d\u7a33\u5b9a\u6027\uff0c\u5728 16:25 \u98d9\u5347\u81f3 159.21 \u8bcd\u5143/\u79d2\uff0c\u968f\u540e\u5728 16:50 \u66b4\u8dcc\u81f3 8.22 \u8bcd\u5143/\u79d2\u3002\u672c\u5206\u6790\u53d7\u9650\u4e8e\u6570\u636e\u7f3a\u5931\uff1b\u6bcf\u4e2a\u6a21\u578b\u5728\u9884\u671f\u7684 48 \u4e2a\u6837\u672c\u4e2d\u8bb0\u5f55\u4e86 45 \u4e2a\u6709\u6548\u6837\u672c\uff0c\u8986\u76d6\u7387\u4e3a 93.8%\u3002","total_duration_ns":944132086,"translated_at":"2026-08-16T02:21:07.499145+00:00","translation_eval_count":172,"translation_model":"glm-5.2","translation_prompt_eval_count":259,"translation_total_duration_ns":910601009,"translation_wall_seconds":1.061,"valid_point_count":270,"wall_seconds":1.123},{"coverage_pct":68.8,"eval_count":156,"generated_at":"2026-08-14T19:03:02.718819+00:00","generated_label":"Aug 14, 2026 \u00b7 19:03 UTC","id":3,"model":"glm-5.2","period_end":"2026-08-14T19:03:01+00:00","period_label":"Aug 14, 2026 \u00b7 15:03 UTC to Aug 14, 2026 \u00b7 19:03 UTC","period_start":"2026-08-14T15:03:01+00:00","prompt_eval_count":3344,"summary":"Across the four-hour window, glm-5.2 delivered the strongest average throughput at 181.14 token/s, while nemotron-3-ultra was the weakest at 31.55 token/s. The most operationally significant volatility occurred in nemotron-3-ultra, which experienced a severe downward trend of 25.8 percent, plummeting to a minimum of 2.64 token/s. Similarly, deepseek-v4-flash showed high instability with a 44.4 percent coefficient of variation and repeated sharp drops below 15 token/s. A key limitation is that the dataset contains only 198 valid points out of an expected 288, resulting in 68.8 percent coverage. This missing data prevents a complete assessment of the full period.","summary_en":"Across the four-hour window, glm-5.2 delivered the strongest average throughput at 181.14 token/s, while nemotron-3-ultra was the weakest at 31.55 token/s. The most operationally significant volatility occurred in nemotron-3-ultra, which experienced a severe downward trend of 25.8 percent, plummeting to a minimum of 2.64 token/s. Similarly, deepseek-v4-flash showed high instability with a 44.4 percent coefficient of variation and repeated sharp drops below 15 token/s. A key limitation is that the dataset contains only 198 valid points out of an expected 288, resulting in 68.8 percent coverage. This missing data prevents a complete assessment of the full period.","summary_zh":"\u5728\u56db\u5c0f\u65f6\u7684\u65f6\u95f4\u7a97\u53e3\u5185\uff0cglm-5.2 \u63d0\u4f9b\u4e86\u6700\u5f3a\u7684\u5e73\u5747\u541e\u5410\u91cf\uff0c\u8fbe\u5230 181.14 \u8bcd\u5143/\u79d2\uff0c\u800c nemotron-3-ultra \u8868\u73b0\u6700\u5f31\uff0c\u4e3a 31.55 \u8bcd\u5143/\u79d2\u3002\u6700\u5177\u64cd\u4f5c\u610f\u4e49\u7684\u6ce2\u52a8\u51fa\u73b0\u5728 nemotron-3-ultra \u4e2d\uff0c\u8be5\u6a21\u578b\u7ecf\u5386\u4e86 25.8% \u7684\u4e25\u91cd\u4e0b\u884c\u8d8b\u52bf\uff0c\u9aa4\u964d\u81f3 2.64 \u8bcd\u5143/\u79d2\u7684\u6700\u4f4e\u70b9\u3002\u540c\u6837\uff0cdeepseek-v4-flash \u8868\u73b0\u51fa\u9ad8\u5ea6\u7684\u4e0d\u7a33\u5b9a\u6027\uff0c\u5176\u53d8\u5f02\u7cfb\u6570\u4e3a 44.4%\uff0c\u5e76\u591a\u6b21\u51fa\u73b0\u8dcc\u7834 15 \u8bcd\u5143/\u79d2\u7684\u6025\u5267\u4e0b\u964d\u3002\u4e00\u4e2a\u5173\u952e\u9650\u5236\u662f\uff0c\u8be5\u6570\u636e\u96c6\u5728\u9884\u671f\u7684 288 \u4e2a\u6570\u636e\u70b9\u4e2d\u4ec5\u5305\u542b 198 \u4e2a\u6709\u6548\u70b9\uff0c\u5bfc\u81f4\u8986\u76d6\u7387\u4e3a 68.8%\u3002\u8fd9\u4e9b\u7f3a\u5931\u6570\u636e\u4f7f\u5f97\u65e0\u6cd5\u5bf9\u6574\u4e2a\u65f6\u95f4\u6bb5\u8fdb\u884c\u5b8c\u6574\u8bc4\u4f30\u3002","total_duration_ns":1109998781,"translated_at":"2026-08-16T02:21:06.434487+00:00","translation_eval_count":170,"translation_model":"glm-5.2","translation_prompt_eval_count":261,"translation_total_duration_ns":832035277,"translation_wall_seconds":0.975,"valid_point_count":198,"wall_seconds":1.308},{"coverage_pct":43.8,"eval_count":142,"generated_at":"2026-08-14T18:03:02.892298+00:00","generated_label":"Aug 14, 2026 \u00b7 18:03 UTC","id":2,"model":"glm-5.2","period_end":"2026-08-14T18:03:01+00:00","period_label":"Aug 14, 2026 \u00b7 14:03 UTC to Aug 14, 2026 \u00b7 18:03 UTC","period_start":"2026-08-14T14:03:01+00:00","prompt_eval_count":2406,"summary":"Over the four-hour window, glm-5.2 is the strongest model with an average throughput of 173.55 token/s, while nemotron-3-ultra is the weakest at 37.89 token/s. Operationally, deepseek-v4-flash exhibits the most significant volatility, dropping to extreme lows of 8.22 token/s and 11.88 token/s, yielding a high coefficient of variation of 41.4 percent. All models have exactly 21 samples each, resulting in a dataset coverage of only 43.8 percent. This missing-data limitation restricts visibility into the first 140 minutes of the period, meaning the calculated averages may not represent full operational capacity.","summary_en":"Over the four-hour window, glm-5.2 is the strongest model with an average throughput of 173.55 token/s, while nemotron-3-ultra is the weakest at 37.89 token/s. Operationally, deepseek-v4-flash exhibits the most significant volatility, dropping to extreme lows of 8.22 token/s and 11.88 token/s, yielding a high coefficient of variation of 41.4 percent. All models have exactly 21 samples each, resulting in a dataset coverage of only 43.8 percent. This missing-data limitation restricts visibility into the first 140 minutes of the period, meaning the calculated averages may not represent full operational capacity.","summary_zh":"\u5728\u56db\u5c0f\u65f6\u7684\u65f6\u95f4\u7a97\u53e3\u5185\uff0cglm-5.2 \u662f\u8868\u73b0\u6700\u5f3a\u7684\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 173.55 \u8bcd\u5143/\u79d2\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u4e3a 37.89 \u8bcd\u5143/\u79d2\u3002\u5728\u8fd0\u884c\u65b9\u9762\uff0cdeepseek-v4-flash \u8868\u73b0\u51fa\u6700\u663e\u8457\u7684\u6ce2\u52a8\u6027\uff0c\u964d\u81f3 8.22 \u8bcd\u5143/\u79d2\u548c 11.88 \u8bcd\u5143/\u79d2\u7684\u6781\u4f4e\u503c\uff0c\u5bfc\u81f4\u5176\u53d8\u5f02\u7cfb\u6570\u9ad8\u8fbe 41.4%\u3002\u6240\u6709\u6a21\u578b\u5404\u6709\u6070\u597d 21 \u4e2a\u6837\u672c\uff0c\u5bfc\u81f4\u6570\u636e\u96c6\u8986\u76d6\u7387\u4ec5\u4e3a 43.8%\u3002\u8fd9\u79cd\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\u4f7f\u5f97\u5bf9\u8be5\u65f6\u6bb5\u524d 140 \u5206\u949f\u7684\u60c5\u51b5\u65e0\u6cd5\u89c2\u6d4b\uff0c\u8fd9\u610f\u5473\u7740\u8ba1\u7b97\u51fa\u7684\u5e73\u5747\u503c\u53ef\u80fd\u65e0\u6cd5\u4ee3\u8868\u5b8c\u6574\u7684\u8fd0\u884c\u80fd\u529b\u3002","total_duration_ns":1154661116,"translated_at":"2026-08-16T02:21:05.455547+00:00","translation_eval_count":139,"translation_model":"glm-5.2","translation_prompt_eval_count":247,"translation_total_duration_ns":797456418,"translation_wall_seconds":0.964,"valid_point_count":126,"wall_seconds":1.333},{"coverage_pct":31.2,"eval_count":142,"generated_at":"2026-08-14T17:31:27.604225+00:00","generated_label":"Aug 14, 2026 \u00b7 17:31 UTC","id":1,"model":"glm-5.2","period_end":"2026-08-14T17:31:26+00:00","period_label":"Aug 14, 2026 \u00b7 13:31 UTC to Aug 14, 2026 \u00b7 17:31 UTC","period_start":"2026-08-14T13:31:26+00:00","prompt_eval_count":1938,"summary":"Across the rolling four-hour window, glm-5.2 is the strongest model by average throughput at 164.88 token/s, while nemotron-3-ultra is the weakest at 38.09 token/s. The most operationally significant volatility occurs in deepseek-v4-flash, which exhibits severe throughput instability; despite an average of 80.49 token/s, it drops to extreme lows of 8.22 token/s and 11.88 token/s. This evaluation is constrained by a major missing-data limitation. The dataset contains only 90 valid observations out of an expected 288, representing 31.2% coverage, leaving the majority of the period unmonitored.","summary_en":"Across the rolling four-hour window, glm-5.2 is the strongest model by average throughput at 164.88 token/s, while nemotron-3-ultra is the weakest at 38.09 token/s. The most operationally significant volatility occurs in deepseek-v4-flash, which exhibits severe throughput instability; despite an average of 80.49 token/s, it drops to extreme lows of 8.22 token/s and 11.88 token/s. This evaluation is constrained by a major missing-data limitation. The dataset contains only 90 valid observations out of an expected 288, representing 31.2% coverage, leaving the majority of the period unmonitored.","summary_zh":"\u5728\u6eda\u52a8\u7684\u56db\u5c0f\u65f6\u7a97\u53e3\u5185\uff0cglm-5.2 \u662f\u5e73\u5747\u541e\u5410\u91cf\u6700\u5f3a\u7684\u6a21\u578b\uff0c\u8fbe\u5230 164.88 \u8bcd\u5143/\u79d2\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u4e3a 38.09 \u8bcd\u5143/\u79d2\u3002\u6700\u5177\u64cd\u4f5c\u610f\u4e49\u7684\u6ce2\u52a8\u51fa\u73b0\u5728 deepseek-v4-flash \u4e2d\uff0c\u5176\u8868\u73b0\u51fa\u4e25\u91cd\u7684\u541e\u5410\u91cf\u4e0d\u7a33\u5b9a\u6027\uff1b\u5c3d\u7ba1\u5e73\u5747\u503c\u4e3a 80.49 \u8bcd\u5143/\u79d2\uff0c\u4f46\u5b83\u8dcc\u81f3 8.22 \u8bcd\u5143/\u79d2\u548c 11.88 \u8bcd\u5143/\u79d2\u7684\u6781\u7aef\u4f4e\u70b9\u3002\u6b64\u9879\u8bc4\u4f30\u53d7\u5230\u91cd\u5927\u7f3a\u5931\u6570\u636e\u5c40\u9650\u6027\u7684\u5236\u7ea6\u3002\u6570\u636e\u96c6\u5728\u9884\u671f\u7684 288 \u4e2a\u89c2\u6d4b\u503c\u4e2d\u4ec5\u5305\u542b 90 \u4e2a\u6709\u6548\u89c2\u6d4b\u503c\uff0c\u4ee3\u8868 31.2% \u7684\u8986\u76d6\u7387\uff0c\u5bfc\u81f4\u8be5\u65f6\u6bb5\u7684\u5927\u90e8\u5206\u65f6\u95f4\u5904\u4e8e\u672a\u76d1\u63a7\u72b6\u6001\u3002","total_duration_ns":918251329,"translated_at":"2026-08-16T02:20:54.941934+00:00","translation_eval_count":147,"translation_model":"glm-5.2","translation_prompt_eval_count":247,"translation_total_duration_ns":776915802,"translation_wall_seconds":0.941,"valid_point_count":90,"wall_seconds":1.077}]}
