GLM-5.2 是智谱迄今能力最强的开源模型,面向长程智能体工程打造,并支持真正可用的 1M token 上下文窗口。它能在超长任务中完整保持项目状态,减少反复压缩和丢弃上下文的损耗——任务越长,越能记得住、推得动。
| 服务商 | 上下文 | 最大输出 | 价格区间 | 输入(M/Tokens) | 输出(M/Tokens) | 缓存(M/Tokens) | 延迟 | 吞吐 | 稳定性 |
|---|---|---|---|---|---|---|---|---|---|
qianfan-cn | — | — | — | ¥3.2 | ¥11.2 | ¥0.8 | 985 ms | 88tps | 样本不足 |
wanqing-cn | — | — | — | ¥4.4 | ¥15.4 | ¥1.1 | 2.7 s | 94.43tps | 样本不足 |
bailian-cn-beijing | 1M | 128K | — | ¥4.8 | ¥16.8 | ¥1.2 | 624 ms | 68tps | 样本不足 |
siliconflow-cn | — | — | — | ¥4.8 | ¥16.8 | ¥1.2 | 1.7 s | 82.33tps | 样本不足 |
topenrouter-cn | — | — | — | ¥4 | ¥14 | ¥2 | 978 ms | 62.67tps | 样本不足 |
fangzhou-cn | — | — | — | ¥7.6 | ¥26.6 | ¥2 | 1.9 s | 82.75tps | 样本不足 |
通过智能路由聚合多家服务商,保障调用稳定
curl https://api.tokplus.ai/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $TOKGEN_API_KEY" \
-d '{
"model": "z-ai/glm-5.2",
"messages": [
{"role": "user", "content": "你好,介绍一下你自己"}
],
"stream": true
}'