chat-glm
GitHub调用glm-5.2模型处理超长上下文或复杂推理任务,适用于代码库、长文档等场景。注意其成本高,仅在必要时使用,并直接输出结果。
Trigger Scenarios
Install
npx skills add zerogpu/zerogpu-router --skill chat-glm -g -y
SKILL.md
Frontmatter
{
"name": "chat-glm",
"description": "Chat with glm-5.2, a 753B MoE flagship with a 262K-token context window. Use when the input is too large for the other chat skills — an entire repository, a book-length document, a long agent transcript, though chat-deepseek holds four times as much — or for long-horizon reasoning the smaller models cannot hold together. This is the most expensive model on the platform, roughly 7x the cost of `chat`, so prefer chat for anything that fits in its 131K context.",
"allowed-tools": "Bash(zerogpu chat_completions *)",
"argument-hint": "<text>"
}
Call glm-5.2. $ARGUMENTS is the raw prompt — pass it verbatim, no escaping or quoting required (the heredoc below handles every shell metacharacter, newline, quote, and paren safely):
zerogpu chat_completions -m glm-5.2 <<'ZGPU_END_OF_INPUT'
$ARGUMENTS
ZGPU_END_OF_INPUT
Reach for this only when the size or horizon of the task actually needs it. At $1.10 / $3.50 per 1M input/output tokens, glm-5.2 costs about seven times /zerogpu-router:chat on input and six times on output (gpt-oss-120b, $0.15 / $0.60), and over fifty times the 1.2B edge models. For a prompt that fits in 131K tokens, /zerogpu-router:chat is the right call. For coding and agentic work, /zerogpu-router:chat-deepseek is far cheaper and holds four times the context.
Output is the assistant's answer as plain text — the model's reasoning trace comes back in a separate field and is not printed. Relay the answer as-is — do not rewrite or expand it.
Savings note: only if the command output literally contains a line starting with 💰 ZeroGPU savings, append that exact line, unchanged, as the last line of your reply. If no such line is present, say nothing about savings and do not mention or suggest /zerogpu-router:cost-savings — this note is intentionally occasional, not shown every time.
Version History
-
4e1b070
Current 2026-09-22 08:06
同步仪表盘API模型:上下文窗口从1048576调整为262144;更新CLI调用方式为chat_completions端点以独立于CLI版本;移除已下架的follow-up技能。
-
87b63e9
2026-08-27 16:47
修正了 glm-5.2 与 gpt-oss-120b 的成本对比数据,更新为实际价格的约7倍和6倍,此前错误引用导致路由决策偏差。
- cee9321 2026-08-02 21:14


