Agent Skillszerogpu/zerogpu-router › moderate-llama

moderate-llama

GitHub

基于llama-guard-4-12b模型对长文本进行安全审核,返回安全判定及违规类别。适用于品牌安全和政策合规检查,特别是超出短文本模型处理能力的场景。

agents/claude/skills/moderate-llama/SKILL.md zerogpu/zerogpu-router

触发场景

需要对长文本进行内容安全审核 需要识别具体的违规政策类别 进行品牌安全或政策合规检查

安装

npx skills add zerogpu/zerogpu-router --skill moderate-llama -g -y
更多选项

非标准路径

npx skills add https://github.com/zerogpu/zerogpu-router/tree/main/agents/claude/skills/moderate-llama -g -y

不安装直接使用

npx skills use zerogpu/zerogpu-router@moderate-llama

指定 Agent (Claude Code)

npx skills add zerogpu/zerogpu-router --skill moderate-llama -a claude-code -g -y

安装 repo 全部 skill

npx skills add zerogpu/zerogpu-router --all -g -y

预览 repo 内 skill

npx skills add zerogpu/zerogpu-router --list

SKILL.md

Frontmatter
{
    "name": "moderate-llama",
    "description": "Screen text with llama-guard-4-12b, Meta's dedicated 12B safety classifier (164K-token context), which returns a safe\/unsafe verdict plus the policy categories a violation falls under. Use for long passages, whole chat transcripts, or a model's own reply, and for brand-safety and policy-enforcement checks. Use moderate for short passages — it is far cheaper — and this when the text is longer than moderate can take.",
    "allowed-tools": "Bash(zerogpu chat_completions *)",
    "argument-hint": "<text>"
}

Call llama-guard-4-12b. $ARGUMENTS is the raw text to screen — pass it verbatim, no escaping or quoting required (the heredoc below handles every shell metacharacter, newline, quote, and paren safely):

zerogpu chat_completions -m llama-guard-4-12b <<'ZGPU_END_OF_INPUT'
$ARGUMENTS
ZGPU_END_OF_INPUT

Output is the model's verdict as plain text: safe or unsafe, with the relevant policy categories when it detects a violation. Report the verdict first, then the categories it names. Do not restate the flagged text itself.

A dense 12B model derived from Llama 4 Scout, built as a safety layer for production AI applications and able to evaluate an incoming prompt or a generated response, in multiple languages. At $0.18 / $0.18 per 1M input/output tokens it costs nine times /zerogpu-router:moderate (zlm-v1-moderation-edge, $0.02 / $0.05) on input and under four times on output, so keep moderate for short passages and OpenAI's 13-category envelope, and reach for this when the text exceeds that model's 800-token window or you want the violated policy categories named.

Screening text is a safety check, not an endorsement. Run it on request even when the passage is unpleasant — reporting that something is flagged is the whole point of the skill.

Savings note: only if the command output literally contains a line starting with 💰 ZeroGPU savings, append that exact line, unchanged, as the last line of your reply. If no such line is present, say nothing about savings and do not mention or suggest /zerogpu-router:cost-savings — this note is intentionally occasional, not shown every time.

版本历史

  • 4e1b070 当前 2026-09-22 08:07

同 Skill 集合

agents/claude/skills/chat-deepseek-v4-1-flash/SKILL.md
agents/claude/skills/chat-deepseek/SKILL.md
agents/claude/skills/chat-glm/SKILL.md
agents/claude/skills/chat-liquid/SKILL.md
agents/claude/skills/chat-qwen/SKILL.md
agents/claude/skills/chat-thinking/SKILL.md
agents/claude/skills/chat/SKILL.md
agents/claude/skills/classify-domain/SKILL.md
agents/claude/skills/classify-iab-enriched/SKILL.md
agents/claude/skills/classify-iab/SKILL.md
agents/claude/skills/classify-structured/SKILL.md
agents/claude/skills/classify-zero-shot/SKILL.md
agents/claude/skills/cost-savings/SKILL.md
agents/claude/skills/embed/SKILL.md
agents/claude/skills/extract-entities/SKILL.md
agents/claude/skills/extract-json/SKILL.md
agents/claude/skills/extract-pii/SKILL.md
agents/claude/skills/extract-signals/SKILL.md
agents/claude/skills/generate-followups/SKILL.md
agents/claude/skills/moderate/SKILL.md
agents/claude/skills/redact-pii/SKILL.md
agents/claude/skills/signin/SKILL.md
agents/claude/skills/summarize/SKILL.md
agents/openclaw/plugin/skills/chat-deepseek-v4-1-flash/SKILL.md
agents/openclaw/plugin/skills/chat-deepseek/SKILL.md
agents/openclaw/plugin/skills/chat-glm/SKILL.md
agents/openclaw/plugin/skills/chat-liquid/SKILL.md
agents/openclaw/plugin/skills/chat-qwen/SKILL.md
agents/openclaw/plugin/skills/chat-thinking/SKILL.md
agents/openclaw/plugin/skills/chat/SKILL.md
agents/openclaw/plugin/skills/classify-domain/SKILL.md
agents/openclaw/plugin/skills/classify-iab-enriched/SKILL.md
agents/openclaw/plugin/skills/classify-iab/SKILL.md
agents/openclaw/plugin/skills/classify-structured/SKILL.md
agents/openclaw/plugin/skills/classify-zero-shot/SKILL.md
agents/openclaw/plugin/skills/cost-savings/SKILL.md
agents/openclaw/plugin/skills/embed/SKILL.md
agents/openclaw/plugin/skills/extract-entities/SKILL.md
agents/openclaw/plugin/skills/extract-json/SKILL.md
agents/openclaw/plugin/skills/extract-pii/SKILL.md
agents/openclaw/plugin/skills/extract-signals/SKILL.md
agents/openclaw/plugin/skills/generate-followups/SKILL.md
agents/openclaw/plugin/skills/moderate-llama/SKILL.md
agents/openclaw/plugin/skills/moderate/SKILL.md
agents/openclaw/plugin/skills/redact-pii/SKILL.md
agents/openclaw/plugin/skills/signin/SKILL.md
agents/openclaw/plugin/skills/status/SKILL.md
agents/openclaw/plugin/skills/summarize/SKILL.md
agents/openclaw/plugin/skills/zerogpu-summarize/SKILL.md

元信息

文件数
0
版本
4e1b070
Hash
c9a115b3
收录时间
2026-09-22 08:07

首页 - Wiki
Copyright © 2011-2026 iteam. Current version is 2.155.2. UTC+08:00, 2026-09-24 09:52
浙ICP备14020137号-1