embed
GitHub将文本转换为384维嵌入向量,用于语义搜索、RAG检索、聚类或相似度比较。支持指定模型并返回标准化JSON输出。
Trigger Scenarios
Install
npx skills add zerogpu/zerogpu-router --skill embed -g -y
SKILL.md
Frontmatter
{
"name": "embed",
"description": "Turn text into a 384-dimensional embedding vector for semantic search, RAG retrieval, clustering, deduplication, or similarity comparison. Use when the user asks to embed text, build or query a vector index, or compare passages by meaning rather than by keyword.",
"allowed-tools": "Bash(zerogpu embeddings *)",
"argument-hint": "<text> [-m <model>]"
}
Embed the text in the request below. Run this with the Bash tool, pasting the text into the heredoc verbatim — no escaping, the quoted heredoc handles every shell metacharacter, newline, quote, and paren:
zerogpu embeddings -m all-minilm-l6-v2 <<'ZGPU_END_OF_INPUT'
<the text, verbatim>
ZGPU_END_OF_INPUT
Defaults to all-minilm-l6-v2 (22.7M parameters, 512-token window), the general-purpose choice for semantic similarity over short chunks. Use -m bge-small-en-v1.5 instead (33M parameters, 512-token window) when the request asks for it, for English retrieval, or when retrieval quality is the bottleneck. Both cost $0.004 per 1M input tokens and bill nothing on output, and both return 384-dimensional vectors, so they are interchangeable in an existing index.
Output is OpenAI's embeddings envelope as JSON: data[].embedding holds the vector, data[].index maps it back to its input, and usage reports input tokens. Do not print the raw vector back to the user — it is 384 floats and unreadable. Say what was embedded, which model produced it, and how many dimensions came back. Print or write the numbers only if the user explicitly asks for the vector or is piping it somewhere.
These models are served only by the Embeddings API, so this skill calls them there rather than through Chat Completions. Inputs longer than the model's window are truncated, so chunk long documents and embed the chunks.
Savings note: only if the command output literally contains a line starting with 💰 ZeroGPU savings, append that exact line, unchanged, as the last line of your reply. If no such line is present, say nothing about savings and do not mention or suggest /zerogpu-router:cost-savings — this note is intentionally occasional, not shown every time.
Request: $ARGUMENTS
Version History
-
4e1b070
Current 2026-09-22 08:07
版本从2.3.0升级至3.0.0。主要变更包括:嵌入服务价格由每百万输入令牌0.50美元降至0.004美元;默认模型all-minilm-l6-v2的上下文窗口从256扩展至512;bge-small-en-v1.5参数量微调;移除了generate-followups技能;更新CLI依赖以独立于模型变化。
- 87b63e9 2026-08-27 16:48


