Agent Skillsdiegosouzapw/OmniRoute › omni-inference

omni-inference

GitHub

提供OpenAI兼容的AI推理API网关,支持聊天、嵌入、图像、音频及多模型路由。包含会话租赁管理、WebSocket实时通信及Ollama/Anthropic兼容接口,是AI代理的核心集成层。

skills/omni-inference/SKILL.md diegosouzapw/OmniRoute

Trigger Scenarios

需要调用大语言模型进行对话或生成内容 需要集成多种AI模型提供商(如Chat, Embeddings, TTS) 需要实现AI代理的后端推理服务

Install

npx skills add diegosouzapw/OmniRoute --skill omni-inference -g -y
More Options

Use without installing

npx skills use diegosouzapw/OmniRoute@omni-inference

指定 Agent (Claude Code)

npx skills add diegosouzapw/OmniRoute --skill omni-inference -a claude-code -g -y

安装 repo 全部 skill

npx skills add diegosouzapw/OmniRoute --all -g -y

预览 repo 内 skill

npx skills add diegosouzapw/OmniRoute --list

SKILL.md

Frontmatter
{
    "name": "omni-inference",
    "description": "The core OpenAI-compatible inference endpoints: chat completions, embeddings, images, audio (TTS\/STT), moderations, rerank, and the Responses API. The primary integration surface for AI agents."
}

Overview

The core OpenAI-compatible inference endpoints: chat completions, embeddings, images, audio (TTS/STT), moderations, rerank, and the Responses API. The primary integration surface for AI agents.

Authentication

All requests require a valid Bearer token or session cookie. Obtain a token via POST /api/auth/login or configure REQUIRE_API_KEY=false for local development.

Endpoints

POST /api/v1/session-leases

Acquire, renew, or release an exclusive managed connection lease

Requires an API key with lease:exclusive and an explicit non-empty allowedConnections policy. The opaque owner is bound to the authenticated API key; the lease owns an eligible connection, not a provider or model. Managed inference requests present the owner and exact generation headers. Temporary foreign occupancy returns 429 WAITING_FOR_CAPACITY with Retry-After.

curl -X POST https://localhost:20128/api/v1/session-leases \
  -H "Authorization: Bearer $OMNIROUTE_TOKEN"
  -H "Content-Type: application/json" \
  -d '{}'

POST /api/v1/chat/completions

Create chat completion

OpenAI-compatible chat completions endpoint. Routes to configured providers.

curl -X POST https://localhost:20128/api/v1/chat/completions \
  -H "Authorization: Bearer $OMNIROUTE_TOKEN"
  -H "Content-Type: application/json" \
  -d '{}'

GET /api/v1/ws

Chat completion over WebSocket (handshake + upgrade)

OpenAI-compatible chat over a WebSocket connection. GET with ?handshake=1 returns the connection descriptor (auth path, message protocol and live-event channels) as JSON; a plain GET without an Upgrade returns 426 Upgrade Required. After upgrading, the client exchanges JSON frames — {type:"request", id, payload:{model, messages}} to start a completion and {type:"cancel", id} to abort it. A separate live channel (default port LIVE_WS_PORT=20129, path /live) streams dashboard events on the requests, combo and credentials topics with a 15s heartbeat. Requires an API key.

curl https://localhost:20128/api/v1/ws \
  -H "Authorization: Bearer $OMNIROUTE_TOKEN"

POST /api/v1/providers/{provider}/chat/completions

Create chat completion (provider-specific)

Routes to a specific provider by name.

curl -X POST https://localhost:20128/api/v1/providers/{provider}/chat/completions \
  -H "Authorization: Bearer $OMNIROUTE_TOKEN"
  -H "Content-Type: application/json" \
  -d '{}'

POST /api/v1/api/chat

Ollama-compatible chat endpoint

Provides compatibility with Ollama's /api/chat format.

curl -X POST https://localhost:20128/api/v1/api/chat \
  -H "Authorization: Bearer $OMNIROUTE_TOKEN"
  -H "Content-Type: application/json" \
  -d '{}'

POST /api/v1/messages

Create message (Anthropic-compatible)

Anthropic Messages API endpoint. Routes to Claude providers.

curl -X POST https://localhost:20128/api/v1/messages \
  -H "Authorization: Bearer $OMNIROUTE_TOKEN"
  -H "Content-Type: application/json" \
  -d '{}'

POST /api/v1/messages/count_tokens

Count tokens for a message

curl -X POST https://localhost:20128/api/v1/messages/count_tokens \
  -H "Authorization: Bearer $OMNIROUTE_TOKEN"
  -H "Content-Type: application/json" \
  -d '{}'

POST /api/v1/responses

Create response (OpenAI Responses API)

OpenAI Responses API endpoint.

curl -X POST https://localhost:20128/api/v1/responses \
  -H "Authorization: Bearer $OMNIROUTE_TOKEN"
  -H "Content-Type: application/json" \
  -d '{}'

POST /api/v1/embeddings

Create embeddings

curl -X POST https://localhost:20128/api/v1/embeddings \
  -H "Authorization: Bearer $OMNIROUTE_TOKEN"
  -H "Content-Type: application/json" \
  -d '{}'

GET /api/v1/multimodal-embeddings

List embedding models (Jina multimodal-embeddings alias)

curl https://localhost:20128/api/v1/multimodal-embeddings \
  -H "Authorization: Bearer $OMNIROUTE_TOKEN"

POST /api/v1/multimodal-embeddings

Create embeddings (Jina multimodal-embeddings alias)

Same handler as POST /api/v1/embeddings. Provided so Jina-compatible clients that call /v1/multimodal-embeddings do not receive HTTP 404 unknown_route.

curl -X POST https://localhost:20128/api/v1/multimodal-embeddings \
  -H "Authorization: Bearer $OMNIROUTE_TOKEN"
  -H "Content-Type: application/json" \
  -d '{}'

POST /api/v1/providers/{provider}/embeddings

Create embeddings (provider-specific)

curl -X POST https://localhost:20128/api/v1/providers/{provider}/embeddings \
  -H "Authorization: Bearer $OMNIROUTE_TOKEN"
  -H "Content-Type: application/json" \
  -d '{}'

POST /api/v1/images/generations

Generate images

curl -X POST https://localhost:20128/api/v1/images/generations \
  -H "Authorization: Bearer $OMNIROUTE_TOKEN"
  -H "Content-Type: application/json" \
  -d '{}'

POST /api/v1/providers/{provider}/images/generations

Generate images (provider-specific)

curl -X POST https://localhost:20128/api/v1/providers/{provider}/images/generations \
  -H "Authorization: Bearer $OMNIROUTE_TOKEN"
  -H "Content-Type: application/json" \
  -d '{}'

POST /api/v1/audio/speech

Generate speech audio

Text-to-speech endpoint. Routes to configured TTS providers.

curl -X POST https://localhost:20128/api/v1/audio/speech \
  -H "Authorization: Bearer $OMNIROUTE_TOKEN"
  -H "Content-Type: application/json" \
  -d '{}'

POST /api/v1/audio/transcriptions

Transcribe audio

Audio-to-text transcription endpoint.

curl -X POST https://localhost:20128/api/v1/audio/transcriptions \
  -H "Authorization: Bearer $OMNIROUTE_TOKEN"
  -H "Content-Type: application/json" \
  -d '{}'

POST /api/v1/moderations

Create moderation

Content moderation endpoint. Routes to configured moderation providers.

curl -X POST https://localhost:20128/api/v1/moderations \
  -H "Authorization: Bearer $OMNIROUTE_TOKEN"
  -H "Content-Type: application/json" \
  -d '{}'

POST /api/v1/rerank

Rerank documents

Document reranking endpoint.

curl -X POST https://localhost:20128/api/v1/rerank \
  -H "Authorization: Bearer $OMNIROUTE_TOKEN"
  -H "Content-Type: application/json" \
  -d '{}'

GET /api/v1

API v1 root endpoint

Returns basic API info and status.

curl https://localhost:20128/api/v1 \
  -H "Authorization: Bearer $OMNIROUTE_TOKEN"

GET /api/v1/providers/{provider}/models

List models for a specific provider

Returns only models for the selected provider with provider prefix removed from each model id.

curl https://localhost:20128/api/v1/providers/{provider}/models \
  -H "Authorization: Bearer $OMNIROUTE_TOKEN"

GET /api/v1/management/proxy-subscriptions

List proxy subscriptions

Lists all operator-supplied proxy subscription links. Also starts the background auto-refresh scheduler (idempotent) so enabled subscriptions stay in sync. Credentials embedded in url are redacted in the response.

curl https://localhost:20128/api/v1/management/proxy-subscriptions \
  -H "Authorization: Bearer $OMNIROUTE_TOKEN"

POST /api/v1/management/proxy-subscriptions

Create a proxy subscription

Creates a subscription record. If mode is rule, at least one entry in ruleProviders is required. updateIntervalMinutes defaults to 60 and enabled defaults to false when omitted or not exactly true.

curl -X POST https://localhost:20128/api/v1/management/proxy-subscriptions \
  -H "Authorization: Bearer $OMNIROUTE_TOKEN"
  -H "Content-Type: application/json" \
  -d '{}'

GET /api/v1/management/proxy-subscriptions/{id}

Get a proxy subscription

curl https://localhost:20128/api/v1/management/proxy-subscriptions/{id} \
  -H "Authorization: Bearer $OMNIROUTE_TOKEN"

PATCH /api/v1/management/proxy-subscriptions/{id}

Update a proxy subscription

Partial update — only fields present in the body are changed (name/url/mode/ruleProviders/localCoreEndpoint/updateIntervalMinutes/enabled).

curl -X PATCH https://localhost:20128/api/v1/management/proxy-subscriptions/{id} \
  -H "Authorization: Bearer $OMNIROUTE_TOKEN"
  -H "Content-Type: application/json" \
  -d '{}'

DELETE /api/v1/management/proxy-subscriptions/{id}

Delete a proxy subscription

Removes the subscription record and unbinds/drops its synced proxy_registry rows.

curl -X DELETE https://localhost:20128/api/v1/management/proxy-subscriptions/{id} \
  -H "Authorization: Bearer $OMNIROUTE_TOKEN"

GET /api/v1/management/proxy-subscriptions/{id}/nodes

Get a subscription's last-parsed node summary

Returns the last-parsed node list without re-fetching the (possibly slow) subscription URL.

curl https://localhost:20128/api/v1/management/proxy-subscriptions/{id}/nodes \
  -H "Authorization: Bearer $OMNIROUTE_TOKEN"

POST /api/v1/management/proxy-subscriptions/{id}/refresh

Refresh a proxy subscription

Re-fetches and re-parses the subscription URL, syncs its nodes into proxy_registry, and (re)binds the pool.

curl -X POST https://localhost:20128/api/v1/management/proxy-subscriptions/{id}/refresh \
  -H "Authorization: Bearer $OMNIROUTE_TOKEN"
  -H "Content-Type: application/json" \
  -d '{}'

POST /api/v1/ocr

Document OCR

Multi-provider document OCR endpoint (Mistral OCR–compatible request and response shape). Accepts a JSON body referencing a document/image and returns extracted text. model selects the provider via a provider/model prefix (e.g. mistral/mistral-ocr-latest, azure-document-intelligence/prebuilt-read, vertex-deepseek-ocr/deepseek-ocr-maas); a bare model id (e.g. mistral-ocr-latest) resolves to its registered provider, and an omitted model defaults to Mistral. Azure Document Intelligence is asynchronous upstream — the handler polls the returned operation until it succeeds or fails before responding, so this endpoint can take longer to return for that provider. Success responses carry the X-OmniRoute-* cost-telemetry headers.

curl -X POST https://localhost:20128/api/v1/ocr \
  -H "Authorization: Bearer $OMNIROUTE_TOKEN"
  -H "Content-Type: application/json" \
  -d '{}'

POST /api/v1/audio/translations

Translate audio to English

OpenAI Whisper–compatible audio translation (multipart/form-data). Unlike /api/v1/audio/transcriptions, output is always English regardless of the source language. Success responses carry the X-OmniRoute-* cost-telemetry headers.

curl -X POST https://localhost:20128/api/v1/audio/translations \
  -H "Authorization: Bearer $OMNIROUTE_TOKEN"
  -H "Content-Type: application/json" \
  -d '{}'

GET /api/v1/providers/suggested-models

Suggested media models

Read-only server-side proxy to the public HuggingFace Hub models search API, used by the dashboard to suggest models for a media provider kind without exposing an HF token client-side. Never accepts or returns credentials.

curl https://localhost:20128/api/v1/providers/suggested-models \
  -H "Authorization: Bearer $OMNIROUTE_TOKEN"

GET /api/v1/provider-plugin-manifest

Provider plugin manifest

Returns the manifest describing installed provider plugins.

curl https://localhost:20128/api/v1/provider-plugin-manifest \
  -H "Authorization: Bearer $OMNIROUTE_TOKEN"

Payloads

See the full OpenAPI specification at GET /api/openapi/spec or docs/openapi.yaml for detailed request/response schemas.

Chat completions

Requires OMNIROUTE_URL and OMNIROUTE_KEY. See entry-point SKILL for setup.

Endpoints

  • POST $OMNIROUTE_URL/v1/chat/completions — OpenAI format
  • POST $OMNIROUTE_URL/v1/messages — Anthropic Messages format
  • POST $OMNIROUTE_URL/v1/responses — OpenAI Responses API

Discover

curl $OMNIROUTE_URL/v1/models | jq '.data[].id'

Combos (e.g. auto, cost-optimized, subscription) auto-fallback through multiple providers.

OpenAI format example

curl -X POST $OMNIROUTE_URL/v1/chat/completions \
  -H "Authorization: Bearer $OMNIROUTE_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "claude-opus-4-7",
    "messages": [{"role": "user", "content": "Refactor this function"}],
    "stream": true
  }'

Anthropic format example

curl -X POST $OMNIROUTE_URL/v1/messages \
  -H "Authorization: Bearer $OMNIROUTE_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "claude-opus-4-7",
    "max_tokens": 4096,
    "messages": [{"role": "user", "content": "Hi"}]
  }'

Tool use

Supports OpenAI tools array and Anthropic tools block. Tool results auto-compressed via RTK (47 filters: git-diff, grep, test-jest, terraform-plan, docker-logs, etc.) — 20-40% token savings. Disable per-request with X-Omniroute-Rtk: off header.

Reasoning / thinking

Anthropic extended thinking and OpenAI Responses reasoning blocks are forwarded verbatim. Cached automatically via reasoning cache.

Errors

  • 401 → invalid API key
  • 400 invalid_model → model not in registry; check /v1/models
  • 503 circuit_open → provider circuit breaker tripped; retry later or use combo
  • 429 rate_limited → honor Retry-After; consider using a combo for auto-fallback

Image generation

Requires OMNIROUTE_URL and OMNIROUTE_KEY. See entry-point SKILL for setup.

Endpoints

  • POST $OMNIROUTE_URL/v1/images/generations — Text-to-image
  • POST $OMNIROUTE_URL/v1/images/edits — Image edit (mask)
  • POST $OMNIROUTE_URL/v1/images/variations — Variations

Discover

curl $OMNIROUTE_URL/v1/models/image | jq '.data[]'

Returns { id, owned_by, sizes:[...], capabilities:[...] } per model.

Generate example

curl -X POST $OMNIROUTE_URL/v1/images/generations \
  -H "Authorization: Bearer $OMNIROUTE_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "dall-e-3",
    "prompt": "a red bicycle on a wet street, photoreal",
    "n": 1,
    "size": "1024x1024",
    "response_format": "b64_json"
  }'

Response: { created, data: [{ url? or b64_json, revised_prompt }] }

Errors

  • 400 invalid_size → not supported by this model; check /v1/models/image
  • 400 content_policy_violation → blocked by provider safety
  • 503 → provider unavailable; try another model in /v1/models/image

Text-to-speech

Requires OMNIROUTE_URL and OMNIROUTE_KEY. See entry-point SKILL for setup.

Endpoint

  • POST $OMNIROUTE_URL/v1/audio/speech — returns binary audio (mp3/opus/wav/flac)

Discover

curl $OMNIROUTE_URL/v1/models/tts | jq '.data[]'

Each entry includes voices:[...] for the available voice names per provider.

Example

curl -X POST $OMNIROUTE_URL/v1/audio/speech \
  -H "Authorization: Bearer $OMNIROUTE_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "tts-1",
    "input": "Hello from OmniRoute.",
    "voice": "alloy",
    "response_format": "mp3"
  }' --output speech.mp3

Voices

Voice names vary by provider. Check /v1/models/tts — each entry has voices:[...]. Common OpenAI voices: alloy, echo, fable, onyx, nova, shimmer.

Errors

  • 400 invalid_voice → voice not supported by this model
  • 400 input_too_long → input exceeds model character limit
  • 503 → provider unavailable; try another model in /v1/models/tts

Speech-to-text

Requires OMNIROUTE_URL and OMNIROUTE_KEY. See entry-point SKILL for setup.

Endpoints

  • POST $OMNIROUTE_URL/v1/audio/transcriptions — multipart upload, returns text
  • POST $OMNIROUTE_URL/v1/audio/translations — transcribe + translate to English

Discover

curl $OMNIROUTE_URL/v1/models/stt | jq '.data[]'

Example

curl -X POST $OMNIROUTE_URL/v1/audio/transcriptions \
  -H "Authorization: Bearer $OMNIROUTE_KEY" \
  -F "file=@audio.mp3" \
  -F "model=whisper-1" \
  -F "response_format=verbose_json"

Response: { text, language, duration, segments?:[{ start, end, text }] }

Supported formats

Audio: mp3, mp4, mpeg, mpga, m4a, wav, webm. Response formats: json, text, srt, verbose_json, vtt.

Errors

  • 400 invalid_file_format → unsupported audio format
  • 400 file_too_large → exceeds provider limit (usually 25MB)
  • 503 → provider unavailable; try another model in /v1/models/stt

Embeddings

Requires OMNIROUTE_URL and OMNIROUTE_KEY. See entry-point SKILL for setup.

Endpoint

  • POST $OMNIROUTE_URL/v1/embeddings

Discover

curl $OMNIROUTE_URL/v1/models/embedding | jq '.data[]'

Each entry: { id, owned_by, dimensions, max_input_tokens }.

Example

curl -X POST $OMNIROUTE_URL/v1/embeddings \
  -H "Authorization: Bearer $OMNIROUTE_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "text-embedding-3-large",
    "input": ["first text", "second text"],
    "encoding_format": "float"
  }'

Response: { data:[{ embedding:[...], index }], usage:{ prompt_tokens, total_tokens } }

Batch input

input accepts a string or array of strings (up to provider batch limit, typically 2048 items).

Errors

  • 400 input_too_long → input exceeds max_input_tokens for this model
  • 400 invalid_encoding_format → use float or base64
  • 503 → provider unavailable; try another model in /v1/models/embedding

Web search

Requires OMNIROUTE_URL and OMNIROUTE_KEY. See entry-point SKILL for setup.

Endpoint

  • POST $OMNIROUTE_URL/v1/web/search — unified search format

Discover

curl $OMNIROUTE_URL/v1/models/web | jq '.data[] | select(.kind == "webSearch")'

Example

curl -X POST $OMNIROUTE_URL/v1/web/search \
  -H "Authorization: Bearer $OMNIROUTE_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "tavily/search",
    "query": "OmniRoute github latest release",
    "max_results": 5,
    "include_answer": true
  }'

Response: { answer?, results:[{ url, title, content, score }] }

Parameters

Field Type Description
model string Provider model from /v1/models/web
query string Search query
max_results number Max results (default: 5)
include_answer boolean Include AI-synthesized answer
search_depth string basic or advanced (Tavily)

Errors

  • 400 query_too_long → shorten the search query
  • 503 → provider unavailable; try another model in /v1/models/web

Web fetch

Requires OMNIROUTE_URL and OMNIROUTE_KEY. See entry-point SKILL for setup.

Endpoint

  • POST $OMNIROUTE_URL/v1/web/fetch

Discover

curl $OMNIROUTE_URL/v1/models/web | jq '.data[] | select(.kind == "webFetch")'

Example

curl -X POST $OMNIROUTE_URL/v1/web/fetch \
  -H "Authorization: Bearer $OMNIROUTE_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "jina/reader",
    "url": "https://anthropic.com",
    "format": "markdown"
  }'

Response: { url, title, markdown, links?:[...], images?:[...] }

Parameters

Field Type Description
model string Provider from /v1/models/web (e.g. jina/reader, firecrawl/scrape)
url string URL to fetch
format string markdown (default), html, text

Errors

  • 400 invalid_url → URL must be http/https
  • 403 blocked → provider blocked by target site; try a different model
  • 503 → provider unavailable; try another model in /v1/models/web

Version History

  • 3d7ed7a Current 2026-08-20 06:19

    修复因CLI quota子命令更新未重新生成导致技能文件过期的问题;新增Vertex AI DeepSeek OCR提供商支持。

  • 1cafd32 2026-07-25 11:47

Same Skill Collection

skills/cli-a2a/SKILL.md
skills/cli-backup-sync/SKILL.md
skills/cli-batches/SKILL.md
skills/cli-chat/SKILL.md
skills/cli-compression/SKILL.md
skills/cli-contexts/SKILL.md
skills/cli-cost-usage/SKILL.md
skills/cli-eval/SKILL.md
skills/cli-health/SKILL.md
skills/cli-keys/SKILL.md
skills/cli-mcp/SKILL.md
skills/cli-models/SKILL.md
skills/cli-plugins-skills/SKILL.md
skills/cli-policy-audit/SKILL.md
skills/cli-providers/SKILL.md
skills/cli-resilience/SKILL.md
skills/cli-routing/SKILL.md
skills/cli-serve/SKILL.md
skills/cli-setup/SKILL.md
skills/cli-skill-collector/SKILL.md
skills/cli-tunnel/SKILL.md
skills/config-codex-cli/SKILL.md
skills/omni-agents-a2a/SKILL.md
skills/omni-api-keys/SKILL.md
skills/omni-auth/SKILL.md
skills/omni-budget/SKILL.md
skills/omni-cache/SKILL.md
skills/omni-cli-tools/SKILL.md
skills/omni-combos-routing/SKILL.md
skills/omni-compression/SKILL.md
skills/omni-context-rtk/SKILL.md
skills/omni-db-backups/SKILL.md
skills/omni-github-skills/SKILL.md
skills/omni-mcp/SKILL.md
skills/omni-models/SKILL.md
skills/omni-providers/SKILL.md
skills/omni-proxies/SKILL.md
skills/omni-resilience/SKILL.md
skills/omni-settings/SKILL.md
skills/omni-sync-cloud/SKILL.md
skills/omni-tunnels/SKILL.md
skills/omni-usage-logs/SKILL.md
skills/omni-version-manager/SKILL.md
skills/omni-webhooks/SKILL.md
skills/ponytail/SKILL.md

Metadata

Files
0
Version
3d7ed7a
Hash
3a2de6e4
Indexed
2026-07-25 11:47

inicio - Wiki
Copyright © 2011-2026 iteam. Current version is 2.155.2. UTC+08:00, 2026-08-22 03:17
浙ICP备14020137号-1 $mapa de visitantes$