jetstream-operations

GitHub

用于 NATS JetStream 系统的运维操作,涵盖故障排查、性能调优及监控告警。适用于诊断消费者延迟、消息投递失败、集群健康等问题,提供基于 nats CLI 的诊断流程和参考文件指引。

jetstream-operations/SKILL.md lithqube/nats-jetstream-claude-skills

Trigger Scenarios

JetStream 系统故障排查 消费者延迟或消息丢失诊断 性能调优与吞吐量优化 监控指标与告警配置

Install

npx skills add lithqube/nats-jetstream-claude-skills --skill jetstream-operations -g -y
More Options

Non-standard path

npx skills add https://github.com/lithqube/nats-jetstream-claude-skills/tree/main/jetstream-operations -g -y

Use without installing

npx skills use lithqube/nats-jetstream-claude-skills@jetstream-operations

指定 Agent (Claude Code)

npx skills add lithqube/nats-jetstream-claude-skills --skill jetstream-operations -a claude-code -g -y

安装 repo 全部 skill

npx skills add lithqube/nats-jetstream-claude-skills --all -g -y

预览 repo 内 skill

npx skills add lithqube/nats-jetstream-claude-skills --list

SKILL.md

Frontmatter
{
    "name": "jetstream-operations",
    "description": "Use this skill whenever users are operating, troubleshooting, or monitoring a running NATS JetStream system — including diagnosing consumer lag, messages not delivering, stream-full errors, performance tuning, Prometheus metrics, Grafana dashboards, alerting rules, JetStream advisory subjects, nats CLI usage, cluster health and leader election issues, or client connection\/reconnection problems. Use this skill when something is broken or slow, or when the user wants to observe their system. Do NOT use for designing new streams\/consumers (use jetstream-architecture) or deploying\/configuring infrastructure (use jetstream-deployment)."
}

JetStream Operations

Troubleshoot, monitor, and tune NATS JetStream including consumer lag, delivery failures, performance optimization, and observability.

For designing new streams or consumers, defer to the jetstream-architecture skill. For deploying or configuring NATS infrastructure, defer to the jetstream-deployment skill. If the user is operating a Synadia Agent fabric, the same tooling applies: enumerate live agents with nats req '$SRV.PING.agents' '', inspect endpoints/metadata with $SRV.INFO.agents, and watch agent liveness on the heartbeat subjects agents.hb.*.*.* (an agent is offline after ~3× its advertised interval_s). Tapping agents.> shows live prompt/response traffic. For how the protocol and those subjects are defined, see the nats-agent-fabric skill.

Reference Files

Read these files when they're relevant — don't load all of them upfront:

  • operations/troubleshooting.md — step-by-step diagnosis for messages not delivering, consumer lag, stream-full errors, cluster split-brain, and client disconnections. Read whenever something is broken.
  • operations/performance.md — publish throughput tuning, fetch batch sizing, MaxAckPending, parallel workers, FileStorage vs MemoryStorage, OS/TCP tuning, built-in benchmark commands. Read for performance or throughput questions.
  • operations/monitoring.md — NATS HTTP endpoints, Prometheus metrics, nats-exporter setup, advisory subjects, Grafana dashboard recommendations, Prometheus alert rules. Read for observability and alerting questions.
  • operations/cli-reference.md — full nats CLI reference for streams, consumers, pub/sub, server commands, and diagnostic workflows. Read when the user needs specific CLI commands.

Workflow

Step 1: Identify symptoms — what is the user observing? (lag, missing messages, errors, slow throughput)

Step 2: Inspect current state — use nats CLI to examine streams, consumers, and server status.

Step 3: Diagnose root cause — match symptoms to known patterns (ack timeout, filter mismatch, resource limits).

Step 4: Apply fix — provide specific configuration changes or code fixes.

Step 5: Verify resolution — confirm the fix with CLI commands and metrics.

Step 6: Set up monitoring — recommend metrics, alerts, and advisory subscriptions to prevent recurrence.

Core Principles

  • Always start diagnosis with nats stream report and nats consumer report
  • Check num_ack_pending first when consumers appear stuck — it's the most common bottleneck
  • Monitor JetStream advisory subjects for real-time operational events, but subscribe to the specific ones you need rather than $JS.EVENT.ADVISORY.> — that wildcard is the whole-account firehose and includes an API audit advisory published on every JetStream API response. Advisories are fire-and-forget, so anything you must not miss (max-deliveries, terminated) belongs in a stream you consume durably, not a bare subscription
  • Use nats server report jetstream to check cluster-wide JetStream health
  • Set alerts on consumer pending count, not just publish rate
  • Prefer nats CLI over raw API calls for operational tasks
  • Always check the NATS server logs — they surface warnings before failures
  • When in doubt, compare stream sequence numbers with consumer ack floor

Version History

  • 5001f06 Current 2026-09-22 13:05

    修正关于死信队列的错误描述,明确 NATS 无死信队列且超限丢弃;更新 EOL 服务器版本至 2.14;优化对 Advisory 事件订阅模式的说明。

  • de601e5 2026-07-25 05:48

Same Skill Collection

jetstream-architecture/SKILL.md
jetstream-deployment/SKILL.md
nats-agent-fabric/SKILL.md
nuxt-nats/SKILL.md

Metadata

Files
0
Version
5001f06
Hash
6fc4f01e
Indexed
2026-07-25 05:48

ホーム - Wiki
Copyright © 2011-2026 iteam. Current version is 2.155.2. UTC+08:00, 2026-10-04 02:08
浙ICP备14020137号-1