Agent Skillslangwatch/langwatch › debug-with-langwatch

debug-with-langwatch

GitHub

使用 LangWatch CLI 进行生产环境 Agent 故障排查的结构化工作流。通过定位错误 Trace、检查 Span 详情及监控评估分数,快速诊断并根因分析 LLM 应用异常。

skills/_compiled/native/debug-with-langwatch/SKILL.md langwatch/langwatch

Trigger Scenarios

生产环境 Agent 报错或失败 需要排查 LLM 响应质量差或延迟高 Agent 行为异常需根因分析

Install

npx skills add langwatch/langwatch --skill debug-with-langwatch -g -y
More Options

Non-standard path

npx skills add https://github.com/langwatch/langwatch/tree/main/skills/_compiled/native/debug-with-langwatch -g -y

Use without installing

npx skills use langwatch/langwatch@debug-with-langwatch

指定 Agent (Claude Code)

npx skills add langwatch/langwatch --skill debug-with-langwatch -a claude-code -g -y

安装 repo 全部 skill

npx skills add langwatch/langwatch --all -g -y

预览 repo 内 skill

npx skills add langwatch/langwatch --list

SKILL.md

Frontmatter
{
    "name": "debug-with-langwatch",
    "license": "MIT",
    "metadata": {
        "category": "recipe"
    },
    "description": "Root-cause production errors and misbehaving agent runs with LangWatch. Finds errored traces, inspects spans, checks monitor and evaluator scores, then narrows to a root cause. Use when something is failing or misbehaving in production (errors, bad answers, latency spikes).",
    "compatibility": "Requires the `langwatch` CLI with a valid `LANGWATCH_API_KEY`. Works with any coding agent."
}

Debug Production Issues with LangWatch

A structured diagnostic workflow: errored traces → span inspection → monitor/evaluator scores → root cause. Work the steps in order; each narrows the search space for the next.

If traces themselves look broken (empty inputs/outputs, disconnected spans), switch to the debug-instrumentation recipe instead. That is an instrumentation problem, not an application problem.

Prerequisites

Step 0: Point the CLI at the Right Project

langwatch status

A fast sanity check that the API key, endpoint, and project are the ones you mean to debug. Fix auth first (see the setup-lw recipe): every later step reads from this project.

Step 1: Find the Errored Traces

langwatch trace search --errors-only --limit 25 -o json
langwatch trace search --errors-only -q "timeout" --start-date 2026-01-01 -o json
  • --errors-only is how you find failures. An error is recorded on the span, not in the trace's searchable text, so -q "error" finds nothing and reads like a clean project.
  • --start-date/--end-date bound the window (ISO strings or epoch ms; default is the last 24h).
  • -q does a text search over one phrase: the error message, a user id, a thread id. AND, OR and NOT are matched as words, not as operators.
  • The result is { "traces": [...], "pagination": { "totalHits": N } }. Pull fields out with --jq instead of reading the whole payload:
langwatch trace search --limit 50 -o json --jq ".traces[].traceId"
langwatch trace search -q "refund" -o json --jq ".traces | length"

Look for: traces with error statuses, empty or truncated outputs, outliers in latency or cost, and repeats of the same failure across users/threads (a pattern, not a one-off).

Step 2: Inspect the Failing Spans

langwatch trace get <traceId>            # human-readable digest
langwatch trace get <traceId> -o json    # full span hierarchy

Read the span tree top-down:

  • Which span failed? The error is usually in one span (an LLM call, a tool call), not the whole trace. Note its input: a bad input upstream often explains a failure downstream.
  • What did the model see? Check the prompt/messages on the failing LLM span. Missing context, truncated history, and stale retrieved documents are the usual suspects.
  • Retries and timeouts: repeated identical spans suggest retry loops; a long-running span before the failure suggests a timeout.

Step 3: Check Monitors and Evaluator Scores

Production quality signals live in monitors (online evaluation) and their evaluators:

langwatch monitor list -o json           # which monitors exist, are they enabled/firing?
langwatch monitor get <id> -o json       # one monitor's config and recent state
langwatch evaluator list -o json         # the evaluators the monitors run
  • A firing monitor names the failure mode (toxicity, hallucination, PII). Corroborate it against the spans from Step 2.
  • No monitor for the failure mode you found? That is a gap worth closing once the root cause is fixed (langwatch monitor create).

For a quantitative view of the blast radius:

langwatch analytics query -m trace-count -a sum --group-by metadata.model -o json

Step 4: Root Cause and Verify

  1. Form a hypothesis from the failing span's input + the monitor's failure mode: prompt change, model change, bad retrieval, code regression. git log on the agent's code and prompts tells you what changed when the failures started.
  2. Apply the fix (prompt, code, or configuration).
  3. Generate fresh traffic, then re-run Step 1: the errored traces should stop appearing.
  4. If the failure was a regression, add a scenario so it stays fixed. The scenarios skill covers this.

Discovery

The full command surface, with per-command usage hints, is one command away:

langwatch commands -o json     # machine-readable catalog of every command
langwatch help-tree            # compact annotated tree (fits in context)
langwatch <group> --help       # flags for one group

Pass --agent to any command for compact single-line JSON with colour and spinners off (the CLI sets it by itself under Claude Code, Cursor, Copilot CLI and Amazon Q).

Version History

  • 6f9d4a4 Current 2026-08-28 21:09

    移除了文档查阅和命令发现相关的 Prerequisites 步骤,精简了前置条件部分。

  • 12615f1 2026-08-20 10:01

Same Skill Collection

.claude/skills/browser-pair/SKILL.md
.claude/skills/browser-test/SKILL.md
.claude/skills/code-review/SKILL.md
.claude/skills/feature-map/SKILL.md
.claude/skills/haven-setup/SKILL.md
.claude/skills/langwatch-kanban/SKILL.md
plugins/langwatch/skills/langwatch/SKILL.md
services/langy-agent/skills/github/SKILL.md
skills/_compiled/native/agent-best-practices/SKILL.md
skills/_compiled/native/agent-performance/SKILL.md
skills/_compiled/native/connect-agent/SKILL.md
skills/_compiled/native/context-sweet-spot/SKILL.md
skills/_compiled/native/datasets/SKILL.md
skills/_compiled/native/debug-instrumentation/SKILL.md
skills/_compiled/native/drive-the-ui/SKILL.md
skills/_compiled/native/eval-triage/SKILL.md
skills/_compiled/native/evaluate-multimodal/SKILL.md
skills/_compiled/native/evaluations/SKILL.md
skills/_compiled/native/experiments/SKILL.md
skills/_compiled/native/generate-rag-dataset/SKILL.md
skills/_compiled/native/github/SKILL.md
skills/_compiled/native/level-up/SKILL.md
skills/_compiled/native/lwql-charts/SKILL.md
skills/_compiled/native/online-evaluations/SKILL.md
skills/_compiled/native/prompt-optimization/SKILL.md
skills/_compiled/native/prompts/SKILL.md
skills/_compiled/native/provider-cost-comparison/SKILL.md
skills/_compiled/native/scenarios/SKILL.md
skills/_compiled/native/setup-lw/SKILL.md
skills/_compiled/native/test-cli-usability/SKILL.md
skills/_compiled/native/test-compliance/SKILL.md
skills/_compiled/native/tracing/SKILL.md

Metadata

Files
0
Version
6f9d4a4
Hash
199b669a
Indexed
2026-08-20 10:01

trang chủ - Wiki
Copyright © 2011-2026 iteam. Current version is 2.155.2. UTC+08:00, 2026-09-02 06:27
浙ICP备14020137号-1 $bản đồ khách truy cập$