Agent Skillsintercom/2x-skills › cc-cost-analysis

cc-cost-analysis

GitHub

基于OpenTelemetry遥测数据分析Claude Code使用成本,涵盖模型支出、会话经济、Token类型及上下文膨胀等维度,提供查询模板与成本公式。

plugins/claude-code-tools/skills/cc-cost-analysis/SKILL.md intercom/2x-skills

Trigger Scenarios

analyze Claude Code costs break down the Claude Code bill find expensive sessions identify context bloat

Install

npx skills add intercom/2x-skills --skill cc-cost-analysis -g -y
More Options

Non-standard path

npx skills add https://github.com/intercom/2x-skills/tree/main/plugins/claude-code-tools/skills/cc-cost-analysis -g -y

Use without installing

npx skills use intercom/2x-skills@cc-cost-analysis

指定 Agent (Claude Code)

npx skills add intercom/2x-skills --skill cc-cost-analysis -a claude-code -g -y

安装 repo 全部 skill

npx skills add intercom/2x-skills --all -g -y

预览 repo 内 skill

npx skills add intercom/2x-skills --list

SKILL.md

Frontmatter
{
    "name": "cc-cost-analysis",
    "description": "Analyze Claude Code usage costs from OpenTelemetry data — per-user spend, expensive sessions, context bloat, and model\/token cost breakdowns — with a structured framework, ready-to-use query shapes, and cost-model formulas. Triggers on \"analyze Claude Code costs\", \"break down the Claude Code bill\", \"find expensive sessions\", \"identify context bloat\"."
}

Claude Code Cost Analysis

A framework and ready-to-use query shapes for analyzing Claude Code usage costs across a team or org.

Prerequisites: you need Claude Code telemetry

Claude Code can export usage as OpenTelemetry metrics and events (CLAUDE_CODE_ENABLE_TELEMETRY=1 plus an OTLP endpoint — see the Claude Code monitoring docs). Point that export at an observability backend (Honeycomb, Datadog, an OTLP collector into a warehouse, etc.). This skill analyzes that data.

The example queries in references/queries.md are written for Honeycomb (environment/dataset names are placeholders — substitute wherever you send telemetry). The analysis dimensions and cost formulas are backend-agnostic — the same shapes translate to any store that has per-api_request rows with token counts, cost_usd, model, and session.id.

A few attribute names in the examples (user.email, speed, skill.name, normalized_file_path) are enrichments that depend on your telemetry pipeline and Claude Code version — verify column names against your own data (get_dataset_columns or the equivalent) before relying on them. Numeric fields are sometimes stored as strings; cast with FLOAT(...) before aggregating (see the query examples).

Analysis Dimensions

Each dimension answers a different question. Run whichever are relevant — there is no required order, though broader dimensions provide context for narrower ones.

Dimension Question Answered Query Section
Model mix Which models account for what cost? references/queries.md Section 1
Session economics Where in sessions does cost accumulate? Section 2
Token types Cache writes vs reads vs output share? Section 3
Tool result sizes Which tools bloat the context? Section 4
Compaction Are sessions hitting context limits? Section 5
Per-user profiles Who are the heaviest users and why? Section 6
Instructions & skills How much do CLAUDE.md/rules/skills cost? Section 7
Command-level waste Which specific commands are wasteful vs bundled with real work? references/command-attribution.md

Apply cost formulas from references/cost-model.md when computing dollar estimates from token counts.

Illustrative Findings

These are shape-of-the-problem findings from one large deployment — useful to calibrate expectations, not ground truth for your org. Always verify against fresh queries against your own data.

Cost distribution

Dimension Typical finding
Model dominance The most capable model tends to dominate cost far out of proportion to its share of calls
Session length A minority of long sessions (turns 100+) can account for the majority of that model's cost
Token type shares Cache write ~47%, cache read ~44%, output ~9% of weighted cost
Caching savings Prompt caching typically cuts the bill ~80% vs fully uncached input
Instruction overhead CLAUDE.md, rules, and skills are collectively a small fraction (~1%) of total spend

Session turn phases

Phase Turns Driver
Cache creation 0-3 System prompt + instructions written to cache
Steady state 4-50 Caching working well — lowest per-call cost
Context growth 50-99 Cache reads climbing as context accumulates
Long-session tail 100+ Large accumulated context re-read every turn

Tool result sizes (typical ranges)

Category Avg/call Examples
Screenshots 100-500 KB browser/screenshot tools
Search results 10-30 KB search MCP tools
Database results 5-10 KB SQL / log query tools
File reads 3-7 KB Read tool (very high volume)

Context window behavior

  • A small fraction of sessions trigger auto compaction.
  • The most expensive sessions rarely compact — they grow to a large context and stay just below the window limit.
  • Sessions on smaller context windows compact more often and cost far less per call than sessions on the largest window with equivalent workloads.

Key Analysis Patterns

Matched session comparison

The most compelling way to demonstrate context-size impact on cost is finding two sessions with similar API-call counts but different context behaviors — one with high average cache_read (bloated, likely on the largest window) and one low (lean, compacting or on a smaller window). The per-session query (Section 6a adapted to break down by session.id) surfaces these pairs.

User segmentation by context usage

Classify users by what % of their calls exceed a large cache_read threshold (Section 6a, calls_over_200k field):

  • 0% over threshold — never needs the largest context, would work on a smaller window
  • 1-10% — rarely needs it, benefits most from proactive compaction
  • 60%+ — genuinely needs the largest context (investigate what tools drive it via Section 6b)

Important Caveats

Quality vs cost tradeoff: Cost optimizations (model downshift, aggressive compaction, smaller context windows) may degrade output quality. Measure quality objectively before changing configuration. Cost savings mean nothing if the tool becomes less useful.

Caching IS helping: Prompt caching reduces the bill substantially vs uncached input. The remaining cost is driven by the volume of cached tokens, not by caching being inefficient.

Instructions are a small share: CLAUDE.md files, rules, and skills are a small fraction of total spend. The real cost drivers are session length and context accumulation from tool results.

Turn-counts overcount command-level waste: a wasteful command (polling, re-auth, a bare sleep, a no-op) bundled into a turn that also does real work bills cache-read for that turn regardless — the marginal waste is close to zero. Classifying "turns containing pattern X" as waste massively overcounts. Separate standalone waste from bundled work by inspecting the actual command strings, and always back a dollar figure with atoms → examples → a reproducible query → an independent cross-check. See references/command-attribution.md.

Automated eval/test traffic pollutes usage totals: if you (or your telemetry pipeline) run automated eval or regression suites through Claude Code, that traffic is not real engineer usage. Tag or otherwise identify it in your export and exclude it from normal cost/context/compaction analysis; run a separate, intentionally-scoped query when you want to measure what the eval suite itself costs.

Generating Reports

Capture findings in markdown reports. See references/report-templates.md for standard structures. Name files {report-type}-{YYYY-MM-DD}.md.

Reference Files

  • references/queries.md — Example Honeycomb queries organized by analysis dimension. Load when composing a query for a specific analysis dimension.
  • references/cost-model.md — Token pricing ratios, per-call/per-session cost formulas, caching economics. Load when a finding needs a dollar estimate, not just a raw count.
  • references/report-templates.md — Standard report structures for each analysis type. Load before writing up findings.
  • references/column-gotchas.md — Data-quality gotchas specific to cost analysis. Load before trusting a cost or sequence column you haven't queried before.
  • references/command-attribution.md — Attributing cost to specific command patterns (waste vs bundled with real work). Load when classifying whether a costly command is standalone waste or bundled with real work.

Scripts

  • scripts/compute_percentiles.py — Compute p10/p25/p50/p75/p90/p99 from per-user query result sets

Version History

  • a1639a2 Current 2026-09-22 02:25

    新增命令级浪费分析维度,修正成本归因逻辑,明确缓存写入、自动评估流量排除及压缩成本估算规则。

  • 59213af 2026-07-19 09:00

Same Skill Collection

plugins/claude-code-tools/skills/audit-memory/SKILL.md
plugins/claude-code-tools/skills/permissions-analyzer/SKILL.md
plugins/claude-code-tools/skills/tool-misses/SKILL.md
plugins/code-review-tools/skills/thermo-nuclear-code-review/SKILL.md
plugins/pr-tools/skills/attach-github-assets/SKILL.md
plugins/pr-tools/skills/create-pr/SKILL.md
plugins/security-tools/skills/secure-github-actions/SKILL.md
plugins/skill-tools/skills/skill-review/SKILL.md
plugins/test-tools/skills/fix-flaky-tests/SKILL.md

Metadata

Files
0
Version
a1639a2
Hash
2ca3061e
Indexed
2026-07-19 09:00

Accueil - Wiki
Copyright © 2011-2026 iteam. Current version is 2.155.2. UTC+08:00, 2026-09-23 05:31
浙ICP备14020137号-1