system-profile

GitHub

用于对脚本、进程、GPU或内存等进行性能分析。支持自动选择工具或编写代码插桩,深入调查CPU、内存、通信及GPU计算瓶颈,并输出摘要报告。

skills/skills-codex/system-profile/SKILL.md wanshuiyin/Auto-claude-code-research-in-sleep

Trigger Scenarios

用户提到profile、benchmark、bottleneck等关键词 需要分析系统或应用的性能瓶颈

Install

npx skills add wanshuiyin/Auto-claude-code-research-in-sleep --skill system-profile -g -y
More Options

Non-standard path

npx skills add https://github.com/wanshuiyin/Auto-claude-code-research-in-sleep/tree/main/skills/skills-codex/system-profile -g -y

Use without installing

npx skills use wanshuiyin/Auto-claude-code-research-in-sleep@system-profile

指定 Agent (Claude Code)

npx skills add wanshuiyin/Auto-claude-code-research-in-sleep --skill system-profile -a claude-code -g -y

安装 repo 全部 skill

npx skills add wanshuiyin/Auto-claude-code-research-in-sleep --all -g -y

预览 repo 内 skill

npx skills add wanshuiyin/Auto-claude-code-research-in-sleep --list

SKILL.md

Frontmatter
{
    "name": "system-profile",
    "description": "Profile a target (script, process, GPU, memory, interconnect) for performance analysis. Use when user says \"profile\", \"benchmark\", \"bottleneck\", or wants performance analysis.",
    "argument-hint": "<target, e.g. \"train.py\", \"gpu\", \"pid 1234\", \"vllm serving\">"
}

System Profile

Profile the specified target and summarize the results. Target: $ARGUMENTS

Instructions

You are a profiling assistant. Based on the user's target, choose appropriate profiling strategies, including writing instrumentation code when needed, then run profiling, analyze results, and produce a summary.

Step 1: Determine the profiling target

Parse $ARGUMENTS to understand what to profile. Examples:

  • A Python script or module
  • A running process (PID or service name)
  • A specific function or code block
  • An entire framework or system (e.g., "autogen", "vllm serving") — profile its end-to-end execution, identify bottlenecks across components
  • "gpu" / "interconnect" / "memory" for focused profiling

If $ARGUMENTS is empty or unclear, ask the user.

Step 2: Choose profiling methods

Select from external tools and/or code instrumentation as appropriate. Don't limit yourself to the examples below — use whatever makes sense for the target.

External tools (check availability first):

  • CPU: cProfile, py-spy, line_profiler, perf stat, /usr/bin/time -v
  • Memory: tracemalloc, memory_profiler, memray
  • GPU: nvidia-smi, nvidia-smi dmon, nvitop, torch.profiler, nsys
  • Interconnect: nvidia-smi topo -m, nvidia-smi nvlink, NCCL_DEBUG=INFO
  • System: strace -c, iostat, vmstat

Code instrumentation — when external tools are insufficient, write and insert profiling code into the target. Typical scenarios:

  • Timing specific code blocks (wall time vs CPU time)
  • Measuring CPU-GPU or GPU-GPU transfer size, frequency, and bandwidth
  • Tracking memory allocation across CPU and GPU to detect redundancy
  • Wrapping NCCL collectives to measure latency and throughput
  • Adding CUDA event timing around kernels

Design the instrumentation based on what you observe in the code — don't use a fixed template.

Step 3: Key dimensions to investigate

Depending on the target, focus on some or all of these:

CPU overhead

  • Context switching (voluntary / involuntary)
  • CPU utilization: ratio of CPU time to wall time
  • Per-function execution time hotspots

Memory overhead

  • CPU and GPU memory usage (allocated vs reserved vs peak)
  • Redundant replication: same data living on both CPU and GPU
  • Per-device allocation balance in multi-GPU setups

Interconnect & communication

  • CPU-GPU transfer: frequency, per-transfer size, total volume, bandwidth achieved
  • GPU-GPU transfer: P2P bandwidth, NVLink vs PCIe topology impact
  • NCCL collectives: operation type, message size distribution, latency
  • Communication-to-computation ratio

GPU compute

  • SM utilization, kernel launch overhead
  • Memory bandwidth utilization vs peak

Step 4: Instrumentation guidelines

When inserting code into the target:

  1. Read and understand the target code first
  2. Prefer wrapping (decorator, context manager, standalone runner) over inline edits
  3. If inline edits are necessary, mark them clearly (e.g., # [PROFILE] comments)
  4. Minimize observer effect — don't instrument tight inner loops; sample instead
  5. Collect results into a structured log, don't scatter print statements

Step 5: Run profiling

  1. Check available tools and hardware topology
  2. Run the chosen methods, capture all output
  3. Save artifacts (flamegraphs, traces, logs) to ./profile_output/

Step 6: Produce the report

Part A — Profiling results (structured tables by dimension, as applicable):

  • CPU overhead table
  • Memory overhead table (with redundancy column)
  • Interconnect table (transfer type / frequency / size / latency / bandwidth)
  • Hotspots / bottleneck identification
  • Actionable recommendations ranked by expected impact

Part B — Instrumentation changelog (MANDATORY): List every file that was modified or created for profiling purposes:

File Change type What was added/modified Line(s)
... modified ... ...
... created ... —

This allows the user to review and revert all instrumentation changes. Offer to clean up (remove all instrumentation) when the user is done.

Version History

  • f4f20f9 Current 2026-08-20 05:07
  • 53562a7 2026-07-25 10:47

Same Skill Collection

skills/ablation-planner/SKILL.md
skills/alphaxiv/SKILL.md
skills/analyze-results/SKILL.md
skills/arxiv/SKILL.md
skills/auto-paper-improvement-loop/SKILL.md
skills/auto-review-loop-llm/SKILL.md
skills/auto-review-loop-minimax/SKILL.md
skills/auto-review-loop/SKILL.md
skills/citation-audit/SKILL.md
skills/claims-drafting/SKILL.md
skills/comm-lit-review/SKILL.md
skills/deepxiv/SKILL.md
skills/dse-loop/SKILL.md
skills/embodiment-description/SKILL.md
skills/exa-search/SKILL.md
skills/experiment-audit/SKILL.md
skills/experiment-bridge/SKILL.md
skills/experiment-plan/SKILL.md
skills/experiment-queue/SKILL.md
skills/feishu-notify/SKILL.md
skills/figure-description/SKILL.md
skills/figure-spec/SKILL.md
skills/formula-derivation/SKILL.md
skills/gemini-search/SKILL.md
skills/grant-proposal/SKILL.md
skills/idea-creator/SKILL.md
skills/idea-discovery-robot/SKILL.md
skills/idea-discovery/SKILL.md
skills/interview-cheatsheet/SKILL.md
skills/invention-structuring/SKILL.md
skills/jurisdiction-format/SKILL.md
skills/kill-argument/SKILL.md
skills/mermaid-diagram/SKILL.md
skills/meta-apply/SKILL.md
skills/meta-optimize/SKILL.md
skills/monitor-experiment/SKILL.md
skills/novelty-check/SKILL.md
skills/openalex/SKILL.md
skills/overleaf-sync/SKILL.md
skills/paper-claim-audit/SKILL.md
skills/paper-compile/SKILL.md
skills/paper-figure/SKILL.md
skills/paper-illustration-image2/SKILL.md
skills/paper-illustration/SKILL.md
skills/paper-plan/SKILL.md
skills/paper-poster-html/SKILL.md
skills/paper-poster/SKILL.md
skills/paper-slides/SKILL.md
skills/paper-talk/SKILL.md
skills/paper-write/SKILL.md

Metadata

Files
0
Version
341f914
Hash
9d31e32d
Indexed
2026-07-25 10:47

ホーム - Wiki
Copyright © 2011-2026 iteam. Current version is 2.155.2. UTC+08:00, 2026-10-04 12:35
浙ICP备14020137号-1