Agent Skills › sgl-project/sglang › kernel-organization

kernel-organization

GitHub

规范 SGLang 内核代码的组织结构,定义 API 放置、分组逻辑、注册方式及测试目录标准,确保内核模块的清晰性与可维护性。

.claude/skills/kernel-organization/SKILL.md sgl-project/sglang

Trigger Scenarios

添加新的内核 API 移动或拆分内核文件 审查内核代码结构

Install

npx skills add sgl-project/sglang --skill kernel-organization -g -y
More Options

Non-standard path

npx skills add https://github.com/sgl-project/sglang/tree/main/.claude/skills/kernel-organization -g -y

Use without installing

npx skills use sgl-project/sglang@kernel-organization

指定 Agent (Claude Code)

npx skills add sgl-project/sglang --skill kernel-organization -a claude-code -g -y

安装 repo 全部 skill

npx skills add sgl-project/sglang --all -g -y

预览 repo 内 skill

npx skills add sgl-project/sglang --list

SKILL.md

Frontmatter
{
    "name": "kernel-organization",
    "description": "Apply the SGLang kernels RFC when adding, moving, splitting, or reviewing kernel APIs, registry metadata, kernel tests, benchmarks, and model-specific implementations. Use with add-jit-kernel, add-sgl-kernel, and write-sglang-test for placement and migration checks."
}

Kernel organization

Read python/sglang/kernels/README.md and test/README.md before choosing a location. RFC #29630 establishes the namespace; subsequent migrations #32148 and #40922 clarify logical grouping and remove the old _jit_ filename prefix.

Choose the owner

  • Put callable SGLang kernel APIs under sglang.kernels.ops.<group>. Runtime and integration tests import from that namespace, including its submodules. The sgl_kernel wheel retains its own public API and packaging tests.
  • Group by computation, not model or GPU: GEMM and GEMV in gemm, expert routing in moe, attention index selection in attention, sampling in sampling, normalization in layernorm. A quantized GEMM is still a GEMM; quantization alone belongs in quantization.
  • Model-specific files/subpackages and tuning data are allowed inside a logical group. Do not add a model bundle or an implementation file directly under ops/. Propose a new logical group only for a distinct responsibility, and update ops._GROUPS, documentation, and tests together.
  • Classify a fused operator by its complete contract. Do not split a fused kernel into separate launches just to separate norm, RoPE, or quantization. Split unrelated public entry points that happen to share a source file.
  • Keep shared CUDA build/runtime infrastructure in kernels/jit; operator wrappers call it from their logical group. Keep process groups, communicator state, model dispatch, and buffer ownership in srt. K3-specific adapters in srt/layers/communication/ need not pretend to be generic interfaces.

Register the public entry point

Add lazy KernelSpec metadata in the owning group's __init__.py, or use the existing BaseFusedOp registration when the operation has interchangeable backends. Group imports must not import GPU implementations or compile kernels.

Use <group>.<name> for the op id and module:callable for the target. Preserve existing backend dispatch and describe actual device/architecture restrictions with CapabilityRequirement; JIT/AOT are not device types. Input shape/dtype checks remain part of the entry point's contract. Registering a torch custom op does not register it in the SGLang kernel inventory. Predicates, private JIT factories, reference helpers, and runtime classes are not separate kernel APIs.

Place and preserve tests

  • Kernel numerical tests: test/registered/kernels/ops/<group>/.
  • CI microbenchmarks: test/registered/kernels/benchmark/<group>/.
  • Inventory/selector tests: test/registered/unit/kernels/.
  • Runtime unit tests: test/registered/unit/<subsystem>/, following the tested runtime module. A CUDA allocation does not make a runtime test a kernel test.
  • Non-CI smoke scripts and benchmarks: test/manual/kernels/; do not leave standalone test/benchmark entry points in production operator modules.
  • Shared test helpers: sglang.test.kernels. Preserve vendored upstream trees and AOT wheel packaging boundaries instead of reorganizing them incidentally.

For a move/split, preserve assertions, parametrization, fixtures, platform skips, execution entry points, and CI stages/runners. Apportion existing time estimates across split files; do not duplicate the original budget for every output file or silently drop a registration. Follow write-sglang-test for CI registration.

Verify a migration

  1. Search all runtime, test, benchmark, documentation, patch-string, and lazy registry references before deleting the old path. Check relative imports and package re-exports as well as direct imports. Do not leave forwarding shims.
  2. Keep tuning files with the GEMM that loads them and preserve relative lookup behavior. Check wheel/package inclusion as well as source-tree execution.
  3. Preserve module singleton state: every runtime caller must import the same new communication adapter, not a second copy of its buffer registry.
  4. Separate relocation commits from semantic changes. Use mechanical-refactor-verify to reproduce moves, and compare test inventories and CI registration coverage before/after. Only delete an experimental path after checking its call sites, flags, source, and dedicated tests together.
  5. Run the namespace/dispatch CPU tests, the registered-test validation hook, pre-commit, and relevant GPU tests when available. State exactly which GPU checks ran; import/AST checks do not prove numerical or performance parity.

Do not infer violations solely from a model name inside a group, a missing _jit_ prefix, use of a runtime utility, or absence of BaseFusedOp inheritance.

Version History

  • 81f27fb Current 2026-09-28 08:36

Same Skill Collection

.claude/skills/add-jit-kernel/SKILL.md
.claude/skills/add-sgl-kernel/SKILL.md
.claude/skills/babysit-pr-to-pass-ci/SKILL.md
.claude/skills/ci-test-audit/SKILL.md
.claude/skills/ci-workflow-guide/SKILL.md
.claude/skills/clean-startup-log/SKILL.md
.claude/skills/compute-mamba-ratio/SKILL.md
.claude/skills/cookbook-add-model/SKILL.md
.claude/skills/cookbook-migrate-model/SKILL.md
.claude/skills/cookbook-review-pr/SKILL.md
.claude/skills/debug-cuda-crash/SKILL.md
.claude/skills/debug-distributed-hang/SKILL.md
.claude/skills/env-var-conventions/SKILL.md
.claude/skills/generate-profile/SKILL.md
.claude/skills/kl-consistency-test/SKILL.md
.claude/skills/large-class-style/SKILL.md
.claude/skills/llm-torch-profiler-analysis/SKILL.md
.claude/skills/mechanical-refactor-verify/SKILL.md
.claude/skills/scripted-runtime-notes/SKILL.md
.claude/skills/sglang-bisect-ci-regression/SKILL.md
.claude/skills/sglang-cherrypick/SKILL.md
.claude/skills/sglang-prod-incident-triage/SKILL.md
.claude/skills/sglang-runtime-context/SKILL.md
.claude/skills/speculative-naming/SKILL.md
.claude/skills/write-sglang-test/SKILL.md

Metadata

Files
0
Version
81f27fb
Hash
04af1fc4
Indexed
2026-09-28 08:36

ホーム - Wiki
Copyright © 2011-2026 iteam. Current version is 2.155.2. UTC+08:00, 2026-10-04 19:56
浙ICP备14020137号-1