Agent Skills
› NVIDIA/dgx-spark-playbooks
› dgx-station-mig
dgx-station-mig
GitHub用于检查、规划和应用 NVIDIA DGX Station GB300 的 MIG 配置。支持状态检查、布局规划及变更执行,强调安全约束与用户确认,防止误操作影响活跃客户端。
触发场景
用户请求启用或禁用 MIG 实例
需要重新配置 GPU 分区
查询 MIG UUID 或排查 MIG 相关故障
安装
npx skills add NVIDIA/dgx-spark-playbooks --skill dgx-station-mig -g -y
SKILL.md
Frontmatter
{
"name": "dgx-station-mig",
"description": "Inspect NVIDIA DGX Station GB300 MIG state and installed-driver profiles, create a digest-bound layout plan, disclose disruption and restoration, and apply an approved still-valid plan. Use when the user asks to enable, disable, partition, reconfigure, inspect, or troubleshoot MIG instances or needs MIG UUIDs. Never assume static profile IDs or terminate GPU clients."
}
DGX Station MIG
Treat inspection and planning as read-only. Treat every apply as disruptive.
Workflow
- Run
scripts/dgx-assist system inspect --json. - Search the playbooks for the user's MIG concern and cite relevant results.
- Read
compatibility.capabilities.mig_inspection. If it is false, stop and explain that the detected release has no MIG inspection profile. - Run
mig inspectandmig profiles. Use only profiles reported by the installed driver. This read-only step is available on the recognized Software 1.0 and Software 2.0 profiles. - Before planning, require
compatibility.capabilities.mig_mutationto betrue. If it is false, report the observed state but do not suggest a layout or work around the restriction. - Require the desired exact layout. Do not choose a layout from model names or nominal memory sums.
- Run
mig plan --layout "<layout>". - Present current and proposed state, exact commands, active clients, privilege, workload impact, reset/reboot possibility, and restoration commands.
- If any client is active, stop. Tell the user dgx-assist will not terminate it.
- Obtain explicit confirmation immediately before mutation.
- Run
mig apply --plan-id "<id>" --yes. - Report every verified MIG UUID. On partial failure, stop and present only the recorded restoration plan.
Safety requirements
- Never use a hardcoded GPU index, profile table, or assumed placement.
- Never destroy instances, reset a GPU, reboot, or change MIG mode outside an approved plan.
- Never kill a GPU process or take over a service.
- Invalidate the plan when clients, mode, profiles, release, driver, or instances change.
- Do not improvise further mutations after partial failure.
- Do not treat Fabric Manager as a normal Station prerequisite.
- Treat
--yesonly as approval already obtained.
Read references/workflow.md before planning or applying a layout.
版本历史
- 1fb66f0 当前 2026-08-20 12:20


