train

GitHub

用于启动机器人环境的PPO训练任务,解析参数生成命令并后台执行。支持Cartpole、H1等环境及自定义超参数,提供进度监控与日志路径反馈。

.claude/skills/train/SKILL.md rohanpsingh/LearningHumanoidWalking

Trigger Scenarios

启动机器人训练 运行PPO算法 配置训练超参数

Install

npx skills add rohanpsingh/LearningHumanoidWalking --skill train -g -y
More Options

Non-standard path

npx skills add https://github.com/rohanpsingh/LearningHumanoidWalking/tree/main/.claude/skills/train -g -y

Use without installing

npx skills use rohanpsingh/LearningHumanoidWalking@train

指定 Agent (Claude Code)

npx skills add rohanpsingh/LearningHumanoidWalking --skill train -a claude-code -g -y

安装 repo 全部 skill

npx skills add rohanpsingh/LearningHumanoidWalking --all -g -y

预览 repo 内 skill

npx skills add rohanpsingh/LearningHumanoidWalking --list

SKILL.md

Frontmatter
{
    "name": "train",
    "description": "Launch a training run for a robot environment using PPO",
    "allowed-tools": "Bash, Read, Glob",
    "argument-hint": [
        "env-name"
    ],
    "disable-model-invocation": true
}

/train — Launch a PPO Training Run

Parse the user's request from $ARGUMENTS and construct a training command.

Command Template

RAY_ADDRESS= uv run python run_experiment.py train --env <ENV> --logdir <LOGDIR> [OPTIONS...]

Available Environments

Name Description
cartpole Cartpole swing-up (simplest, good for testing)
h1 Unitree H1 standing task
jvrc_walk JVRC humanoid basic walking
jvrc_step JVRC humanoid stepping with planned footsteps

Hyperparameters (defaults)

Flag Default Description
--n-itr 20000 Training iterations
--lr 1e-4 Learning rate
--gamma 0.99 Discount factor
--std-dev 0.223 Action noise
--learn-std off Learn action noise (flag)
--entropy-coeff 0.0 Entropy regularization
--clip 0.2 PPO clipping
--minibatch-size 64 Minibatch size
--epochs 3 Optimization epochs per update
--num-procs 12 Parallel workers
--num-envs-per-worker 1 Vectorized envs per worker
--max-grad-norm 0.05 Gradient clipping
--max-traj-len 400 Episode horizon
--eval-freq 100 Eval every N iterations
--seed None Random seed
--device auto Training device (auto/cpu/cuda)
--no-mirror off Disable symmetry wrapper (flag)
--recurrent off Use LSTM policy (flag)
--continued None Path to pretrained weights

Instructions

  1. Determine the environment name from the user's request. If ambiguous, ask.
  2. Use --logdir /tmp/training_runs unless the user specifies a different path.
  3. Only include flags that differ from defaults — keep the command clean.
  4. Show the user the full command you're about to run.
  5. Run the command in the background using run_in_background: true on the Bash tool. Set a generous timeout (600000ms).
  6. After launching, tell the user the logdir path and how to check progress (you can tail the output using the task ID).
  7. If the user asks to check on training, use TaskOutput with block: false to check the latest output.

Cartpole-Specific Defaults

For cartpole, these settings are known to work well with the current defaults (--lr 3e-4 --max-grad-norm 0.5 --lam 0.95 --gamma 0.99):

  • --minibatch-size 256
  • --std-dev 0.15 --learn-std --entropy-coeff 0.01
  • --max-traj-len 500 --n-itr 500 --num-procs 12
  • --no-mirror (cartpole has no body symmetry)

Suggest these defaults when the user trains cartpole, but let them override.

Version History

  • cd8c655 Current 2026-07-24 11:47

Same Skill Collection

.claude/skills/experiment/SKILL.md
.claude/skills/eval/SKILL.md

Metadata

Files
0
Version
cd8c655
Hash
5795cdc5
Indexed
2026-07-24 11:47

inicio - Wiki
Copyright © 2011-2026 iteam. Current version is 2.155.2. UTC+08:00, 2026-08-17 17:47
浙ICP备14020137号-1 $mapa de visitantes$