train
GitHub用于启动基于PPO算法的机器人环境训练任务。支持Cartpole、H1等环境,解析用户请求生成训练命令,处理超参数配置,并在后台执行训练,提供日志路径及进度查询指引。
Trigger Scenarios
用户希望开始机器人模型训练
用户指定特定环境(如cartpole)进行强化学习实验
用户需要调整PPO超参数并运行训练
Install
npx skills add rohanpsingh/LearningHumanoidWalking --skill train -g -y
SKILL.md
Frontmatter
{
"name": "train",
"description": "Launch a training run for a robot environment using PPO",
"allowed-tools": "Bash, Read, Glob",
"argument-hint": [
"env-name"
],
"disable-model-invocation": true
}
/train — Launch a PPO Training Run
Parse the user's request from $ARGUMENTS and construct a training command.
Command Template
RAY_ADDRESS= uv run python run_experiment.py train --env <ENV> --logdir <LOGDIR> [OPTIONS...]
Available Environments
| Name | Description |
|---|---|
cartpole |
Cartpole swing-up (simplest, good for testing) |
h1 |
Unitree H1 standing task |
jvrc_walk |
JVRC humanoid basic walking |
jvrc_step |
JVRC humanoid stepping with planned footsteps |
Hyperparameters (defaults)
| Flag | Default | Description |
|---|---|---|
--n-itr |
20000 | Training iterations |
--lr |
1e-4 | Learning rate |
--gamma |
0.99 | Discount factor |
--std-dev |
0.223 | Action noise |
--learn-std |
off | Learn action noise (flag) |
--entropy-coeff |
0.0 | Entropy regularization |
--clip |
0.2 | PPO clipping |
--minibatch-size |
64 | Minibatch size |
--epochs |
3 | Optimization epochs per update |
--num-procs |
12 | Parallel workers |
--num-envs-per-worker |
1 | Vectorized envs per worker |
--max-grad-norm |
0.05 | Gradient clipping |
--max-traj-len |
400 | Episode horizon |
--eval-freq |
100 | Eval every N iterations |
--seed |
None | Random seed |
--device |
auto | Training device (auto/cpu/cuda) |
--no-mirror |
off | Disable symmetry wrapper (flag) |
--recurrent |
off | Use LSTM policy (flag) |
--continued |
None | Path to pretrained weights |
Instructions
- Determine the environment name from the user's request. If ambiguous, ask.
- Use
--logdir /tmp/training_runsunless the user specifies a different path. - Only include flags that differ from defaults — keep the command clean.
- Show the user the full command you're about to run.
- Run the command in the background using
run_in_background: trueon the Bash tool. Set a generous timeout (600000ms). - After launching, tell the user the logdir path and how to check progress (you can tail the output using the task ID).
- If the user asks to check on training, use
TaskOutputwithblock: falseto check the latest output.
Cartpole-Specific Defaults
For cartpole, these settings are known to work well with the current defaults
(--lr 3e-4 --max-grad-norm 0.5 --lam 0.95 --gamma 0.99):
--minibatch-size 256--std-dev 0.15 --learn-std --entropy-coeff 0.01--max-traj-len 500 --n-itr 500 --num-procs 12--no-mirror(cartpole has no body symmetry)
Suggest these defaults when the user trains cartpole, but let them override.
Version History
- cd8c655 Current 2026-07-24 11:47


