Agent Skills
› aiming-lab/AutoResearchClaw
› nlp-alignment
nlp-alignment
GitHub提供LLM对齐最佳实践,涵盖RLHF、DPO等方法及训练配方,帮助处理模型安全与对齐任务。
Trigger Scenarios
需要优化大语言模型的对齐效果
实施RLHF或DPO等对齐技术
解决模型安全或指令遵循问题
Install
npx skills add aiming-lab/AutoResearchClaw --skill nlp-alignment -g -y
SKILL.md
Frontmatter
{
"name": "nlp-alignment",
"metadata": {
"author": "researchclaw",
"version": "1.0",
"category": "domain",
"priority": "4",
"references": "Ouyang et al., Training language models to follow instructions, NeurIPS 2022; Rafailov et al., DPO, NeurIPS 2023",
"trigger-keywords": "alignment,rlhf,dpo,reward model,preference,instruction tuning,safety",
"applicable-stages": "9,10"
},
"description": "Best practices for LLM alignment techniques including RLHF, DPO, and instruction tuning. Use when working on alignment or safety."
}
LLM Alignment Best Practice
Methods:
- RLHF: Train reward model → PPO fine-tuning (complex but powerful)
- DPO: Direct preference optimization (simpler, no reward model needed)
- GRPO: Group relative policy optimization
- SFT: Supervised fine-tuning as alignment baseline
Training recipe:
- Start with SFT on high-quality instruction data
- DPO: lr=5e-7, beta=0.1, batch_size=64
- PPO: lr=1e-6, clip=0.2, KL coeff=0.02
- Use reference model for KL penalty
- Evaluate on safety benchmarks (TruthfulQA, BBQ, etc.)
Common pitfalls:
- Reward hacking: model finds shortcuts to high reward
- Mode collapse: model generates repetitive outputs
- Catastrophic forgetting: loses general capabilities
Version History
- e2e23c9 Current 2026-07-25 07:48


