Agent Skills
› Prismer-AI/Prismer
› ml-experiment
ml-experiment
GitHub用于设计和运行机器学习实验,涵盖模型训练、基准测试、消融研究及结果评估。支持复现论文、配置超参数搜索、统计分析及可视化,确保实验可复现性与严谨性。
Trigger Scenarios
用户要求训练模型或比较算法
需要运行消融研究以分析组件影响
要求复现特定论文的实验结果
评估机器学习模型性能指标
Install
npx skills add Prismer-AI/Prismer --skill ml-experiment -g -y
SKILL.md
Frontmatter
{
"name": "ml-experiment",
"description": "Design and run machine learning experiments with proper evaluation using jupyter_execute, including training, benchmarking, and ablation studies. Use when the user wants to train models, compare algorithms, run ablation studies, evaluate ML performance, or reproduce paper results."
}
ML Experiment Skill
Description
Design, implement, and evaluate machine learning experiments with reproducible workflows, proper baselines, and statistical analysis.
Tools Used
jupyter_execute- Execute ML code in Python (auto-switches to Jupyter)jupyter_notebook- Manage experiment notebooksupdate_notebook- Set up experiment cellsupdate_latex- Write experiment results to paperslatex_compile- Compile CS conference papers (auto-switches to LaTeX)arxiv_to_prompt- Read related work from arXiv papersupdate_notes- Write experiment logs and analysis summaries
Capabilities
Experiment Design
- Proper train/validation/test splits
- Cross-validation and bootstrap confidence intervals
- Ablation study design
- Hyperparameter search (grid, random, Bayesian)
Implementation
- PyTorch and TensorFlow model building
- Data loading and augmentation pipelines
- Training loops with logging and checkpointing
- Distributed training setup
Evaluation
- Standard metrics per task (accuracy, F1, BLEU, mAP, etc.)
- Statistical significance testing (paired t-test, bootstrap)
- Comparison with baselines
- Error analysis and visualization
Usage Patterns
Run an Experiment
When user says: "Train a model for [task]"
- Clarify dataset, metrics, and baselines
- Implement data loading and preprocessing
- Build model architecture
- Train with proper logging
- Evaluate and compare to baselines
- Report results with confidence intervals
Reproduce a Paper
When user says: "Reproduce [paper title/arXiv ID]"
- Fetch paper using arxiv_to_prompt
- Extract key method details
- Implement core algorithm
- Run experiments matching paper setup
- Compare results to reported numbers
Tool Examples
Train and evaluate a classifier
# via jupyter_execute
import torch
from sklearn.model_selection import train_test_split
from sklearn.metrics import classification_report
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
# ... train model ...
print(classification_report(y_test, predictions))
Run ablation study
# via jupyter_execute
configs = [
{"name": "full", "use_augmentation": True, "use_dropout": True},
{"name": "no_aug", "use_augmentation": False, "use_dropout": True},
{"name": "no_dropout", "use_augmentation": True, "use_dropout": False},
]
results = {c["name"]: train_and_eval(**c) for c in configs}
Validation checkpoints
- Verify data shapes match expected dimensions before training
- Check that loss is decreasing after the first few epochs
- Confirm test set has no overlap with training data
Version History
- 2dbe71f Current 2026-08-20 10:22


