Agent Skillspinchbench/skill › pinchbench

pinchbench

GitHub

用于运行 PinchBench 基准测试,评估 OpenClaw Agent 在日历、邮件、代码等真实任务中的性能。支持指定模型、筛选任务集及结果上传至排行榜,适用于模型能力对比与集成测试。

Trigger Scenarios

测试模型作为 Agent 大脑的能力 对比不同模型的性能表现 提交基准测试结果到排行榜 检查 OpenClaw 对多步工作流的处理能力

Install

npx skills add pinchbench/skill --skill pinchbench -g -y
More Options

Use without installing

npx skills use pinchbench/skill@pinchbench

指定 Agent (Claude Code)

npx skills add pinchbench/skill --skill pinchbench -a claude-code -g -y

安装 repo 全部 skill

npx skills add pinchbench/skill --all -g -y

预览 repo 内 skill

npx skills add pinchbench/skill --list

SKILL.md

Frontmatter
{
    "name": "pinchbench",
    "metadata": {
        "author": "pinchbench",
        "version": "2.0.0-rc1",
        "homepage": "https:\/\/pinchbench.com",
        "repository": "https:\/\/github.com\/pinchbench\/skill"
    },
    "description": "Run PinchBench benchmarks to evaluate OpenClaw agent performance across real-world tasks. Use when testing model capabilities, comparing models, submitting benchmark results to the leaderboard, or checking how well your OpenClaw setup handles calendar, email, research, coding, and multi-step workflows."
}

PinchBench Benchmark Skill

PinchBench measures how well LLM models perform as the brain of an OpenClaw agent. Results are collected on a public leaderboard at pinchbench.com.

Prerequisites

  • Python 3.10+
  • uv package manager
  • OpenClaw instance (this agent)

Quick Start

cd <skill_directory>

# Run benchmark with a specific model
uv run benchmark.py --model anthropic/claude-sonnet-4

# Run only automated tasks (faster)
uv run benchmark.py --model anthropic/claude-sonnet-4 --suite automated-only

# Run specific tasks
uv run benchmark.py --model anthropic/claude-sonnet-4 --suite task_calendar,task_stock

# Skip uploading results
uv run benchmark.py --model anthropic/claude-sonnet-4 --no-upload

Available Tasks (23)

Task Category Description
task_sanity Basic Verify agent works
task_calendar Productivity Calendar event creation
task_stock Research Stock price lookup
task_blog Writing Blog post creation
task_weather Coding Weather script
task_summary Analysis Document summarization
task_events Research Conference research
task_email Writing Email drafting
task_memory Memory Context retrieval
task_files Files File structure creation
task_workflow Integration Multi-step API workflow
task_clawdhub Skills ClawHub interaction
task_skill_search Skills Skill discovery
task_image_gen Creative Image generation
task_humanizer Writing Text humanization
task_daily_summary Productivity Daily digest
task_email_triage Email Inbox triage
task_email_search Email Email search
task_market_research Research Market analysis
task_spreadsheet_summary Analysis Spreadsheet analysis
task_eli5_pdf_summary Analysis PDF simplification
task_openclaw_comprehension Knowledge OpenClaw docs comprehension
task_second_brain Memory Knowledge management

Command Line Options

Option Description
--model Model identifier (e.g., anthropic/claude-sonnet-4)
--suite all, automated-only, or comma-separated task IDs
--output-dir Results directory (default: results/)
--timeout-multiplier Scale task timeouts for slower models
--runs Number of runs per task for averaging
--no-upload Skip uploading to leaderboard
--register Request new API token for submissions
--upload FILE Upload previous results JSON

Token Registration

To submit results to the leaderboard:

# Register for an API token (one-time)
uv run benchmark.py --register

# Run benchmark (auto-uploads with token)
uv run benchmark.py --model anthropic/claude-sonnet-4

Results

Results are saved as JSON in the output directory:

# View task scores
jq '.tasks[] | {task_id, score: .grading.mean}' results/0001_anthropic-claude-sonnet-4.json

# Show failed tasks
jq '.tasks[] | select(.grading.mean < 0.5)' results/*.json

# Calculate overall score
jq '{average: ([.tasks[].grading.mean] | add / length)}' results/*.json

Adding Custom Tasks

Create a markdown file in tasks/ following TASK_TEMPLATE.md. Each task needs:

  • YAML frontmatter (id, name, category, grading_type, timeout)
  • Prompt section
  • Expected behavior
  • Grading criteria
  • Automated checks (Python grading function)

Leaderboard

View results at pinchbench.com. The leaderboard shows:

  • Model rankings by overall score
  • Per-task breakdowns
  • Historical performance trends

Version History

  • 819384a Current 2026-07-25 08:38

Same Skill Collection

.agents/skills/building-dashboards/SKILL.md
.agents/skills/query-metrics/SKILL.md

Metadata

Files
0
Version
819384a
Hash
22e57c17
Indexed
2026-07-25 08:38

inicio - Wiki
Copyright © 2011-2026 iteam. Current version is 2.155.2. UTC+08:00, 2026-08-21 05:17
浙ICP备14020137号-1 $mapa de visitantes$