# Pinchbench

> Run PinchBench benchmarks to evaluate OpenClaw agent performance across real-world tasks. Use when testing model capabilities, comparing models, submitting benchmark results to the leaderboard, or checking how well your OpenClaw setup handles calendar, email, research, coding, and multi-step workflows.

- **Type:** Skill
- **Install:** `agentstack add skill-pinchbench-skill-skill`
- **Verified:** Yes — security-reviewed for prompt injection and unsafe behavior
- **Seller:** [pinchbench](https://agentstack.voostack.com/s/pinchbench)
- **Installs:** 0
- **Category:** [Productivity](https://agentstack.voostack.com/c/productivity)
- **Latest version:** 0.1.0
- **License:** MIT
- **Upstream author:** [pinchbench](https://github.com/pinchbench)
- **Source:** https://github.com/pinchbench/skill
- **Website:** https://pinchbench.com

## Install

```sh
agentstack add skill-pinchbench-skill-skill
```

Requires the [AgentStack CLI](https://agentstack.voostack.com/docs/cli). Works with Claude Code, Cursor, and any MCP-compatible agent.

## About

# PinchBench Benchmark Skill

PinchBench measures how well LLM models perform as the brain of an OpenClaw agent. Results are collected on a public leaderboard at [pinchbench.com](https://pinchbench.com).

## Prerequisites

- Python 3.10+
- [uv](https://docs.astral.sh/uv/) package manager
- OpenClaw instance (this agent)

## Quick Start

```bash
cd 

# Run benchmark with a specific model
uv run benchmark.py --model anthropic/claude-sonnet-4

# Run only automated tasks (faster)
uv run benchmark.py --model anthropic/claude-sonnet-4 --suite automated-only

# Run specific tasks
uv run benchmark.py --model anthropic/claude-sonnet-4 --suite task_calendar,task_stock

# Skip uploading results
uv run benchmark.py --model anthropic/claude-sonnet-4 --no-upload
```

## Available Tasks (23)

| Task | Category | Description |
|------|----------|-------------|
| `task_sanity` | Basic | Verify agent works |
| `task_calendar` | Productivity | Calendar event creation |
| `task_stock` | Research | Stock price lookup |
| `task_blog` | Writing | Blog post creation |
| `task_weather` | Coding | Weather script |
| `task_summary` | Analysis | Document summarization |
| `task_events` | Research | Conference research |
| `task_email` | Writing | Email drafting |
| `task_memory` | Memory | Context retrieval |
| `task_files` | Files | File structure creation |
| `task_workflow` | Integration | Multi-step API workflow |
| `task_clawdhub` | Skills | ClawHub interaction |
| `task_skill_search` | Skills | Skill discovery |
| `task_image_gen` | Creative | Image generation |
| `task_humanizer` | Writing | Text humanization |
| `task_daily_summary` | Productivity | Daily digest |
| `task_email_triage` | Email | Inbox triage |
| `task_email_search` | Email | Email search |
| `task_market_research` | Research | Market analysis |
| `task_spreadsheet_summary` | Analysis | Spreadsheet analysis |
| `task_eli5_pdf_summary` | Analysis | PDF simplification |
| `task_openclaw_comprehension` | Knowledge | OpenClaw docs comprehension |
| `task_second_brain` | Memory | Knowledge management |

## Command Line Options

| Option | Description |
|--------|-------------|
| `--model` | Model identifier (e.g., `anthropic/claude-sonnet-4`) |
| `--suite` | `all`, `automated-only`, or comma-separated task IDs |
| `--output-dir` | Results directory (default: `results/`) |
| `--timeout-multiplier` | Scale task timeouts for slower models |
| `--runs` | Number of runs per task for averaging |
| `--no-upload` | Skip uploading to leaderboard |
| `--register` | Request new API token for submissions |
| `--upload FILE` | Upload previous results JSON |

## Token Registration

To submit results to the leaderboard:

```bash
# Register for an API token (one-time)
uv run benchmark.py --register

# Run benchmark (auto-uploads with token)
uv run benchmark.py --model anthropic/claude-sonnet-4
```

## Results

Results are saved as JSON in the output directory:

```bash
# View task scores
jq '.tasks[] | {task_id, score: .grading.mean}' results/0001_anthropic-claude-sonnet-4.json

# Show failed tasks
jq '.tasks[] | select(.grading.mean < 0.5)' results/*.json

# Calculate overall score
jq '{average: ([.tasks[].grading.mean] | add / length)}' results/*.json
```

## Adding Custom Tasks

Create a markdown file in `tasks/` following `TASK_TEMPLATE.md`. Each task needs:

- YAML frontmatter (id, name, category, grading_type, timeout)
- Prompt section
- Expected behavior
- Grading criteria
- Automated checks (Python grading function)

## Leaderboard

View results at [pinchbench.com](https://pinchbench.com). The leaderboard shows:

- Model rankings by overall score
- Per-task breakdowns
- Historical performance trends

## Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

- **Author:** [pinchbench](https://github.com/pinchbench)
- **Source:** [pinchbench/skill](https://github.com/pinchbench/skill)
- **License:** MIT
- **Homepage:** https://pinchbench.com

Install and usage instructions live in the source repository linked above.

## Pricing

- **Free** — Free

## Security capabilities

Automated source analysis of v0.1.0 — what this tool can access:

- **Network access:** no
- **Filesystem access:** no
- **Shell / process execution:** no
- **Environment & secrets:** no
- **Dynamic code execution:** no

*"Yes" means the capability is present in the source — more access means more to trust, not that it is unsafe.*


## Versions

- **0.1.0** — security scan: passed — Imported from the upstream source.

## Links

- Listing page: https://agentstack.voostack.com/l/skill-pinchbench-skill-skill
- Seller: https://agentstack.voostack.com/s/pinchbench
- Browse the marketplace: https://agentstack.voostack.com/browse

---
Listed on AgentStack — the marketplace for AI agent skills and MCP servers. Every listing is security-reviewed. Creators keep 70%.
