AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
SKILL verified MIT Self-run

Pinchbench

skill-pinchbench-skill-skill · by pinchbench

Run PinchBench benchmarks to evaluate OpenClaw agent performance across real-world tasks. Use when testing model capabilities, comparing models, submitting benchmark results to the leaderboard, or checking how well your OpenClaw setup handles calendar, email, research, coding, and multi-step workflows.

No reviews yet
0 installs
0 views
view→install

Install

$ agentstack add skill-pinchbench-skill-skill

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access No
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets No
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/skill-pinchbench-skill-skill)

Reliability & compatibility

Security review passed
0 installs to date
no reviews yet
2mo ago

Declared compatibility

Claude CodeClaude Desktop

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Pinchbench? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

PinchBench Benchmark Skill

PinchBench measures how well LLM models perform as the brain of an OpenClaw agent. Results are collected on a public leaderboard at pinchbench.com.

Prerequisites

  • Python 3.10+
  • uv package manager
  • OpenClaw instance (this agent)

Quick Start

cd 

# Run benchmark with a specific model
uv run benchmark.py --model anthropic/claude-sonnet-4

# Run only automated tasks (faster)
uv run benchmark.py --model anthropic/claude-sonnet-4 --suite automated-only

# Run specific tasks
uv run benchmark.py --model anthropic/claude-sonnet-4 --suite task_calendar,task_stock

# Skip uploading results
uv run benchmark.py --model anthropic/claude-sonnet-4 --no-upload

Available Tasks (23)

| Task | Category | Description | |------|----------|-------------| | task_sanity | Basic | Verify agent works | | task_calendar | Productivity | Calendar event creation | | task_stock | Research | Stock price lookup | | task_blog | Writing | Blog post creation | | task_weather | Coding | Weather script | | task_summary | Analysis | Document summarization | | task_events | Research | Conference research | | task_email | Writing | Email drafting | | task_memory | Memory | Context retrieval | | task_files | Files | File structure creation | | task_workflow | Integration | Multi-step API workflow | | task_clawdhub | Skills | ClawHub interaction | | task_skill_search | Skills | Skill discovery | | task_image_gen | Creative | Image generation | | task_humanizer | Writing | Text humanization | | task_daily_summary | Productivity | Daily digest | | task_email_triage | Email | Inbox triage | | task_email_search | Email | Email search | | task_market_research | Research | Market analysis | | task_spreadsheet_summary | Analysis | Spreadsheet analysis | | task_eli5_pdf_summary | Analysis | PDF simplification | | task_openclaw_comprehension | Knowledge | OpenClaw docs comprehension | | task_second_brain | Memory | Knowledge management |

Command Line Options

| Option | Description | |--------|-------------| | --model | Model identifier (e.g., anthropic/claude-sonnet-4) | | --suite | all, automated-only, or comma-separated task IDs | | --output-dir | Results directory (default: results/) | | --timeout-multiplier | Scale task timeouts for slower models | | --runs | Number of runs per task for averaging | | --no-upload | Skip uploading to leaderboard | | --register | Request new API token for submissions | | --upload FILE | Upload previous results JSON |

Token Registration

To submit results to the leaderboard:

# Register for an API token (one-time)
uv run benchmark.py --register

# Run benchmark (auto-uploads with token)
uv run benchmark.py --model anthropic/claude-sonnet-4

Results

Results are saved as JSON in the output directory:

# View task scores
jq '.tasks[] | {task_id, score: .grading.mean}' results/0001_anthropic-claude-sonnet-4.json

# Show failed tasks
jq '.tasks[] | select(.grading.mean < 0.5)' results/*.json

# Calculate overall score
jq '{average: ([.tasks[].grading.mean] | add / length)}' results/*.json

Adding Custom Tasks

Create a markdown file in tasks/ following TASK_TEMPLATE.md. Each task needs:

  • YAML frontmatter (id, name, category, grading_type, timeout)
  • Prompt section
  • Expected behavior
  • Grading criteria
  • Automated checks (Python grading function)

Leaderboard

View results at pinchbench.com. The leaderboard shows:

  • Model rankings by overall score
  • Per-task breakdowns
  • Historical performance trends

Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.