# Eval Harness Design

> Use when agent evals need graders, pass@k, regression gates, or reliability evidence.

- **Type:** Skill
- **Install:** `agentstack add skill-jukrap-ai-agent-playbook-eval-harness-design`
- **Verified:** Yes — security-reviewed for prompt injection and unsafe behavior
- **Seller:** [jukrap](https://agentstack.voostack.com/s/jukrap)
- **Installs:** 0
- **Category:** [Agent Skills](https://agentstack.voostack.com/c/agent-skills)
- **Latest version:** 0.1.0
- **License:** MIT
- **Upstream author:** [jukrap](https://github.com/jukrap)
- **Source:** https://github.com/jukrap/ai-agent-playbook/tree/main/skills/delivery/eval-harness-design
- **Website:** https://www.npmjs.com/package/ai-agent-playbook

## Install

```sh
agentstack add skill-jukrap-ai-agent-playbook-eval-harness-design
```

Requires the [AgentStack CLI](https://agentstack.voostack.com/docs/cli). Works with Claude Code, Cursor, and any MCP-compatible agent.

## About

# Eval Harness Design

Use this as the primary delivery skill for eval-driven changes to agent workflows, prompts, MCP tools, skills, or automation surfaces.

## Workflow

1. Define the target behavior, risk class, baseline, failure modes, and release decision before changing the harness.
2. Choose deterministic code, schema, or rule graders before model or human graders.
3. Separate capability evals from regression evals, then set pass@k, pass^k, cost, latency, and repeatability expectations.
4. Store eval definitions and run reports as runtime evidence until reviewed and promoted.

## Reference

Read `references/eval-artifact-contract.md` for eval definition, run report, evidence envelope, and storage boundaries.

Read `references/grader-and-metric-rubric.md` for grader choice, metric thresholds, pass@k usage, and anti-overfitting checks.

## Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

- **Author:** [jukrap](https://github.com/jukrap)
- **Source:** [jukrap/ai-agent-playbook](https://github.com/jukrap/ai-agent-playbook)
- **License:** MIT
- **Homepage:** https://www.npmjs.com/package/ai-agent-playbook

Install and usage instructions live in the source repository linked above.

## Pricing

- **Free** — Free

## Security capabilities

Automated source analysis of v0.1.0 — what this tool can access:

- **Network access:** no
- **Filesystem access:** no
- **Shell / process execution:** no
- **Environment & secrets:** no
- **Dynamic code execution:** no

*"Yes" means the capability is present in the source — more access means more to trust, not that it is unsafe.*


## Versions

- **0.1.0** — security scan: passed — Imported from the upstream source.

## Links

- Listing page: https://agentstack.voostack.com/l/skill-jukrap-ai-agent-playbook-eval-harness-design
- Seller: https://agentstack.voostack.com/s/jukrap
- Browse the marketplace: https://agentstack.voostack.com/browse

---
Listed on AgentStack — the marketplace for AI agent skills and MCP servers. Every listing is security-reviewed. Creators keep 70%.
