# Skill Eval

> Evaluate AI Agent Skills across safety, quality, reliability, and cost efficiency. Audit for security issues (secrets, injection, unsafe installs), test functional correctness with-skill vs without-skill, measure trigger precision, classify cost-efficiency tradeoffs, track version lifecycle, and generate unified grades. Use when evaluating a skill before installing, auditing marketplace skills, p…

- **Type:** Skill
- **Install:** `agentstack add skill-aws-samples-sample-agent-skill-eval-sample-agent-skill-eval`
- **Verified:** Yes — security-reviewed for prompt injection and unsafe behavior
- **Seller:** [aws-samples](https://agentstack.voostack.com/s/aws-samples)
- **Installs:** 0
- **Category:** [Developer Tools](https://agentstack.voostack.com/c/developer-tools)
- **Latest version:** 0.1.0
- **License:** MIT-0
- **Upstream author:** [aws-samples](https://github.com/aws-samples)
- **Source:** https://github.com/aws-samples/sample-agent-skill-eval

## Install

```sh
agentstack add skill-aws-samples-sample-agent-skill-eval-sample-agent-skill-eval
```

Requires the [AgentStack CLI](https://agentstack.voostack.com/docs/cli). Works with Claude Code, Cursor, and any MCP-compatible agent.

## About

# Skill Eval — Agent Skill Evaluation Framework

Evaluate Agent Skills across four dimensions: safety (audit), quality (functional), reliability (trigger), and cost efficiency (Pareto classification).

## Quick Start

```bash
skill-eval audit /path/to/skill          # Is it safe?
skill-eval report /path/to/skill         # Full grade (audit + functional + trigger)
skill-eval functional /path/to/skill     # Quality: with-skill vs without-skill
skill-eval trigger /path/to/skill        # Reliability: activation precision
```

## Decision Tree

- **"Is this skill safe?"** → `skill-eval audit `
- **"Full evaluation with grade"** → `skill-eval report `
- **"Full repo security review"** → `skill-eval audit  --include-all`
- **"Write eval cases"** → `skill-eval init `, then edit `evals/`
- **"Compare two versions"** → `skill-eval compare  `
- **"Check for regressions"** → `skill-eval snapshot `, then `skill-eval regression `
- **"Track changes"** → `skill-eval lifecycle  --save --label v1.0`

## Commands

| Command | Purpose |
|---------|---------|
| `audit` | Security & structure scan (secrets, permissions, spec compliance) |
| `functional` | Quality eval — runs prompts with and without skill, grades output |
| `trigger` | Reliability eval — tests activation precision for relevant/irrelevant queries |
| `report` | Unified grade combining audit (40%) + functional (40%) + trigger (20%) |
| `compare` | Side-by-side comparison of two skills on the same eval cases |
| `snapshot` | Save current audit as regression baseline |
| `regression` | Check for score regressions against baseline |
| `lifecycle` | Version tracking and change detection |
| `init` | Generate eval scaffold from SKILL.md frontmatter |

For detailed flags and examples, see `references/cli-reference.md`.

## Eval File Format

Functional evals (`evals/evals.json`):
```json
[{"id": "case-1", "prompt": "...", "assertions": ["contains 'expected'"], "files": ["files/input.csv"]}]
```

Trigger queries (`evals/eval_queries.json`):
```json
[{"query": "relevant question", "should_trigger": true}, {"query": "unrelated question", "should_trigger": false}]
```

## Scoring

Grades: A (90+), B (80-89), C (70-79), D (60-69), F (<60). Findings deduct: CRITICAL −25, WARNING −10, INFO −2.

For the full security check reference and OWASP mapping, see `references/security-checks.md`.

## Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

- **Author:** [aws-samples](https://github.com/aws-samples)
- **Source:** [aws-samples/sample-agent-skill-eval](https://github.com/aws-samples/sample-agent-skill-eval)
- **License:** MIT-0

Install and usage instructions live in the source repository linked above.

## Pricing

- **Free** — Free

## Security capabilities

Automated source analysis of v0.1.0 — what this tool can access:

- **Network access:** no
- **Filesystem access:** no
- **Shell / process execution:** no
- **Environment & secrets:** no
- **Dynamic code execution:** no

*"Yes" means the capability is present in the source — more access means more to trust, not that it is unsafe.*


## Versions

- **0.1.0** — security scan: passed — Imported from the upstream source.

## Links

- Listing page: https://agentstack.voostack.com/l/skill-aws-samples-sample-agent-skill-eval-sample-agent-skill-eval
- Seller: https://agentstack.voostack.com/s/aws-samples
- Browse the marketplace: https://agentstack.voostack.com/browse

---
Listed on AgentStack — the marketplace for AI agent skills and MCP servers. Every listing is security-reviewed. Creators keep 70%.
