# Eval

> >-

- **Type:** Skill
- **Install:** `agentstack add skill-mifunedev-agro-eval`
- **Verified:** Yes — security-reviewed for prompt injection and unsafe behavior
- **Seller:** [mifunedev](https://agentstack.voostack.com/s/mifunedev)
- **Installs:** 0
- **Category:** [Cloud & Infrastructure](https://agentstack.voostack.com/c/cloud-infrastructure)
- **Latest version:** 0.1.0
- **License:** Apache-2.0
- **Upstream author:** [mifunedev](https://github.com/mifunedev)
- **Source:** https://github.com/mifunedev/agro/tree/main/.agro/skills/eval
- **Website:** https://agro.mifune.dev

## Install

```sh
agentstack add skill-mifunedev-agro-eval
```

Requires the [AgentStack CLI](https://agentstack.voostack.com/docs/cli). Works with Claude Code, Cursor, and any MCP-compatible agent.

## About

# Eval

The runner for the harness **fitness function**. It discovers `.agro/evals/probes/*.sh`,
runs each against *real state*, and writes the `.agro/evals/RESULTS.md` scoreboard. A
rectification is provably "done" when its probe is green; a recurrence shows up as
a **REGRESSION** (was-PASS, now-fail) naming the `# source:` lesson. The full
contract — 3-state exit oracle, header convention, correction-surface triage — is
in [`.agro/evals/README.md`](../../../.agro/evals/README.md).

## Usage

```bash
bash .claude/skills/eval/run.sh                 # run the whole suite, rewrite RESULTS.md
bash .claude/skills/eval/run.sh --probe     # run one probe, update only its row
bash .claude/skills/eval/run.sh --tier A        # run only Tier-A probes
```

Exit-code oracle (per probe): `0`=PASS, `1`=REGRESSION, `2`=SKIPPED (not
applicable — excluded from pass-rate), `124`=TIMEOUT, other=ERROR. Each probe is
wrapped in `timeout 30s`. Runner aggregate exit (the process `$?` of `run.sh`
itself): `0` when no new green→red regression occurred this run, `1` when one or
more new regressions were detected (`${#regressions[@]} > 0`). When invoked via the
Bash tool as `bash .claude/skills/eval/run.sh`, the agent caller reads `$?` directly
to gate on success — the printed `REGRESSIONS (...)` stdout block and per-probe stderr
lines remain the human-readable signal. Note: the `eval-weekly` cron is an intentional
legacy caller that appends `|| true` then greps stdout; it does not consume the exit
code by design — this is not a bug.

## What the runner does

1. **Discover + run** every probe matching the filters; extract `# tier:` /
   `# source:` via the exact header grep.
2. **Compute the delta** vs the prior `RESULTS.md` row. **First run** (no prior
   row) emits `new-pass`/`new-fail` and raises NO regression without prior state.
3. **Surface regressions** — any `PASS → (REGRESSION|TIMEOUT|ERROR)` transition is
   printed first, naming the probe's `source`.
4. **Rewrite `RESULTS.md` atomically** — build the full scoreboard into a temp
   sibling file (`RESULTS.md.tmp.$$`) and replace the live file in one `mv -f`
   (never truncate-then-append in place), so a crash or concurrent run can't leave
   a partial scoreboard. Overwrite the row for each probe run; carry prior rows for
   probes not run this invocation from a **pre-write snapshot (`RESULTS_ORIG`)**
   captured before the rewrite — not the live file — so a filtered run never erases
   untouched rows and the scoreboard stays complete.

## When NOT to use

- **Tier-B behavioral evals** (sub-agent + LLM-judge of judgment-call behavior)
  are deferred — `/eval` is deterministic only. Never hard-gate on a noisy metric.
- For *scoring* context files for staleness/budget, that is `/audit context` and
  `/audit skills` — `/eval` checks behavior/state, not prose quality.

## Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

- **Author:** [mifunedev](https://github.com/mifunedev)
- **Source:** [mifunedev/agro](https://github.com/mifunedev/agro)
- **License:** Apache-2.0
- **Homepage:** https://agro.mifune.dev

Install and usage instructions live in the source repository linked above.

## Pricing

- **Free** — Free

## Security capabilities

Automated source analysis of v0.1.0 — what this tool can access:

- **Network access:** no
- **Filesystem access:** no
- **Shell / process execution:** no
- **Environment & secrets:** no
- **Dynamic code execution:** no

*"Yes" means the capability is present in the source — more access means more to trust, not that it is unsafe.*


## Versions

- **0.1.0** — security scan: passed — Imported from the upstream source.

## Links

- Listing page: https://agentstack.voostack.com/l/skill-mifunedev-agro-eval
- Seller: https://agentstack.voostack.com/s/mifunedev
- Browse the marketplace: https://agentstack.voostack.com/browse

---
Listed on AgentStack — the marketplace for AI agent skills and MCP servers. Every listing is security-reviewed. Creators keep 70%.
