# Run Iteration Eval

> Run the Logic-Lens content-eval pipeline for one iteration and produce a scored summary.json — use to measure a skill change. Wraps scripts/run-content-evals.sh (runner, costs tokens) and scripts/grade-iteration.py (grader, free, re-runnable). ALWAYS sync the plugin cache first. Use when the user wants to "run the evals", "score this iteration", "measure the skill change", "smoke-test before the…

- **Type:** Skill
- **Install:** `agentstack add skill-hyhmrright-logic-lens-run-iteration-eval`
- **Verified:** Yes — security-reviewed for prompt injection and unsafe behavior
- **Seller:** [hyhmrright](https://agentstack.voostack.com/s/hyhmrright)
- **Installs:** 0
- **Category:** [AI & ML](https://agentstack.voostack.com/c/ai-and-ml)
- **Latest version:** 0.1.0
- **License:** MIT
- **Upstream author:** [hyhmrright](https://github.com/hyhmrright)
- **Source:** https://github.com/hyhmrright/logic-lens/tree/main/.claude/skills/run-iteration-eval
- **Website:** https://github.com/hyhmrright/logic-lens

## Install

```sh
agentstack add skill-hyhmrright-logic-lens-run-iteration-eval
```

Requires the [AgentStack CLI](https://agentstack.voostack.com/docs/cli). Works with Claude Code, Cursor, and any MCP-compatible agent.

## About

# run-iteration-eval

Measures a skill change by running the content cases in `evals/content/v2/evals-v2.json` through
`claude -p` and grading the outputs. Outputs land in `skills-workspace/iteration-/`.

The runner and grader are split on purpose: **running** calls Claude and costs tokens; **grading**
is pure regex Python and is free to re-run on outputs that already exist. Never re-run the runner
just to re-score — re-grade instead.

## Steps

1. **Sync the cache first — non-negotiable.** The runner loads the skill from the plugin cache,
   not `skills/`. Run the `sync-skill-cache` skill (or its script directly). If you skip this, the
   eval grades the previously-published skill and the entire run is wasted:
   ```bash
   bash .claude/skills/sync-skill-cache/scripts/sync-cache.sh
   ```

2. **Pick a scope.** Full runs cost real tokens; scope down while iterating:
   ```bash
   SMOKE=1 bash scripts/run-content-evals.sh              # one case per mode (~$0.10) — fast sanity
   CASES="200 201 202" bash scripts/run-content-evals.sh  # only the cases a diagnosis flagged
   TAG=myfix bash scripts/run-content-evals.sh            # full run, named tag
   bash scripts/run-content-evals.sh                      # full run, tag = git short SHA
   ```
   The runner is idempotent — a case with an existing `output.md` is skipped. Delete the
   `eval-/` dir to force a re-run of that case.

3. **Read `summary.json`** in the iteration dir. It carries overall pass rate plus the per-mode and
   **per-subscore (logic vs format)** breakdown. The `logic` subscore reflects reasoning quality;
   `format` reflects Output-Skeleton compliance and is the historical bottleneck with high
   single-run variance. Judge a change on the right subscore — a format wobble is not a reasoning
   regression.

4. **Re-grade without re-running** (free) after editing the grader or to recompute on existing
   outputs:
   ```bash
   python3 scripts/grade-iteration.py skills-workspace/iteration-
   ```

5. **Single-iteration grade without the full suite** — when you already have outputs and only want
   the score table, `grade-iteration.py ` is the cheapest path (see `scripts/README.md`).

## Variance caveat

logic-review single-run scores are variance-dominated (see project memory). One run is a signal,
not a verdict — for a decision near the margin, run the affected cases 2–3× or widen the case set
before concluding a change helped or hurt. Hand the result to `iteration-guard` for the
ship/rollback call rather than eyeballing a single number.

## Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

- **Author:** [hyhmrright](https://github.com/hyhmrright)
- **Source:** [hyhmrright/logic-lens](https://github.com/hyhmrright/logic-lens)
- **License:** MIT
- **Homepage:** https://github.com/hyhmrright/logic-lens

Install and usage instructions live in the source repository linked above.

## Pricing

- **Free** — Free

## Security capabilities

Automated source analysis of v0.1.0 — what this tool can access:

- **Network access:** no
- **Filesystem access:** no
- **Shell / process execution:** no
- **Environment & secrets:** no
- **Dynamic code execution:** no

*"Yes" means the capability is present in the source — more access means more to trust, not that it is unsafe.*


## Versions

- **0.1.0** — security scan: passed — Imported from the upstream source.

## Links

- Listing page: https://agentstack.voostack.com/l/skill-hyhmrright-logic-lens-run-iteration-eval
- Seller: https://agentstack.voostack.com/s/hyhmrright
- Browse the marketplace: https://agentstack.voostack.com/browse

---
Listed on AgentStack — the marketplace for AI agent skills and MCP servers. Every listing is security-reviewed. Creators keep 70%.
