Install
$ agentstack add skill-crewforth-crewforth-eval-grader ✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ✓ Network access No
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
Verified badge
Passed review? Show it. Paste this badge into your README, it links to the public security report.
Reliability & compatibility
Declared compatibility
Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.
We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.
How agent discovery & health will work →About
Eval Grader
Trigger phrases: "eval", "grader", "measure output quality", "LLM-as-judge", "score the output"
Measure every change; don't vibe it. When you iterate on a prompt, an agent, or any generative output (docs, slides, UI, a summary, an extraction), a two-layer grader over a fixed task set turns "feels better" into a signed number you can trust.
This is the external, machine-grounded verifier the iterate skill asks for — a model grading its own output inflates; a separate grader on a fixed suite does not.
> Crewforth adaptation (local, .claude/): use when tuning a generative task; the scorecard goes to docs/EVAL.md > (§4.3). Stack-agnostic — graders are ordinary code + judge calls. §4 Prohibitions apply.
Two layers
- Layer 1 — code graders (deterministic, near-free, run every time): structural metrics over the artifact —
did it produce a valid result? plus counts, sizes, schema validity, "wall-of-text" / clutter flags. They catch gross regressions a judge shouldn't be spent on. Ground truth is computed from the source, not hand-authored.
- Layer 2 — LLM-as-judge graders (semantic): one call per dimension (clarity · correctness-vs-source ·
completeness…), scored on an explicit rubric. Steer against leniency — "use the full 0-5 range, not only 3-5"; judge with a different model family to avoid self-preference; randomize A/B order to kill position bias.
Each grader is one scorecard column; adding a metric = appending one grader.
pass-slow — grade cost alongside correctness
A result is not just right/wrong. An efficiency grader downgrades a correct output that ran over a turn/token budget to pass-slow — so "correct but too expensive" is visible, not hidden inside a green pass.
The loop
- A fixed task set (
tasks), each with an input and a measurable expectation. - Run all graders over each task's output → a scorecard.
- Pin a baseline once; every later run shows signed deltas vs that baseline, not vs the previous run — so
re-running the same round shows real movement, not noise.
- Change one thing, re-run, read the deltas. Keep what moves the number up.
Noise floor
State it. At n=20 tasks, one task ≈ 5 points — deltas smaller than that are not meaningful. If the cheapest option already hits the ceiling, say so plainly instead of chasing a fractional gain.
Micro-test before you commit to a wording
Changing an instruction — a skill's phrasing, a rule in the discipline, an agent's trigger — is a change to behaviour, and the temptation is to reason about whether it reads better. Reading better and working better are different properties. Test it cheaply first:
- Sample it a handful of times, not once. Same prompt, same conditions.
- Against a no-guidance control — the identical task with the instruction absent. Without the control you
learn what the model does, not what your wording adds.
- Read every result by hand. At this size there is no statistic to hide behind; a score computed over four
runs is a number pretending to be evidence.
- Treat run-to-run variance as a warning, not noise to average away. If the same arm swings across runs,
the wording is not doing reliable work — and any delta you measure is smaller than the variance you have not controlled.
The failure this prevents, observed in Crewforth's own evals: a case scored 7/9 against 9/9 — the guidance apparently making things worse — and an identical second round came back 9/9 to 9/9. Two checks of variance inverted the finding. Had the first round been reported, a good rule would have been removed on noise.
Corollary: a delta smaller than the observed spread between identical runs is not a result. Say "below the noise floor" and either raise n or accept that the change is unmeasurable at this scale — both are honest; quoting the number is not.
The grader architecture, a starter grader catalogue, and the judge-bias checklist live in references/method.md.
DoD
- A fixed task set + a two-layer grader; a pinned baseline; every change reported as a signed delta with the noise floor stated.
- Any wording change was micro-tested against a no-guidance control, with every run read rather than averaged.
Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: crewforth
- Source: crewforth/crewforth
- License: MIT
- Homepage: https://crewforth.com/
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet, be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.