AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
SKILL verified MIT Self-run

Agentsop Domain Eval Set

skill-agentsope-skillalchemy-agentsop-domain-eval-set · by agentsope

Build and govern a 50-200 example domain-specific held-out benchmark sampled from real traffic. Distinct from public benchmarks (MMLU/HumanEval/GSM8K via lm-evaluation-harness) which measure GENERAL capability. Only a held-out domain set predicts whether THIS system works on YOUR data. Collect real examples, label, hold out (never train/prompt on it), size 50-200, version it, refresh on drift.

No reviews yet
0 installs
28 views
0.0% view→install

Install

$ agentstack add skill-agentsope-skillalchemy-agentsop-domain-eval-set

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access No
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets No
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/skill-agentsope-skillalchemy-agentsop-domain-eval-set)

Reliability & compatibility

Security review passed
0 installs to date
no reviews yet
2mo ago

Declared compatibility

Claude CodeClaude Desktop

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Agentsop Domain Eval Set? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

domain-eval-set — Your Held-Out Domain Benchmark

> "Compiled program beats baseline on a held-out test set (not the val set used in optimization)." > — DSPy SOP exit criterion [dspy.ai/learn/optimization/overview/]

> "Build the eval loop before optimizing anything. Every subsequent change must be gated on these numbers." > — LlamaIndex SOP Stage 2

This is an ENHANCE overlay skill. It produces one artifact — a versioned, sealed, human-labeled set of 50–200 examples drawn from your domain — that other skills consume: [[agentsop-regression-gate]] enforces it on every PR, [[agentsop-metric-design]] defines the scoring function applied to each example, and [[lm-evaluation-harness]] runs the complementary public-capability axis. The core claim: public benchmarks tell you the model is smart in general; only a held-out domain set tells you it works on your task. The latter is the one that predicts production.


1. 何时激活 (When to Activate)

Activate when any of these is true:

  • "Does THIS system work on OUR data?" — someone is about to ship or trust an

LLM/RAG/agent system and the only evidence is vibes, a demo, or a public benchmark number. You need a quantitative answer on the real distribution.

  • A public-benchmark number is being used as a deployment gate. Someone cites

"92% on MMLU" or "passes HumanEval" to justify go-live. That measures general capability, not your task fit (AP-1). Force a domain set into the decision.

  • A model / prompt / retriever / chunking change needs a regression gate and

no domain test set exists yet to gate against. You must build the set before [[agentsop-regression-gate]] can do its job.

  • Switching models (GPT-4o → a cheaper or local model). The public-bench gap

may be small while the domain gap is large, or vice versa. Only your held-out set tells you which.

  • Production complaints don't match your eval scores. Either the set is stale

(refresh, OP-DE06) or it never reflected the domain (rebuild from real traffic).

Do NOT activate for:

  • Pure capability comparison / academic reporting. "Which model is best at

MMLU/GSM8K?" → that is [[lm-evaluation-harness]], not this skill.

  • One-off throwaway prototypes where no decision rides on quality and nothing

ships. Don't build a benchmark for a script you'll delete tomorrow.

  • Tasks with an objective oracle already (compiler passes, exact DB match,

schema validity gives ≥95% of signal) — the "eval set" is just running the oracle; you don't need curated held-out examples. Don't gold-plate.


2. 核心心智模型 (Core Mental Model)

"Public benchmarks measure general capability. A 50–200 example held-out domain set measures YOUR task. Only the latter predicts production."

Two orthogonal axes, constantly confused:

| Axis | What it measures | Tool | Predicts production? | |---|---|---|---| | General capability | Reasoning, knowledge, coding in general, on shared public tasks | [[lm-evaluation-harness]] (MMLU, HumanEval, GSM8K, TruthfulQA) | No — a proxy at best | | Domain task fit | Whether the system answers your users on your data | this skill (held-out domain set) | Yes — this is the signal |

A model can score 90% on MMLU and 40% on your insurance-claims triage. A model can score below SOTA on HumanEval and be perfect at your internal codebase's patterns. The public number and the domain number are nearly uncorrelated once you're past a basic capability floor. The public bench is a sanity check; the domain set is the decision.

Three corollaries (each maps to an SOP stage):

  1. Real beats synthetic. The set is sampled from real domain traffic

(tickets, queries, logs, transactions), stratified, with edge cases pulled deliberately. Auto-generated QA pairs (LlamaIndex DatasetGenerator) are a fine bootstrap, but a model can ace generated questions and still fail real user phrasing. Generated sets do not replace a real held-out set (§7).

  1. Held out means SEALED. The held-out split is never shown to the optimizer,

never pasted into a prompt as a few-shot demo, never used to pick chunk size or reranker, never in the fine-tune data. The moment it leaks, the number is inflated and meaningless (AP-2, OP-DE07). Per DSPy: the test set must be distinct from the val set used in optimization [dspy.ai/learn/optimization/overview/].

  1. Small but significant. 50–200 examples. Below ~30 you are "memorizing, not

training" [dspy.ai/learn/optimization/overview/] and differences are noise. The set is small enough to label by hand and large enough to detect ~5–10pp regressions and to slice by segment.


3. SOP (Standard Operating Procedure)

0. Confirm activation (§1) — is the question "does this work on OUR data"?
1. COLLECT  — sample real domain examples; stratify; pull edge cases       (OP-DE01)
2. LABEL    — gold answer / reference / pass-fail; 2 annotators on subset  (OP-DE02)
3. HOLD OUT — split train/dev/test; SEAL the test split                    (OP-DE03)
4. SIZE     — land at 50-200; per-segment counts                           (OP-DE04)
5. VERSION  — hash + date + rubric; freeze as an artifact                  (OP-DE05)
6. LEAK-AUDIT — diff held-out vs demos / train / fine-tune data            (OP-DE07)
7. PAIR     — report alongside public bench; gate on the domain set        (OP-DE08)
   (later) REFRESH on domain shift                                         (OP-DE06)

Stage 1 — Collect from real traffic

Pull from where the real distribution lives: support tickets, search/query logs, user transcripts, transaction records, bug reports. Stratify so the set covers the production mix — by query type (lookup / summary / compare), by segment (tenant, language, product area), by difficulty. Then deliberately over-sample edge cases and known failures — the head of the distribution is easy; the tail is where systems break.

Target a raw pool ≥ 2× the final size (you'll drop ambiguous items in labeling). Record provenance and timestamp per example (needed later for drift refresh).

Exit: a candidate pool ≥ 2× target, with provenance, spanning the real mix.

Stage 2 — Label and curate

Attach ground truth per example: a gold answer, an acceptable reference response (not "the unique correct" one for open-ended tasks — see [[agentsop-metric-design]]), or a pass/fail label. For RAG, also label the gold passage so RetrieverEvaluator(["mrr","hit_rate"]) can run [LlamaIndex OP-10].

Have two annotators label a subset, measure agreement, resolve disagreements, and drop genuinely ambiguous items — an example two experts can't agree on will only add noise. Record the rubric. (This is the data-side analogue of DSPy's "human-validate the metric on ≥20 spot-checks" discipline [DSPy Case C].)

Exit: labeled set with inter-annotator agreement noted, rubric recorded, ambiguous items logged as rejected.

Stage 3 — Hold-out discipline (the load-bearing stage)

Split into train / dev / test. The test (held-out) split is sealed:

  • NEVER shown to an optimizer (DSPy trainset, MIPRO/GEPA).
  • NEVER pasted into a prompt as a few-shot demo.
  • NEVER used to pick chunk size / reranker / hybrid alpha / model.
  • NEVER in fine-tune data.

Store it in a separate file/location with an access note. Per DSPy, the exit-gate test set must be "distinct from the val set used in optimization" [dspy.ai/learn/optimization/overview/]. The dev split is what you tune against; the test split is the one number you trust at decision time.

Exit: sealed held-out test split + train/dev splits; access policy written.

Stage 4 — Size for 50–200

  • 50 — minimum for a coarse production go/no-go signal.
  • 100–200 — stable enough to detect ~5–10pp regressions and to slice per

segment (each slice needs its own ≥~30 to be meaningful).

  • ** 200

synthetic for the decision gate.

  1. Two-annotator a subset (OP-DE02) so you trust the gold labels; drop the

ambiguous ones rather than padding the count.

  1. Seal those 50 (OP-DE03), version them (OP-DE05). Iterate against the

synthetic dev set; report the go/no-go on the 50 real held-out.

  1. Grow it on real traffic post-launch (OP-DE06) — pilot usage is the cheapest

source of new labeled examples.

结果: A 50-example human-labeled, sealed held-out set built from the hardest real contracts predicts production far better than 200 synthetic questions or any public legal benchmark. The synthetic set still earns its keep — as the dev set you tune against, never as the number you trust.

可提取的操作: OP-DE01, OP-DE02, OP-DE03, OP-DE04. Lesson: spend scarce labels on a small REAL held-out set; let synthetic generation cover the dev set; never let a public bench be the gate.

Dilemma 2 — "Our eval set went stale; scores are green but production is on fire"

困境: A support-triage classifier shows 0.91 on eval_v1 (built 9 months ago) and every PR passes [[agentsop-regression-gate]]. Yet production accuracy collapsed and users are escalating. The eval set says everything is fine.

约束: eval_v1 is versioned and trusted; nobody wants to "move the goalposts". The domain shifted — a new product line generates a third of current tickets, and none of those ticket types existed when eval_v1 was built. Rebuilding costs annotator time.

决策步骤:

  1. Diagnose drift, not regression. Slice production traffic by ticket type and

compare against eval_v1's segment counts. The new product line is ~33% of live traffic and 0% of the eval set → the eval set no longer represents the domain. The green score is measuring an obsolete distribution.

  1. Do NOT just lower the threshold — the metric isn't wrong, the data is

stale. (Compare metric-design AP-8: changing the yardstick mid-stream without re-grounding.)

  1. Refresh (OP-DE06): sample recent real tickets — especially the new product

line and recent escalations — label them, retire ticket types that no longer occur, and build eval_v2.

  1. Bump the version (OP-DE05), keep eval_v1 for back-comparison. Re-score the

current system on eval_v2: it drops to 0.63 — now matching reality.

  1. Re-gate [[agentsop-regression-gate]] on eval_v2. Add a drift check to the

refresh cadence: quarterly, compare live segment mix vs eval segment mix; if any segment drifts >X%, trigger a refresh.

结果: The "green-but-on-fire" gap was a stale held-out set, not a model regression. A versioned refresh (eval_v2) restored the eval as a true production predictor; the back-comparison against eval_v1 documented exactly how much the domain moved.

可提取的操作: OP-DE06 RefreshOnDomainShift, OP-DE05 VersionTheSet. Lesson: a held-out set is a snapshot of a moving distribution. Schedule drift checks; an old green score can be the most dangerous number you have.


6. 反模式与边界 (Anti-Patterns & Boundaries)

Anti-patterns

| # | Anti-pattern | Why it's wrong | Fix | |---|---|---|---| | AP-1 | Public bench as proxy for domain performance ("92% MMLU → ship it") | Public benches measure general capability; near-uncorrelated with task fit past a floor | Build a domain held-out set; gate on it (OP-DE08) | | AP-2 | Eval set leaks into prompt / training / trainset | Held-out number is inflated and meaningless; you're testing on the train set | Seal it; leak-audit before release (OP-DE03, OP-DE07) | | AP-3 | Set too small to be significant (<30 examples) | "Memorizing, not training" [dspy.ai/learn/optimization/overview/]; variance swamps signal | Target 50–200 (OP-DE04) | | AP-4 | Synthetic-only held-out (auto-generated QA is the test set) | Tests "what the corpus says", not real user phrasing; flatters the system | Synthetic = dev set bootstrap only; real-labeled = held-out (§7, Dilemma 1) | | AP-5 | Never refreshing as the domain drifts | Green scores on an obsolete distribution; "green but on fire" (Dilemma 2) | Schedule drift checks; refresh + version (OP-DE06) | | AP-6 | Unversioned set silently edited | Can't compare across system versions; results table is meaningless | Hash + date + rubric; bump on deliberate refresh (OP-DE05) | | AP-7 | No stratification / edge cases (only easy head-of-distribution) | Passes eval, fails the tail where systems actually break | Stratify by segment/type; over-sample edge cases (OP-DE01) | | AP-8 | Tuning chunk size / reranker / model against the held-out set | That makes it a val set, not held-out; the trust is gone | Tune on dev; touch held-out only at decision time (OP-DE03) |

Boundaries (when this skill is the wrong tool)

  • You only need general capability comparison / academic reporting. "Which

model is best at reasoning?" → [[lm-evaluation-harness]] (MMLU/GSM8K/etc.), not this skill. This skill is for your task, not the leaderboard.

  • An objective oracle already exists (compiler passes, exact DB match, schema

validity gives ≥95% of signal). The "eval set" is just running the oracle on inputs — you don't need curated human-labeled held-out examples. Don't gold-plate.

  • Nothing ships and no decision rides on quality (throwaway prototype). The

cost of building and labeling a real set has no payoff.

  • You have zero access to real domain data and no path to any (pre-product,

cold start). Bootstrap with synthetic + public benches transparently, label as soon as pilot traffic appears, and treat early numbers as provisional.

  • Scoring each example is itself the hard part (open-ended generation, no

reference answer) — building the set is necessary but not sufficient. Pair with [[agentsop-metric-design]] to define a defensible, calibrated scoring function.


7. 跨框架对照 (Cross-Framework Mapping)

When does each kind of eval set apply? They are complementary axes, not substitutes — a mature pipeline uses all three.

| Concept | Held-out domain set (this skill) | [[lm-evaluation-harness]] (public) | LlamaIndex DatasetGenerator (synthetic) | |---|---|---|---| | What it measures | Task fit on your data | General capability | Coverage of your corpus's content | | Data source | Real traffic, human-labeled | Public academic datasets (MMLU, HumanEval, GSM8K, TruthfulQA, HellaSwag) | LLM-generated QA from your docs | | Size | 50–200 | thousands (fixed by benchmark) | arbitrary (default num=50) | | Contamination risk | You control it (leak-audit) | High — public benches leak into pretraining | Low (your private corpus) but synthetic | | Predicts production? | Yes (the decision gate) | No (capability floor / sanity check) | Partially (dev-set iteration, not the gate) | | When to use | Go/no-go on shipping to your users; per-PR regression gate | Model selection on raw capability; academic reporting; training-progress tracking | Bootstrap a dev set fast before you've labeled real data | | Invocation | eval_vN.jsonl + scoring fn from [[agentsop-metric-design]] | lm_eval --tasks mmlu,gsm8k,... | DatasetGenerator.from_documents(docs).generate_dataset_from_nodes(num=50) |

Decision rubric:

Q1. Are you deciding whether to SHIP / SWITCH on YOUR users' data?
    YES → held-out domain set is the gate (this skill). Public bench = sanity check only.
Q2. Are you comparing raw model capability or reporting academic numbers?
    YES → lm-evaluation-harness (MMLU/HumanEval/GSM8K). Not this skill.
Q3. Do you have NO real labeled data yet but a corpus exists?
    YES → DatasetGenerator to bootstrap a DEV set; label real held-out as soon as traffic appears.
Q4. Is there an objective oracle (tests/schema/exact-match)?
    YES → run the oracle; no curated set needed.
DEFAULT → build + version a 50-200 real held-out set; gate via [[agentsop-regression-gate]];
          score via [[agent

…

## Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

- **Author:** [agentsope](https://github.com/agentsope)
- **Source:** [agentsope/SkillAlchemy](https://github.com/agentsope/SkillAlchemy)
- **License:** MIT

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.