Install
$ agentstack add skill-tollens-ai-quality-strategy-skills-tooling-adequacy ✓ scanned · ✓ verified — works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ✓ Network access No
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
About
Tooling Adequacy
This skill answers the second of the four quality questions — "How do we know if what we have is good?" — for testing. It checks whether the test strategy's tools and judging methods can actually answer its learning needs (the questions the strategy wants testing to answer).
Answering a learning need takes two distinct capabilities, and both can fail independently:
- an instrument — to exercise the system and observe a result (run it, drive it, capture output);
- an oracle — to judge whether that observed result is correct.
A learning need is only answerable if both are adequate. A perfect instrument with no oracle means you can run the thing a million times and still not know if it's right; a perfect oracle with no instrument means you know what "right" looks like but can't produce a result to check.
This skill exists because of two reliable agent failure modes. (1) When agents collapse "how do we know?" into "is it good?", they defer to whatever tooling already exists, assume it's adequate, and build findings on an instrument that was blind to the question. (2) When there's no obvious oracle, agents fall back on the old-world reflex — "no oracle, so this isn't really testable" — and silently drop the learning need. But in the agent era an oracle is usually cheap to build.
Resolving file paths — do this first
This skill is part of the quality-strategy plugin. Before anything else, resolve two absolute paths and use them throughout:
- PLUGIN_ROOT — the plugin's install directory:
${CLAUDE_PLUGIN_ROOT}(Claude Code expands this to an absolute path when it loads this file; read it off and note it down). The grounding files this skill reads —PHILOSOPHY.md,skills/test-strategy/FRAMINGS.md— live under it. - PROJECT_DIR — the absolute path of the project whose tooling and oracles you're assessing (normally the current working directory; confirm with the user if it's ambiguous). The strategy docs normally live under
$PROJECT_DIR/quality/— but/quality-strategyasks at session start where the strategy should be saved, so they may live elsewhere. If$PROJECT_DIR/quality/is absent, get the docs home instead of assuming: from the orchestrator's brief when this skill was dispatched, else by asking the user; if the path you're given ends in/quality, its parent is the home. From then on treat$PROJECT_DIRas that docs home wherever a path below says$PROJECT_DIR/quality/...— one substitution, made once, before you act on any path.
File references below use the $PLUGIN_ROOT and $PROJECT_DIR placeholders. Substitute the resolved absolute paths before you act on them. The Read tool does no variable expansion and resolves relative paths against the current working directory, not this skill's directory — so an unsubstituted placeholder or a bare relative path will fail.
When to use
- From
/test-strategy— offered during the per-ility discussion when whether the existing means can answer a question at all becomes the live issue (the instrument can't exercise it, or nothing can judge the output). Input: the questions (learning needs) under discussion with their proposed methods. Output: a tooling-and-oracle adequacy assessment plus a list of build items (things to build or acquire) the test strategy records against its agreed moves — and/tooling-strategyconsumes. - Standalone — to audit an existing codebase's test tooling and oracles against what the team wants to find out. Input: the questions the team wants answered (ask the user) plus the repo's test infrastructure.
This skill judges adequacy and names the gaps; it does not plan the build. /tooling-strategy consumes its build items (together with /oracle-adequacy's, from the quality side) and turns them into a prioritised build plan.
What you need
- Grounding. Read
$PLUGIN_ROOT/PHILOSOPHY.mdand$PLUGIN_ROOT/skills/test-strategy/FRAMINGS.md. Framings #1 (investigation, not checking), #4 (asking and testing are parallel), #5 (don't import old-world costs — central to the oracle reframe below), #6 (economics: checking is cheap, investigation is the bottleneck) and #9 (smells — humans are the oracle for some questions) are load-bearing here. - The learning needs and their methods. From
/test-strategy: the questions the agreed moves answer, in$PROJECT_DIR/quality/test-strategy.md. Standalone: ask the user what they want to find out, and capture it as a short list of questions. - The test infrastructure inventory. From
$PROJECT_DIR/quality/test-pre-read.mdif it exists (a legacy artefact of the older test-strategy shape); otherwise a quick filesystem pass (test dirs, CI config, test commands, config files). Do not read source code — judge the means of finding out from the strategy, the infrastructure inventory, and conversation, to keep the independent perspective testing relies on (FRAMINGS #3).
The work, in order
1. Map each learning need to its instrument and its oracle
For each learning need (standalone: each thing the team wants to learn), state the two things its proposed method depends on:
- Instrument — how you'll exercise the system and observe a result. Be specific: "run Y under condition Z, capturing W."
- Oracle — how you'll decide the observed result is correct. Name it explicitly. A method that observes a result but can't say what "correct" looks like has no oracle yet — that's a finding, not a detail.
2. Assess the instrument — Adequate / Inadequate / Missing
- Adequate — exists and can exercise/observe this question at the fidelity required.
- Inadequate — exists but can't reliably produce the needed observation (the integration suite only runs on a perfect network; unit tests can't drive real concurrency).
- Missing — nothing exists; it must be built or acquired.
3. Assess the oracle — Adequate / Inadequate / Missing — and don't accept "there's no oracle"
Classify the oracle on the same scale. The oracle is whatever lets you decide a result is correct. The kinds, from cheapest signal to richest:
- Specified — a spec, contract, or known-correct value states the expected answer.
- Property / metamorphic — invariants that must hold even when you don't know the exact answer ("intervals never shrink on a correct recall"; "encode∘decode = identity"; "A→B→A round-trips").
- Differential / simulated — an independent, deliberately-simple, obviously-correct (if slow) reference implementation that you diff the real system against. An agent can often write one cheaply — the classic simulated oracle. Also: a prior version, or a competitor, as the reference.
- Golden master / snapshot — a captured known-good output, re-judged whenever it changes.
- Human or agent-judge — for trust, feel, and quality questions a person is the oracle; agents lack smells, so humans stay the oracle there (FRAMINGS #9). An agent-judge panel can scale fuzzy judgements where that's appropriate, but name it as the oracle and note its limits.
Kill the old-world reflex (FRAMINGS #5). "There's no oracle for this, so we can't test it" was often true when building an oracle meant expensive human work; it's usually false now. When the oracle is Missing or Inadequate, the default move is to propose building one — most often a property or a simulated/reference oracle — as a build item, not to drop the learning need.
4. Verdict and build items
A learning need is answerable only if both instrument and oracle are Adequate. For every instrument or oracle marked Inadequate or Missing, name the build item: what must be built or acquired, and which learning need(s) it unblocks. Oracle-build items (write a reference implementation; define the properties; write down the expected outputs) count just as much as instrument-build items. These are what the test strategy records as blocked-on-tooling against its agreed moves.
5. Catch the mismatches (FRAMINGS #1, #6, #9)
- Checking aimed at investigation — tooling/oracles that only verify known expected behaviour where the learning need actually calls for finding out what's true.
- Automation aimed at judgement — a learning need about trust/feel/quality assigned a purely automated oracle. The human is the oracle there; say so.
Push back when
- Every instrument or oracle comes back "Adequate" with no reasons. That's the Q2-collapse failure mode — re-run, and demand a specific reason for each item against this question.
- A Missing oracle is treated as "untestable." Challenge it: "under agent costs, could we write a simple reference implementation, or state a property, that judges this?" (FRAMINGS #5).
- A trust/feel/quality learning need is handed a purely automated oracle (FRAMINGS #9). The human is the instrument and the oracle; pretending a tool covers it is the inadequacy.
- The user wants to treat "we have a test suite / CI" as evidence of adequacy. Adequacy is per-question, not per-repo — a suite can be excellent for some learning needs and blind to others, and a green suite says nothing about questions it has no oracle for.
This skill is DONE when
- [ ] Every learning need names both its instrument and its oracle.
- [ ] Each is classified Adequate / Inadequate / Missing, and every "Adequate" carries a specific reason for this question (not "tests exist").
- [ ] Every Missing/Inadequate oracle has either a proposed constructed oracle (property / simulated-reference / golden-master / human-or-agent-judge) or an explicit, defended statement that none is feasible.
- [ ] A build item is named for every Inadequate/Missing axis (instrument and oracle), each tied to the learning need(s) it unblocks.
- [ ] Checking-vs-investigation and automation-vs-judgement mismatches are flagged.
- [ ] No source code was read.
- [ ] (When run from
/test-strategy) the build items are returned to the orchestrator so the affected moves can be marked blocked-on-tooling, and the assessment is written to the scratch path the brief names (see Output).
Output
When run from /test-strategy, return the assessment to the orchestrator and write it to the scratch path the orchestrator's brief names (e.g. quality/.scratch/tooling-adequacy-.md) — hard evidence this Q2 check actually ran; /test-strategy-review FAILs a claimed-but-absent audit. The orchestrator's brief carries the absolute docs-home path — a sealed dispatch can't ask where the docs live, so write where the brief says, never a path derived from your own working directory. Standalone, surface it in the conversation and offer to write it to $PROJECT_DIR/quality/tooling-adequacy-.md. Shape:
# Tooling & oracle adequacy
*Can we actually find these things out — and would we know if the answer were wrong?*
| Learning need | Instrument (exercise/observe) | Oracle (judge correctness) | Answerable? | Build items |
|---|---|---|---|---|
| … | … — Adequate/Inadequate/Missing | … — Adequate/Inadequate/Missing | Yes / Blocked | … |
## Build items (gating)
- **** — unblocks .
(Or: "None — every learning need has an adequate instrument and oracle, with the per-item reasons above.")
Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: tollens-ai
- Source: tollens-ai/quality-strategy-skills
- License: Apache-2.0
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet — be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.