AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
SKILL verified MIT Self-run

Create Skill Autoresearch

skill-a-tokyo-agent-skills-harness-create-skill-autoresearch · by a-tokyo

>-

No reviews yet
0 installs
15 views
0.0% view→install

Install

$ agentstack add skill-a-tokyo-agent-skills-harness-create-skill-autoresearch

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access Used
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets No
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/skill-a-tokyo-agent-skills-harness-create-skill-autoresearch)

Reliability & compatibility

Security review passed
0 installs to date
no reviews yet
1mo ago

Declared compatibility

Claude CodeClaude Desktop

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Create Skill Autoresearch? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

Create Skill via Autoresearch Factory

A factory for forging production-grade agent skills through gold-standard-driven autoresearch, multi-agent verification, and structured consensus.

The factory orchestrates 4 agent roles through 5 phases:

| Phase | What Happens | Agent Role | |-------|-------------|------------| | 1. Interview | Discover purpose, gold standards, scope | ORCHESTRATOR | | 2. Research | Study domain materials, build dossier, propose rubric | RESEARCHER (N parallel) | | 3. Draft | Design structure, generate SKILL.md, measure baseline | BUILDER | | 4. Autoresearch | Iterate skill against gold standards (LLM-as-judge, or an objective real-world metric for procedural skills — see 3.4) | BUILDER + autoresearch skill | | 5. Verify | Premortem, panel scoring, consensus, ship/iterate | PANEL (3 subagents) |

Key constraint: BUILDER and PANEL never share context. Panel receives only the skill output, gold standards, and rubric -- no bias from the building process.

Relation to create-skill

This factory extends the official single-pass skill creators (Anthropic's Skills best-practices and skill-creator; Cursor's create-skill) rather than replacing them. It adds what a one-shot generator cannot: a research dossier, gold-standard benchmarking, an autonomous improvement loop, and independent multi-agent verification. The skills it produces follow the same official conventions -- see [references/skill-authoring-best-practices.md](references/skill-authoring-best-practices.md).

Companion skills

The factory orchestrates these sibling skills at runtime: autoresearch (Phase 4 improvement loop), premortem (Phase 5 risk pass), and handoff (cross-session continuity); the Phase 5 panel/consensus design draws on llm-council. In this harness they are vendored under .agents/skills/. If you install this skill standalone, install those alongside it. The factory's craft layer ([references/skill-craft-principles.md](references/skill-craft-principles.md)) is distilled from writing-great-skills (mattpocock/skills, MIT), vendored in this repo's .agents/skills/ alongside Anthropic's skill-creator.


Phase 1: Interview

Discover what the user needs through structured questions. Do not assume -- ask.

1.1 Purpose and Domain

> What skill do you want to build? What problem does it solve? > > Describe the domain, the target user (which agent will use this skill), > and what "success" looks like when the skill is used correctly.

Record: SKILL_PURPOSE, DOMAIN, TARGET_USER, SUCCESS_CRITERIA.

1.2 Gold Standards

> Do you have examples of "what good looks like"? > > Gold standards are the benchmark. They can be: > - Input/output pairs: given this input, the skill should produce output like this > - Reference artifacts: existing documents, code, or outputs that represent ideal quality > - Previously solved problems: tasks that were completed successfully by humans > - Quality reports: existing evaluations or benchmarks > > Where are they? What format are they in? How many do you have?

Record: GOLD_STANDARD_SOURCE, GOLD_STANDARD_FORMAT, GOLD_STANDARD_COUNT.

Minimum: 3 gold standards. Fewer than 3 is a risk -- warn the user and suggest alternatives (create synthetic examples, find additional reference materials).

1.3 Study Materials

> What materials should I study to understand this domain? > > Examples: documentation, existing code, transcripts, design docs, > reference implementations, specifications, style guides.

Record: STUDY_MATERIALS (list of paths/URLs).

1.4 Scope and Constraints

> Any constraints on the skill itself? > > - Must it follow specific conventions? (e.g., existing team patterns) > - Are there skills it should integrate with? > - Any anti-patterns to avoid? > - Target line count? (default: - Invocation mode: should the agent fire this skill on its own (model-invoked, pays > permanent context load) or only the human (user-invoked, disable-model-invocation: true)?

Record: CONSTRAINTS, INTEGRATION_SKILLS, ANTI_PATTERNS, INVOCATION_MODE.

1.5 Existing Skill Check

> Is there an existing skill for this domain that we're upgrading? > > If yes, that skill becomes a study material AND a baseline. The factory will > research it, measure it against the rubric, then improve it -- not start from scratch.

If an existing skill is found, record EXISTING_SKILL path and set the factory mode to upgrade (baseline from existing) rather than greenfield (baseline from scratch).

1.6 Confirm and Create Workspace

Summarize all parameters in a table. Ask the user to confirm.

Once confirmed, create the build workspace at builds// with three ownership zones:

  • input/ -- where the user drops gold standards + study materials, in any structure
  • work/ -- everything the factory generates: manifest.yaml, research/, evaluation/, experiments/, handoffs/
  • output// -- the finished skill (SKILL.md + references/) in its own named dir, publish-ready

Do not ask the user to hand-author a manifest. Scan whatever is in input/, classify each item as a gold standard (exemplar input/output pair or reference artifact) vs a study material, and write your derived index to work/manifest.yaml with train/validation/test tags. Present the derived manifest for the user to confirm or correct. See [references/pipeline-phases.md](references/pipeline-phases.md) for intake formats and the manifest schema.


Phase 2: Research

Study the domain thoroughly before writing any skill code.

2.1 Spawn Researcher Subagents

Cluster study materials by relatedness, then launch one explore subagent per cluster. Clustering heuristic:

  • By source type: existing skills in one cluster, gold standard outputs in another, planning docs in a third
  • By subtopic: if materials cover distinct areas (e.g., backend vs frontend), split by area
  • Cap at 5-7 clusters: more than 7 creates synthesis overhead without proportional depth gain
  • Minimum 2 clusters: a single cluster means no parallelism benefit

Each subagent:

  1. Reads the assigned material deeply
  2. Distills findings into a research note in work/research/
  3. Identifies patterns, conventions, and quality signals relevant to the skill

Naming: work/research/01-.md, work/research/02-.md, etc.

2.2 Synthesize Research

After all researchers complete, synthesize findings into work/research/00-synthesis.md:

  • Cross-cutting patterns
  • Key conventions the skill must follow
  • Quality signals that distinguish good from bad output
  • Potential rubric dimensions

2.3 Propose Rubric

Based on research, draft work/evaluation/rubric.yaml:

name: -rubric
dimensions:
  - name: 
    weight: 
    scale: "1-10"
    criteria: ""
  # ... 5-10 dimensions
target_score: 0.85
max_iterations: 20
plateau_window: 5

Always include these universal dimensions (adjust weights per domain):

  • correctness: Instructions are technically accurate and executable
  • completeness: All necessary sections and edge cases covered
  • clarity: A naive agent can follow without ambiguity
  • consistency: Aligns with existing codebase conventions
  • predictability: Drives the same process every run -- completion criteria checkable and exhaustive, no vague gates, no no-op lines (see [references/skill-craft-principles.md](references/skill-craft-principles.md))

Add 3-5 domain-specific dimensions from the research synthesis (5-10 dimensions total).

Present the rubric to the user for review. Iterate until confirmed.

See [references/rubric-templates.md](references/rubric-templates.md) for templates.


Phase 3: Draft

Design before writing. Write before measuring.

3.1 Design Document

Create work/experiments/DESIGN.md with:

  • Skill name and description (following create-skill conventions)
  • Structural decisions: section count, reference file split, progressive disclosure plan
  • Invocation mode, information-hierarchy plan (steps vs reference; inline vs disclosed,

licensed by branching), and candidate leading words -- see [references/skill-craft-principles.md](references/skill-craft-principles.md)

  • Integration points with other skills
  • Key terminology and voice decisions

3.2 Grill the Design

Before writing any skill code, challenge the design adversarially:

  • What would make this skill fail in practice?
  • Are the structural decisions justified or assumed?
  • Does the design match what the gold standards demonstrate?
  • Are there simpler alternatives?

Present concerns to the user. Iterate until the design survives scrutiny.

3.3 Generate SKILL.md Draft

Following the design and the official skill-authoring rules (see [references/skill-authoring-best-practices.md](references/skill-authoring-best-practices.md)), run this pre-flight checklist before writing -- these are hard constraints, not preferences:

  • name: 100 lines
  • Concrete examples over abstract instructions; consistent terminology; forward-slash paths
  • Description craft: leading word front-loaded, one trigger per branch, no synonym padding;

user-invoked skills get a one-line human-facing description ([references/skill-craft-principles.md](references/skill-craft-principles.md))

Write the draft to output//SKILL.md (reference files in output//references/).

3.4 Build Evaluation Script

Create work/evaluation/evaluate.sh that:

  1. Takes a gold standard test case path as argument
  2. Extracts the input from the test case
  3. Invokes the skill on the input -- since skills are markdown instructions (not

executables), this means calling an LLM with the SKILL.md as a system prompt and the test case input as the user message. Use curl to an OpenAI-compatible API, or a language-specific SDK. Capture the LLM's output.

  1. Compares the output to the gold standard reference using an LLM-as-judge
  2. Emits METRIC = lines to stdout
  3. Emits METRIC overall_score= as the primary metric

See self-test/evaluation/evaluate.sh in the agent-skills-harness repo for a complete reference implementation.

The LLM judge should:

  • Use structured JSON output for per-dimension scoring
  • Score each dimension independently (prevent halo effects)
  • Require evidence (verbatim quotes) for extreme scores
  • Use a different model family from the builder when possible

Deterministic vs LLM-judge evaluation: Not every dimension needs an LLM judge. Prefer deterministic checks where possible:

  • Line count, frontmatter validation, link integrity → shell/grep checks
  • Pattern coverage (does output mention X?) → regex matching
  • Structural conformance → programmatic validation

Use LLM-as-judge only for dimensions that require subjective judgment (clarity, quality match, curation). Mix both in evaluate.sh: deterministic checks emit METRIC lines directly, LLM judges handle the rest. If no LLM API is available, fall back to deterministic-only scoring and log a warning.

Procedural / agentic skills (prefer this when it applies): Some skills don't generate an artifact in one shot — they instruct an agent to perform a multi-step task on a real artifact (migrate a framework version, refactor a module, scaffold infra, rebuild a repo). For these, the single-call "SKILL.md as system prompt + input as user message" model in 3.4 is the wrong harness. Evaluate them by execution against a real artifact with an objective real-world metric instead:

  1. Pick a real test repo where "correct" has a ground-truth signal — ideally one where a correct application is a no-op against a captured baseline. (Example: a Tailwind v3→v4 migration should be visually identical, so committed golden screenshots become the gold standard; the metric is pixel-parity + build/lint/typecheck/tests pass + a static "residual v3 markers = 0" grep.)
  2. evaluate.sh orchestrates: reset → fresh-agent applies the skill → measure. Reset the repo to the captured baseline; spawn a fresh subagent told to perform the task following only the skill under test (no other guides, no builder context); then run deterministic real-world checks on the result and emit METRIC lines. The artifact's own ground truth replaces the LLM judge — cheaper, deterministic, and far stronger signal than judging prose.
  3. Fresh agent every run is the point. It measures skill self-sufficiency, not the builder's accumulated context. A clean reset between runs (e.g. git reset --hard && git clean -fd + reinstall) is mandatory or scores drift. Vary the executor model (e.g. a smaller model) as a robustness check — if a smaller model + the skill still hits the target, the skill is robust.
  4. Capture the baseline before drafting. On the unmodified repo, set up the metric (golden snapshots / test suite), confirm it's deterministic (run it twice, expect identical), and confirm the unmigrated state scores 0 so you know the gate discriminates.

This makes the codemod/tool-first pattern natural too: have the skill run any deterministic tool (a codemod, a formatter, a generator) for the mechanical bulk first, and reserve the skill's prose for the judgment the tool can't do — the real-world metric then verifies the whole.

For multi-judge evaluation (recommended when budget allows):

  • Run 2-3 different LLM models as judges on the same output
  • Average their per-dimension scores for a more robust signal
  • Track per-judge variance -- high variance on a dimension indicates the criteria may be ambiguous
  • Configure judges in work/evaluation/judges.yaml:

```yaml judges:

  • model: ""

weight: 1.0

  • model: ""

weight: 1.0 aggregation: "mean" ```

Optionally create work/evaluation/evaluate-checks.sh for correctness gates.

3.5 Measure Baseline

Run evaluate.sh on the test cases with the initial draft. Record baseline scores. This is experiment 0.

Report to the user: > Baseline established: overall_score = [value] > Dimensions: [per-dimension breakdown]


Phase 4: Autoresearch

Invoke the autoresearch skill to iterate the skill draft against the evaluation rubric.

4.1 Configure Autoresearch

Provide these parameters to the autoresearch skill. All paths are relative to the build workspace root (builds//), which is the autoresearch working directory. Autoresearch session files (.md, .jsonl, .tsv, run.log) are created at the workspace root during the active session, then archived to work/experiments/ when the session ends or on handoff.

  • Goal: Improve ` quality as measured by overall_score` (LLM-as-judge against gold standards, or the objective real-world metric for procedural skills — see 3.4)
  • Metric command: ./work/evaluation/evaluate.sh (relative to workspace root)
  • Primary metric: overall_score
  • Direction: higher_is_better
  • In-scope files: output//SKILL.md, output//references/*
  • Out-of-scope files: input/, work/
  • Constraints: Must follow the official skill-authoring rules (= 10:
  • 70% training: Used during each autoresearch experiment
  • 20% validation: Checked adaptively to detect overfitting (see below)
  • 10% test: Held out entirely until Phase 5 verification

If gold standards count 3-9:

  • Leave-one-out rotation: Each experiment evaluates against all but one, rotating which is held out

Record the split in work/evaluation/data-split.yaml.

Cost awareness for large sets (100+ gold standards): Each LLM-as-judge call costs real money. With 70 training cases at ~$0.50/call, that's ~$35/experiment. Mitigate with a sampling strategy: evaluate against a random sample of training cases per experiment (e.g., 10-15), rotating the sample. Run the full training set only when validating kept experiments or at phase boundaries.

Overfitting detection:

Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.