# Clawpathy Autoresearch

> Eval-driven skill tuning. Given a task and an LLM-judge rubric, iteratively rewrites a SKILL.md until a downstream executor agent performs well against the judge. Low-code: all evaluation

- **Type:** Skill
- **Install:** `agentstack add skill-clawbio-clawbio-clawpathy-autoresearch`
- **Verified:** Yes — security-reviewed for prompt injection and unsafe behavior
- **Seller:** [ClawBio](https://agentstack.voostack.com/s/clawbio)
- **Installs:** 0
- **Category:** [AI & ML](https://agentstack.voostack.com/c/ai-and-ml)
- **Latest version:** 0.1.0
- **License:** MIT
- **Upstream author:** [ClawBio](https://github.com/ClawBio)
- **Source:** https://github.com/ClawBio/ClawBio/tree/main/skills/clawpathy-autoresearch
- **Website:** https://clawbio.github.io/ClawBio/

## Install

```sh
agentstack add skill-clawbio-clawbio-clawpathy-autoresearch
```

Requires the [AgentStack CLI](https://agentstack.voostack.com/docs/cli). Works with Claude Code, Cursor, and any MCP-compatible agent.

## About

# clawpathy-autoresearch

Eval-driven skill development. The system iteratively rewrites a `SKILL.md`
so a downstream executor agent performs better at a task class, as judged
by an LLM against a paper/task-specific rubric.

## Core idea

```
  propose (sonnet)  →  execute (sonnet, shell)  →  judge (opus, rubric)
       ↑                                                       │
       └──────── feedback: verdict + recommended edits ────────┘
```

- **Proposer** rewrites SKILL.md based on the last judge verdict.
- **Executor** runs the new SKILL.md end-to-end inside a workspace.
- **Judge** scores methodology (primary) and outputs (secondary) against
  a per-task rubric. Lower is better; 0 = perfect.
- Keep the new SKILL.md only if it strictly beats the best score; else
  revert. Stop on target_score or on `early_stop_n` consecutive regressions.

## You are the orchestrator

You (the agent reading this) don't run the loop yourself. You dispatch
subagents to build the workspace, then hand off to the Python loop.

### Phase 1 — Scout

Dispatch a subagent with `prompts/scout.md` to research the paper/task.
Report key findings to the user in a few lines.

### Phase 2 — Scope (you + user)

Have a conversation. Ask ONE question at a time, multiple-choice where
helpful. Agree on:
- what to reproduce / what success looks like
- which data sources are in-bounds
- what methodology expectations belong in the rubric
- iteration budget and target_score (if any)

Present a summary and get approval.

### Phase 3 — Build

Dispatch a builder subagent with `prompts/builder.md` and the agreed
scope. It writes:
- `task.json`
- `rubric.md` — **the authoritative scoring rubric for the LLM judge**
- `reference/` (optional; judge-only)
- `skill/SKILL.md` — seed

Validate:
```python
from skills.clawpathy_autoresearch import validate_workspace
print(validate_workspace(Path("WORKSPACE")))  # [] means valid
```

### Phase 4 — Loop

```bash
python -m skills.clawpathy_autoresearch WORKSPACE_DIR
# or with custom models:
python -m skills.clawpathy_autoresearch WORKSPACE_DIR \
  --proposer-model sonnet --executor-model sonnet --judge-model opus
```

The loop streams progress to `WORKSPACE/history.jsonl`, snapshots every
iteration's skill to `WORKSPACE/snapshots/iter-NNN.md`, and writes the
executor's full transcript to `WORKSPACE/executor_runs/iter-NNN.log`.

## Workspace layout

```
workspace/
  task.json                  # task metadata + loop knobs
  rubric.md                  # LLM-judge rubric (the heart of the system)
  reference/                 # optional ground truth, judge-only
  skill/SKILL.md             # iterated by the loop
  output/                    # executor outputs (cleared each iter)
  executor_runs/iter-NNN.log # transcripts (judge reads these)
  snapshots/iter-NNN.md      # per-iter SKILL.md snapshots
  history.jsonl              # one row per iter: score, kept, verdict
```

## Key principles

- **LLM judge only.** No deterministic Python scorers. All evaluation goes
  through `judge.md` + opus. This keeps the system low-code and lets the
  rubric carry paper-specific nuance without adding code.
- **Methodology is primary.** The rubric weights "did the agent use sound
  methods?" above "did the numbers match?". Ground-truth match is a signal,
  not the objective — the goal is better SKILL.md files.
- **Never leak ground truth.** `reference/` is judge-only. The executor
  prompt says not to read it, and the judge penalises leakage.
- **No hardcoded answers in SKILL.md.** The proposer prompt and the judge
  both enforce this. The executor must derive results by running methods.
- **Snapshots + strict-better revert.** Score on the first iter becomes the
  floor. Later iters that tie or regress revert to the best.

## Safety

- All processing is local except scout web fetches for public resources.
- ClawBio disclaimer: research/education tool, not a medical device.

## Gotchas

- **Do not skip scoping.** The rubric is paper-specific; a generic rubric
  tunes nothing. Get the user to agree on methodology expectations.
- **Do not write a Python scorer.** Earlier versions of this project did.
  They rewarded API-fetching, not methodology. The judge is the scorer.
- **Do not hand-pick the "best" snapshot yourself.** Trust the loop. If
  the judge is calibrated wrong, fix the rubric, not the history.

## Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

- **Author:** [ClawBio](https://github.com/ClawBio)
- **Source:** [ClawBio/ClawBio](https://github.com/ClawBio/ClawBio)
- **License:** MIT
- **Homepage:** https://clawbio.github.io/ClawBio/

Install and usage instructions live in the source repository linked above.

## Pricing

- **Free** — Free

## Security capabilities

Automated source analysis of v0.1.0 — what this tool can access:

- **Network access:** no
- **Filesystem access:** no
- **Shell / process execution:** no
- **Environment & secrets:** no
- **Dynamic code execution:** no

*"Yes" means the capability is present in the source — more access means more to trust, not that it is unsafe.*


## Versions

- **0.1.0** — security scan: passed — Imported from the upstream source.

## Links

- Listing page: https://agentstack.voostack.com/l/skill-clawbio-clawbio-clawpathy-autoresearch
- Seller: https://agentstack.voostack.com/s/clawbio
- Browse the marketplace: https://agentstack.voostack.com/browse

---
Listed on AgentStack — the marketplace for AI agent skills and MCP servers. Every listing is security-reviewed. Creators keep 70%.
