AgentStack
SKILL verified MIT Self-run

Autoresearch Ml

skill-proyecto26-autoresearch-ai-plugin-autoresearch-ml · by proyecto26

>-

No reviews yet
0 installs
14 views
0.0% view→install

Install

$ agentstack add skill-proyecto26-autoresearch-ai-plugin-autoresearch-ml

✓ scanned · ✓ verified — works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access No
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets No
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

Are you the author of Autoresearch Ml? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

Autoresearch ML: Autonomous LLM Training Optimization

An autonomous experiment loop for single-GPU LLM pretraining. Edit train.py → commit → run 5-minute training → measure val_bpb → keep improvement or revert → repeat forever.

This skill is self-contained — it includes everything needed to set up and run the loop.

Setup Phase

1. Copy Template Assets

Copy the bundled training template to the project directory:

cp ${CLAUDE_SKILL_DIR}/assets/prepare.py .
cp ${CLAUDE_SKILL_DIR}/assets/train.py .
cp ${CLAUDE_SKILL_DIR}/assets/pyproject.toml .
cp ${CLAUDE_SKILL_DIR}/assets/program.md .

2. Install and Prepare

uv sync                    # Install dependencies
uv run prepare.py          # Download data shards, train tokenizer (~2 min)

3. Verify GPU

nvidia-smi
python -c "import torch; print(f'CUDA: {torch.cuda.is_available()}, Device: {torch.cuda.get_device_name()}, VRAM: {torch.cuda.get_device_properties(0).total_mem / 1e9:.1f} GB')"

4. Initialize the Experiment Session

  1. Create a branch: git checkout -b autoresearch/-
  2. Ensure session files are gitignored (critical — git revert will fail if tracked):

``bash echo -e "autoresearch.jsonl\nrun.log" >> .gitignore git add .gitignore && git commit -m "autoresearch: add session files to gitignore" ``

  1. Read prepare.py and train.py thoroughly to understand the codebase
  2. Write autoresearch.md — a living session document recording goal, metrics, files in scope, constraints, and learnings
  3. Write autoresearch.sh — the benchmark script (see Benchmark Script section below)
  4. Commit session files
  5. Run baseline: bash autoresearch.sh
  6. Parse metrics from output (lines matching METRIC name=value)
  7. Record baseline in autoresearch.jsonl:
  • First write a config header: {"type":"config","name":"Optimize val_bpb","metricName":"val_bpb","metricUnit":"bpb","bestDirection":"lower"}
  • Then record the baseline result
  1. Begin the experiment loop

The Experiment Loop

LOOP FOREVER. Never ask "should I continue?" — just keep going.

The user might be asleep, away from the computer, or expects you to work indefinitely. Each experiment takes ~5 minutes, so you can run ~12/hour, ~100 overnight. The loop runs until the user interrupts you, period. If you run out of ideas, think harder — re-read train.py for new angles, try combining previous near-misses, try more radical architectural changes.

Each iteration:

1. Read current git state and autoresearch.md
2. Choose an experimental change to train.py (informed by past results and ASI notes)
3. Edit train.py (the ONLY editable file)
4. git add train.py && git commit -m "experiment: "
5. Run: bash autoresearch.sh > run.log 2>&1
6. Parse METRIC lines from output
7. If output is empty (crash): tail -n 50 run.log to read the stack trace
8. Decide: keep or discard
9. Log result to autoresearch.jsonl (include ASI annotations)
10. If discard/crash: git revert $(git rev-parse HEAD) --no-edit
11. Update autoresearch.md with learnings (every few experiments)
12. Repeat

Decision Rules

  • val_bpb improved (lower)keep (commit stays, branch advances)
  • val_bpb equal or worsediscard (run git revert $(git rev-parse HEAD) --no-edit)
  • Crash (OOM, CUDA error, NaN loss)discard (revert). If it's a simple fix (typo, import), fix and re-run. If the idea is fundamentally broken, log as crash and move on.
  • Simpler code for equal val_bpbkeep (removing complexity is a win)
  • Catastrophic VRAM increase → consider discard even if val_bpb improved slightly

Simplicity Criterion

All else being equal, simpler is better. A 0.001 valbpb improvement that adds 20 lines of hacky code? Probably not worth it. A 0.001 improvement from deleting code? Definitely keep. Equal valbpb with much simpler code? Keep.

Constraints

  • Fixed 5-minute time budget. All experiments are directly comparable — the wall clock is the equalizer.
  • Single file modification. Only train.py changes; prepare.py is immutable. This ensures fair comparison (same data, same evaluation).
  • VRAM is a soft constraint. Using more VRAM is acceptable but note the trade-off (larger model = fewer training steps in 5 minutes).
  • No new packages. You can only use what's already in pyproject.toml.
  • Timeout: If a run exceeds 10 minutes, kill it and treat as a crash.

Don't Thrash

If 3 consecutive experiments fail or get discarded, stop and think about why. Re-read train.py for new angles. Try a fundamentally different approach.

Handling User Messages

If the user sends a message while the loop is running: finish the current cycle, address the feedback, then resume immediately — do not wait for permission.

Logging to autoresearch.jsonl

Each experiment appends one JSON line:

{"run":2,"commit":"def5678","metric":0.993,"metrics":{"peak_memory_mb":44200,"mfu_percent":39.8},"status":"keep","description":"increase LR to 0.04","timestamp":1700000000,"segment":0,"confidence":null,"asi":{"hypothesis":"higher LR converges faster","arch_change":"MATRIX_LR 0.03→0.04"}}

Use the shared logging script:

bash ${CLAUDE_SKILL_DIR}/scripts/log-experiment.sh \
  --run 2 \
  --commit "$(git rev-parse --short HEAD)" \
  --metric 0.993 \
  --status keep \
  --description "increase LR to 0.04" \
  --metrics '{"peak_memory_mb":44200,"mfu_percent":39.8}' \
  --segment 0 \
  --asi '{"hypothesis":"higher LR converges faster"}'

Parse metrics from benchmark output:

bash autoresearch.sh 2>&1 | bash ${CLAUDE_SKILL_DIR}/scripts/parse-metrics.sh

Valid statuses: keep, discard, crash, checks_failed

ASI (Actionable Side Information)

ASI is structured annotation per experiment that survives reverts. When code changes are discarded, only the description and ASI remain — the only structured memory of what happened.

Record ASI for every experiment:

{
  "hypothesis": "Deeper model with fewer steps should compress better",
  "arch_change": "DEPTH 8→12, DEVICE_BATCH_SIZE 128→64",
  "result": "val_bpb improved 0.998→0.992, but 2x VRAM",
  "next_action_hint": "Try intermediate DEPTH=10 for better VRAM tradeoff"
}

Resuming After Context Reset

If autoresearch.jsonl and autoresearch.md exist in the working directory:

  1. Read autoresearch.md for full context (goal, metrics, files, constraints, learnings)
  2. Read autoresearch.jsonl to see all past experiments, current best, and ASI annotations
  3. Check git log to verify current branch state matches expected state
  4. If git state is dirty (unclean shutdown), revert uncommitted changes
  5. Resume the loop from where it left off — no re-setup needed
  6. Resume immediately — do not ask "should I continue?"

Confidence Scoring

After 3+ experiments, assess whether improvements are real or noise:

  • Compute the Median Absolute Deviation (MAD) of all metric values as a noise floor
  • Confidence = |best improvement| / MAD
  • ≥2.0× → likely real improvement
  • 1.0–2.0× → marginal, could be noise
  • run.log 2>&1

valbpb=$(grep "^valbpb:" run.log | tail -1 | awk '{print $2}' || echo "0") memory=$(grep "^peakvrammb:" run.log | tail -1 | awk '{print $2}' || echo "0") mfu=$(grep "^mfu_percent:" run.log | tail -1 | awk '{print $2}' || echo "0")

echo "METRIC valbpb=$valbpb" echo "METRIC peakmemorymb=$memory" echo "METRIC mfu_percent=$mfu"


## Session Files

| File | Purpose |
|------|---------|
| `autoresearch.md` | Living session document — goal, metrics, scope, learnings |
| `autoresearch.sh` | Benchmark script — outputs `METRIC name=value` lines |
| `autoresearch.jsonl` | Append-only experiment log with ASI (survives restarts) |

## Additional Resources

- **`references/gpu-training-guide.md`** — Detailed GPU setup, CUDA configuration, OOM troubleshooting, BPB formula, and performance tuning
- **`scripts/parse-metrics.sh`** — Extract METRIC lines from benchmark output
- **`scripts/log-experiment.sh`** — Append experiment results to autoresearch.jsonl
- **`assets/prepare.py`** — Data preparation (download, tokenizer, dataloader, evaluation)
- **`assets/train.py`** — Model architecture and training loop
- **`assets/program.md`** — Self-contained agent instructions for the ML loop
- **`assets/pyproject.toml`** — Python dependencies (PyTorch, Flash Attention, etc.)

## Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

- **Author:** [proyecto26](https://github.com/proyecto26)
- **Source:** [proyecto26/autoresearch-ai-plugin](https://github.com/proyecto26/autoresearch-ai-plugin)
- **License:** MIT

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet — be the first.

Versions

  • v0.1.0 Imported from the upstream source.