AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
MCP verified MIT Self-run

Ml Researcher Os

mcp-maximussthegreat-ml-researcher-os · by maximussthegreat

Agent skills and workflows for reproducible ML research.

No reviews yet
0 installs
21 views
0.0% view→install

Install

$ agentstack add mcp-maximussthegreat-ml-researcher-os

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access No
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets No
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/mcp-maximussthegreat-ml-researcher-os)

Reliability & compatibility

Security review passed
0 installs to date
no reviews yet
4mo ago

Declared compatibility

Claude CodeClaude DesktopCursorWindsurf

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Ml Researcher Os? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

ml-researcher-os

> Turn your coding agent into a disciplined ML research collaborator.

ml-researcher-os is a skill and workflow pack for developers who use AI coding agents to read papers, design experiments, train models, debug failures, and write research reports.

The goal is not to make an agent sound smart. The goal is to make it behave like a careful junior researcher: skeptical, reproducible, explicit about assumptions, and allergic to fake results.

The 30-second demo

# install path, after the repo is published as a package or cloned locally
gh skill install maximussthegreat/ml-researcher-os

# local setup
python -m pip install -e .
mlro list-skills
mlro init ./research-workspace
mlro extract-claims examples/tiny-paper-replication/paper.md
mlro doctor .
mlro loop init ./research-workspace --metric val_loss --direction lower --budget-minutes 5 --command "python train.py"
mlro audit
mlro improve

Before:

"I implemented the model from the paper."

After:

Hypothesis: the reported gain may depend on sequence length.
Experiment: 3 seeds x 2 baselines x 2 sequence lengths.
Risks: dataset leakage, metric mismatch, weak baseline.
Output: config, training logs, result table, failure notes, report draft.

Why this exists

AI agents can write a lot of code, but ML research fails in quieter places:

  • the paper claim was misunderstood
  • the baseline is too weak
  • the dataset split is leaky
  • the metric is not comparable
  • the agent reports success without checking logs
  • the final writeup hides negative results

ml-researcher-os gives agents a repeatable research protocol instead of a vague prompt.

What will ship first

This repo now includes first drafts of:

  • [paper-claim-extractor](skills/paper-claim-extractor/SKILL.md): extract claims, assumptions, datasets, metrics, and ablations.
  • [experiment-planner](skills/experiment-planner/SKILL.md): convert claims into controlled tests.
  • [training-loop-debugger](skills/training-loop-debugger/SKILL.md): catch common training and evaluation mistakes.
  • [result-reporter](skills/result-reporter/SKILL.md): write tables, limitations, and negative results.
  • [reproducibility-auditor](skills/reproducibility-auditor/SKILL.md): audit rerun readiness.
  • [hf-release-prep](skills/hf-release-prep/SKILL.md): prepare Hugging Face release artifacts.

Repository layout

skills/                  Agent skill specs and prompt contracts
templates/               Experiment, report, model-card, and dataset-card templates
examples/                Small reproducible research demos
docs/                    Architecture, launch notes, and contribution guides
.github/                 Issues, PR templates, and workflows

Demo example

The first demo storyboard is [examples/tiny-paper-replication](examples/tiny-paper-replication). It uses a synthetic paper excerpt so the workflow can be inspected without external data access.

Self-improvement system

This repo is built around a failure-driven improvement loop:

mlro doctor ../my-ml-project --output doctor-report.md

mlro record-failure \
  --title "Agent overclaimed a result" \
  --skill result-reporter \
  --severity high \
  --tag unsupported-claim \
  --observed "The agent claimed reproduction from a stand-in dataset." \
  --expected "The agent should say the claim is only partially supported."

mlro improve --output feedback/IMPROVEMENT_BACKLOG.md
mlro issue-from-failure --failure feedback/failures/example.json
mlro make-regression --failure feedback/failures/example.json
mlro audit

See [docs/self-improvement-loop.md](docs/self-improvement-loop.md) and [docs/repo-doctor.md](docs/repo-doctor.md).

Why people should star it

  • Run mlro doctor on any ML repo and get a reproducibility score in seconds.
  • Run mlro loop to manage fixed-budget agent experiments with keep/reject decisions.
  • Turn vague agent failures into structured JSON with mlro record-failure.
  • Convert failures into GitHub issue drafts with mlro issue-from-failure.
  • Convert failures into regression tasks with mlro make-regression.
  • Keep the project improving from real failures instead of prettier prompts.

Research loop ledger

Inspired by the public momentum around autoresearch-style systems, mlro loop gives agents a tight experiment protocol:

mlro loop init . \
  --metric val_loss \
  --direction lower \
  --budget-minutes 5 \
  --goal "Lower validation loss without changing the data split." \
  --command "python train.py"

mlro loop record . --run baseline --value 1.0
mlro loop record . --run candidate-dropout-0.2 --value 0.91 --changed-file train.py
mlro loop report .

See [docs/research-loop-ledger.md](docs/research-loop-ledger.md) and [docs/landscape-scan-2026-05-23.md](docs/landscape-scan-2026-05-23.md).

Design principles

  1. Reproducibility beats speed.
  2. Every claim needs evidence or a caveat.
  3. Negative results are first-class outputs.
  4. The agent must separate observation, inference, and speculation.
  5. Small examples should run on normal developer hardware.
  6. Every skill change should close or reduce a recorded failure mode.

What this is not

  • Not an autonomous scientist.
  • Not a replacement for domain expertise.
  • Not a benchmark leaderboard for inflated claims.
  • Not a wrapper around one proprietary model.
  • Not a promise that agents can reproduce every paper.

Inspiration and adjacent work

This project is inspired by the broader movement around agent skills, MCP, reproducible ML, and open research tooling. It is designed to complement existing projects, not claim ownership over their ideas.

Useful references to study:

  • Hugging Face ml-intern
  • K-Dense scientific agent skills
  • GitHub agent skills
  • Model Context Protocol
  • Reproducible ML practices from Hugging Face and PyTorch

Roadmap

See [ROADMAP.md](ROADMAP.md).

Contributing

Start with [CONTRIBUTING.md](CONTRIBUTING.md). The best first contributions are:

  • a minimal reproducible ML example
  • a paper-reading checklist
  • a failure case where an agent made a false research claim
  • a benchmark task for agent-skill-bench

License

MIT

Source & license

This open-source MCP server is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.