AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
SKILL verified MIT Self-run

Ml Research Lab

skill-anastasiyaw-codex-claude-code-config-ml-research-lab · by AnastasiyaW

Machine-learning research loop for dataset curation, fine-tuning, evaluation, inference deployment, experiment tracking, and model explainability. Use when working on ML experiments, training data, model benchmarks, RunPod/GPU runs, classifier quality, vLLM/GGUF serving, SHAP-style model explanations, or research-to-code iterations. Do not use for a simple code edit that has no ML dataset, metric…

No reviews yet
0 installs
7 views
0.0% view→install

Install

$ agentstack add skill-anastasiyaw-codex-claude-code-config-ml-research-lab

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access No
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets No
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/skill-anastasiyaw-codex-claude-code-config-ml-research-lab)

Reliability & compatibility

Security review passed
0 installs to date
no reviews yet
14d ago

Declared compatibility

Claude CodeClaude Desktop

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Ml Research Lab? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

ML Research Lab

Use this skill as the compact router for ML work. It is derived from an audit of synthetic-sciences/openscience at commit 531467c, but does not require running OpenScience or loading its full 250+ skill set.

Operating Loop

  1. Freeze the question as a measurable hypothesis.
  2. Identify dataset provenance, labels, splits, leakage risks, and regeneration cost.
  3. Pick the smallest baseline that can disprove the idea.
  4. Define metrics before training. For release claims, require train/val/test split,

no test-set model selection, and multi-seed proof when cost permits.

  1. Run or wire experiment tracking before long jobs start.
  2. Save artifacts: config, command, data manifest, metrics JSON/CSV, logs, model hash,

and a short conclusion.

  1. Compare against baseline, then keep/discard the change from evidence.

Domain Routing

  • Dataset or scrape cleanup: start from data quality, deduplication, leakage checks,

train/eval splits, and regeneration notes.

  • Classical classifier or tabular baseline: use scikit-learn-style pipelines with

preprocessing inside the pipeline and stratified splits for classification.

  • Model debugging or trust: add SHAP/explainability for feature importance, leakage,

bias/proxy features, and misclassified samples.

  • LLM fine-tuning: prefer JSONL chat format, data validation, LoRA/QLoRA baseline,

and tracked runs before scaling.

  • Single-GPU fast LoRA/QLoRA: consider Unsloth only after checking hardware, CUDA,

model support, and export target.

  • Large or production inference: use vLLM for high-throughput GPU serving, GGUF or

llama.cpp for local/Apple/CPU-friendly deployment, and TensorRT-LLM only when the NVIDIA production optimization cost is justified.

  • Research write-up: report method, dataset, exact metric formula, baseline source,

limitations, and failure cases.

Verification Gates

  • Data gate: schema valid, duplicates/leakage checked, split manifest saved.
  • Metric gate: exact metric formula named; if benchmarked, original baseline source

and benchmark code checked.

  • Runtime gate: command/log path and environment captured; GPU memory and errors

checked for long runs.

  • Tracking gate: metrics are retrievable as JSON/CSV or a dashboard link plus local

export.

  • Deployment gate: latency, throughput, memory, and OOM behavior measured before

claiming production readiness.

Adoption Boundary

Do not import broad external skill collections wholesale. Use the inventory script scripts/openscience_skill_inventory.py to rank candidates, inspect the relevant source skill manually, then promote only compact workflows or deterministic scripts that improve our own tests.

Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.