Install
$ agentstack add skill-neuroanalytics-data-science-harness-datalad-init ✓ scanned · ✓ verified — works with Claude Code, Cursor, and more.
Verified badge
Passed review? Show it. Paste this badge into your README — it links to the public security report.
Reliability & compatibility
Declared compatibility
Compatibility is declared by the source manifest. End-to-end runtime verification is coming — see below.
We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps — measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.
How agent discovery & health will work →About
Skill: datalad-init
Create a new DataLad dataset following YODA principles for reproducible, provenance-tracked local analysis projects. The -c yoda configuration is always applied — it sets up the correct directory layout and .gitattributes rules.
Steps
- Determine target path — read
$ARGUMENTS. If a path is given, use it. If empty,
use the current working directory. Show the resolved absolute path and ask the user to confirm before proceeding.
- Pre-flight checks — before running any command:
- Check for an existing
.datalad/directory at the target path. If found, stop:
> "A DataLad dataset already exists at `. Re-initializing would overwrite > dataset configuration. Run datalad status` to check its state."
- Check whether
.git/annex/already exists at the target path:
``bash git annex config annex.backend 2>/dev/null ` If it returns MD5E`, warn: > "The existing annex uses the legacy MD5E backend. Migrating backends after the > fact is complex. Consider initializing a fresh dataset at a new path and > re-ingesting data." Stop unless the user explicitly asks to continue anyway.
- Check whether the target path is inside a plain git repository (not a datalad dataset):
run git rev-parse --git-dir from the target path. If a .git/ is found but no .datalad/, warn: > "Target path is inside a plain git repository. Creating a DataLad dataset here > will nest it inside an untracked git repo, which can cause unexpected behavior. > Consider initializing at the git repo root, or outside the git tree. Proceed anyway?" Wait for user confirmation before continuing.
- Create the dataset — run:
`` datalad create -c yoda ` Show the command before executing. Report the full output. Optionally suggest adding --annex-backend SHA256E (recommended; avoids the legacy MD5E backend which causes compatibility issues with some special remotes): ` datalad create -c yoda --annex-backend SHA256E ``
- Report the YODA structure — after creation, display what was created:
`` / ├── code/ ← version-controlled scripts and notebooks (annexed: never) ├── outputs/ ← results produced by datalad run (annexed: everything) ├── inputs/ ← input data, ideally linked as subdatasets (annexed: everything) └── README.md ← dataset description ` Explain each directory's role briefly (see ${CLAUDEPLUGINROOT}/../references/yoda-layout.md` for full detail).
- Explain YODA principles — briefly state the three principles:
- P1 — Everything is a dataset: input data should be linked as subdatasets, not copied
- P2 — Record data origins: use
datalad download-urlordatalad clonewith provenance - P3 — Never modify a dataset you didn't create: work only in
outputs/andcode/
- Suggest next steps — close with:
- Put scripts and analysis code in
code/ - Link input data as a subdataset:
datalad clone inputs/ - Use
datalad runto execute scripts so commands and outputs are recorded - Use
datalad save -m "..."to record code changes
Constraints
- Never run
datalad createwithout explicit path confirmation from the user. - Never re-initialize an existing DataLad dataset — always stop and explain.
- Always use
-c yoda— never create a bare dataset without the YODA configuration. - Always warn when the target path is inside a plain git repo, and require explicit
confirmation before proceeding.
- Load
${CLAUDE_PLUGIN_ROOT}/../references/yoda-layout.mdif the user asks for more detail
about YODA conventions, .gitattributes behavior, or subdataset patterns.
Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: neuroanalytics
- Source: neuroanalytics/data-science-harness
- License: MIT
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet — be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.