AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
SKILL verified MIT Self-run

Dataset Onboard

skill-zakelfassi-skills-driven-development-dataset-onboard · by zakelfassi

Onboard a new raw dataset into the data platform — sniff the schema, generate a profiling notebook, create the ingestion job, and add an entry to the data dictionary. Use when adding a new data source, when a partner delivers a new CSV/Parquet drop, or when asked to "onboard the {name} dataset".

No reviews yet
0 installs
30 views
0.0% view→install

Install

$ agentstack add skill-zakelfassi-skills-driven-development-dataset-onboard

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access No
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets No
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/skill-zakelfassi-skills-driven-development-dataset-onboard)

Reliability & compatibility

Security review passed
0 installs to date
no reviews yet
3mo ago

Declared compatibility

Claude CodeClaude Desktop

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Dataset Onboard? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

Dataset Onboard

Bring a new raw dataset into the platform with schema documentation, profiling, and an ingestion job.

Inputs

  • Dataset name (snake_case, e.g., customer_events)
  • Source format (csv, parquet, json, avro)
  • Source path or URI (S3, GCS, local mount, API endpoint)
  • Expected frequency (daily, weekly, on-demand)
  • Owner team and contact

Steps

  1. Schema sniff

``python import pandas as pd df = pd.read_csv("{source_path}", nrows=1000) # or read_parquet, etc. print(df.dtypes) print(df.describe(include="all")) print(df.isnull().sum() / len(df)) # null rates `` Document:

  • Column names, inferred types, null rates, example values
  • Detected anomalies (mixed types, encoding issues, unexpected nulls)
  1. Create the data dictionary entry

Edit docs/data-dictionary/{dataset_name}.md: ```markdown # {DatasetName}

Owner: {team} Contact: {email} Source: {uri} Frequency: {frequency}

| Column | Type | Nullable | Description | |--------|------|----------|-------------| | ... | ... | ... | ... | ```

  1. Generate the profiling notebook

``bash cp templates/profiling-notebook.ipynb \ notebooks/profiling/{dataset_name}_profile.ipynb `` Edit the notebook to use the correct source path and column list. Run it to confirm it completes without errors.

  1. Create the ingestion job

`` pipelines/ingestion/{dataset_name}/ ├── ingest.py # main ingestion script ├── schema.py # column definitions and type coercion ├── config.yaml # source path, schedule, destination table └── tests/ └── test_ingest.py # unit test with a small fixture file `` The ingestion script must be idempotent (re-running on the same input produces the same output; no duplicate rows).

  1. Register in the scheduler

Add a DAG entry in dags/{dataset_name}_ingest.py (Airflow) or a flow in flows/ (Prefect). Set the schedule to match {frequency}.

  1. Test the ingestion

``bash python -m pytest pipelines/ingestion/{dataset_name}/tests/ -v python pipelines/ingestion/{dataset_name}/ingest.py --dry-run ``

  1. Run a data quality gate (invoke data-quality-gate skill)

Add at minimum: null checks for required columns, row-count sanity check.

Conventions

  • Raw datasets land in the raw/ schema/layer; never write to staging/ or marts/ from an ingestion job
  • All ingestion jobs accept --dry-run and --date flags
  • Dataset names are snake_case; tables follow the same naming
  • Profiling notebooks live in notebooks/profiling/; they are committed

Edge Cases

  • Malformed source file: Log the error with row number, skip the bad row, emit a data_quality_alert metric. Never silently swallow rows.
  • Schema drift (columns added/removed): The ingestion job must compare the inbound schema to schema.py and fail fast on unexpected changes, rather than silently loading partial data.
  • Large datasets (>1GB): Use chunked reads (chunksize in pandas, or native Parquet partitioning); test with a 10k-row sample first.
  • API source with rate limits: Add exponential back-off and a --resume-from flag that uses a checkpoint file.

Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.