Install
$ agentstack add skill-geledek-enterprise-ai-transformation-skills-process-pilot-design ✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ✓ Network access No
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
Verified badge
Passed review? Show it. Paste this badge into your README, it links to the public security report.
Reliability & compatibility
Declared compatibility
Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.
We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.
How agent discovery & health will work →About
Process — AI Pilot Design
Design a 90-day AI pilot with pre-deployment metrics, stop conditions, and workflow redesign baked in. Seven questions. One-page output ready for sponsor approval.
The discipline: narrow scope, pre-deployment metrics, workflow redesign before technology. AI bolted onto a legacy workflow is the #1 process failure mode (Stanford AI Index 2026 — consult stanford-51-deployments.md).
Output contract (stable): a one-page plan with Day 30 / Day 60 / Day 90 gates, pre-deployment metrics, stop conditions, and a pass/fail verdict criterion.
Role 1: Scope and Problem Framing
Answer the two foundational questions before any technology decisions.
SCOPE CHECK — Scope × Execution Grid (consult 95-5-genai-divide.md):
- Narrow scope + simple execution = fast wins (target this quadrant)
- Narrow scope + complex execution = early pilots (proceed with caution)
- Broad scope = partial pilots or failure (justify explicitly if broad scope is proposed)
Classify the proposed pilot:
- What is the scope? (One function, one workflow, one user group)
- What is the execution complexity? (Simple / Complex)
- Is this upper-left on the grid?
PROBLEM STATEMENT: State the friction this pilot addresses in one sentence. This should be friction-first ("users struggle with X because of Y"), not tech-first ("we want to test AI for Z").
Output: SCOPE | COMPLEXITY | GRID POSITION | FRICTION STATEMENT
Role 2: Pre-Deployment Metrics
Name the metrics BEFORE the pilot launches. If you can't name them now, the pilot isn't ready.
THE MEASUREMENT-GAP WARNING (consult european-fintech-case.md): A leading European fintech replaced 700 agents with one AI system. Tracked volume, response time, cost. Did NOT track resolution quality, repeat-contact rate, CSAT. CSAT dropped 22%. Hiring resumed within months.
Name three metrics that prove this pilot is working:
- [Metric 1 — what it measures, how it will be tracked]
- [Metric 2 — what it measures, how it will be tracked]
- [Metric 3 — what it measures, how it will be tracked]
For each metric, answer:
- Would this naturally get tracked, or does it require additional instrumentation?
- Is this the metric that matters, or the metric that's easy?
Flag any gap between what matters and what will actually be measured.
Output: METRIC 1 | METRIC 2 | METRIC 3 | MEASUREMENT GAP FLAG
Role 3: Blast Radius Assessment
If this pilot fails on day one, who sees it?
BLAST RADIUS OPTIONS (from most to least contained):
- Internal team only — developer and data team see the failure
- Internal users — employees using the tool see the failure
- Paying customers — customer experience degrades
- Regulators — compliance failure triggers regulatory attention
- Public — reputational damage
State the worst plausible blast radius. Set the reliability bar accordingly:
- Internal only → higher tolerance for failure, faster iteration acceptable
- Customers or regulators → human-in-loop required; stop conditions must be tighter
Output: BLAST RADIUS | RELIABILITY BAR | HUMAN-IN-LOOP REQUIREMENT
Role 4: Stop Conditions
Define what causes the pilot to pause for diagnosis at day 30, 60, and 90.
A pilot without stop conditions drifts to perpetual "almost ready." Stop conditions are forcing functions for metric discipline.
For each milestone:
- Day 30: What signal must be present to continue? (Example: adoption rate ≥X%, no critical error events)
- Day 60: What signal must be present to continue to production evaluation? (Example: Metric 1 ≥Y, Metric 3 not deteriorating)
- Day 90: What is the pass/fail verdict criterion? (Example: Metric 2 ≥Z AND no customer-facing errors)
State: if Day 30 stop condition is not met, what happens? (Pause + diagnose / scope reduction / kill)
Output: DAY 30 CONDITION | DAY 60 CONDITION | DAY 90 VERDICT CRITERION | FAILURE RESPONSE
Role 5: Workflow Redesign
AI bolted onto a legacy workflow produces marginal benefit at best. Redesign the workflow before deploying (consult mckinsey-workflow-redesign.md).
Answer the five redesign questions:
- What is this workflow trying to accomplish?
- Which steps exist only because of constraints that AI removes?
- Which steps add value and must be kept?
- Which steps can be eliminated entirely?
- What does the redesigned workflow look like?
Sketch the redesigned workflow (before AI touches the legacy version). This is the target state. AI deployment is the path to this state — not an addition to the current state.
Output: ELIMINATED STEPS | PRESERVED STEPS | REDESIGNED WORKFLOW DESCRIPTION
Role 6: Ownership and Sponsorship
A pilot without a named individual owner is a committee experiment. Committees don't make the decisions that matter when the pilot hits friction.
Name:
- Pilot owner: one person accountable for outcomes (not a team)
- Sponsor: the executive who has committed resources and will review Day 30/60/90 checkpoints
- User champion: the functional lead who will drive adoption within the affected team
Output: PILOT OWNER | SPONSOR | USER CHAMPION
Role 7: Pilot Brief (Synthesis)
Produce a one-page pilot brief from Roles 1–6, ready for sponsor sign-off.
PILOT BRIEF
Problem: [One-sentence friction statement from Role 1] Scope: [Narrow scope + grid position from Role 1] Workflow: [Redesigned workflow from Role 5] Metrics:
- [Metric 1]
- [Metric 2]
- [Metric 3]
Blast radius: [From Role 3] Stop conditions:
- Day 30: [Condition]
- Day 60: [Condition]
- Day 90: [Verdict criterion]
Ownership: [Owner / Sponsor / User champion from Role 6] Duration: 90 days Budget ask: [If known; else "TBD pending sponsor meeting"]
Recommendation: [Proceed / Proceed with modification / Do not proceed — one sentence]
References
All files below live in references/ at the plugin root (${CLAUDE_PLUGIN_ROOT}/references/ when installed as a plugin).
95-5-genai-divide.md— Scope × Execution Grid; 90-day timeline; narrow-scope success patternsstanford-51-deployments.md— invisible-costs; workflow-redesign-first; pre-deployment metrics disciplineeuropean-fintech-case.md— measurement-gap reference casepilot-discipline-ng.md— cross-source pilot discipline synthesismckinsey-workflow-redesign.md— 3× workflow redesign; AI-bolted-onto-legacy failure mode
Reference files are bundled with this skill — Claude resolves them by filename regardless of install layout (single-skill or plugin).
Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: geledek
- Source: geledek/enterprise-ai-transformation-skills
- License: MIT
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet, be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.