Install
$ agentstack add skill-mitodl-agent-kit-deploy-verification ✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ✓ Network access No
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
Verified badge
Passed review? Show it. Paste this badge into your README, it links to the public security report.
Reliability & compatibility
Declared compatibility
Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.
We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.
How agent discovery & health will work →About
Deploy Verification
A config or manifest change merging is not the same as it taking effect. Three separate things can each silently fail while the previous one succeeds: the CD pipeline can be pending or red, the pods can never actually restart even though the pipeline ran, and the restarted pods can still be running the old value if the change didn't reach the path you think it did. This skill walks all three, in order, and demands evidence at each step rather than inferring success from the step before it.
Core rule: report pass/fail per step with the raw evidence (a pod age, a rollout status, a query result, a preview diff) — "looks deployed" or "should be live now" is not a verification, it's a guess with better formatting. If a step can't be checked (no cluster access, no query result yet), say so explicitly rather than skipping it silently.
Step 1 — Confirm the CD pipeline ran and succeeded
Find the deploy workflow for the merged PR's repo and confirm a run completed after the merge commit landed, not just that one exists:
gh run list -R / --workflow .yml --limit 5 \
--json databaseId,status,conclusion,headSha,createdAt
Match headSha (or the commit it built from) against the actual merge SHA — a run against a stale SHA doesn't cover this change. status: in_progress means wait and re-check, not "probably fine." A conclusion: failure stops here: the change never reached the cluster, and nothing past this step is worth checking yet.
Step 2 — Confirm the pods actually rolled
A green pipeline only means the deploy step ran, not that the workload picked it up. Check pod age against the deploy time:
kubectl get pods -n -l app= -o wide
kubectl rollout status deployment/ -n --timeout=5m
Always bound rollout status with --timeout. Its default is --timeout=0s, which kubectl documents as "zero means never" — an unbounded watch, so a stalled rollout hangs the verification instead of reporting anything. A non-zero exit on timeout is itself a finding: the rollout did not complete within the window, which is exactly the "looks deployed but isn't" case this checklist exists to catch. Report it as a failed step, not as a flaky command to retry with a longer wait.
If pod AGE predates the pipeline run, the workload did not restart. This is the single most common silent failure in this checklist: a ConfigMap/Secret change with no checksum/annotation-based restart trigger on the pod template is inert until something else causes a restart. Check whether the Pulumi/Helm resource actually wires a config-hash annotation (or equivalent) into the pod spec — if it doesn't, the merge changed the manifest but nothing forced a rollout.
Don't run kubectl rollout restart yourself to force it — that's a visible, mutating action on a live workload. Report that pods haven't rolled and ask before restarting anything, unless the user has already asked you to force the rollout as part of this task.
Step 3 — Read the config out of the running process
Confirm the value, not just that the pod is new. kubectl get configmap -o yaml only shows what Kubernetes has stored, not what the running process loaded — the two can differ if the app caches config at startup or reads from a mounted file with its own reload semantics.
kubectl exec -n -- env | grep
# or, for a file-mounted config:
kubectl exec -n -- cat
# or, if the app exposes one, a debug/health endpoint that echoes live config
If the app has no way to introspect its running config, say that's the limitation rather than assuming the new pod means the new value is loaded.
Step 4 — Query the metric the change was supposed to move
Use whichever toolhive-swe tier matches the environment (toolhive-swe-ci, toolhive-swe-qa, toolhive-swe-prod — see this repo's agent-config.toml) to query Prometheus directly rather than eyeballing a dashboard:
mcp__toolhive-swe-__grafana_query_prometheus
Compare a post-deploy window against a prior baseline of at least the same length — a 2-hour post-deploy window against a 2-hour pre-deploy window is a fair comparison; a 2-hour post-deploy window against "it looked fine on the dashboard" is not. For anything with daily/weekly seasonality (request volume, error rate), compare against the same time-of-day/day-of-week a week prior, not just the hours immediately before the deploy. State the exact query and the window in the report — a reader should be able to rerun it.
If the metric hasn't moved the expected direction, or hasn't moved at all, say that plainly. Don't extend the window, change the query, or reach for a different metric to manufacture a positive result — report the mismatch and let the user decide whether it's a slow rollout, a wrong hypothesis, or a genuine regression.
Step 5 — Confirm the change didn't leak into an unintended environment
Infra changes are usually per-stack (CI/QA/Prod); a change meant for one stack landing in another is its own bug. Diff every stack that shares the changed code, not just the target one:
pulumi preview --stack
"0 changes" on every stack except the intended target confirms scope. Any unexpected diff on an unrelated stack is a blocker, not a footnote — surface it before calling the deploy verified.
When to stop and ask instead of deciding
- Pods haven't rolled and the reason isn't obvious — don't force a
restart unprompted (Step 2); report and ask.
- The pipeline is still running — wait or say so; don't infer success
from a merge alone.
- The metric contradicts the expected effect — report it as a mismatch,
not as "probably needs more time," unless you have evidence (e.g. traffic volume) that supports that read.
- An unintended-stack diff shows up in Step 5 — this is a scope bug in
the original change, not something to quietly ignore because the intended stack looks correct.
Environment
Requires kubectl configured with a context for the target cluster, the pulumi CLI with access to the relevant stacks, gh for CD pipeline status, and at least one toolhive-swe-{ci,qa,prod} MCP server registered (see agent-config.toml) for the Grafana/Prometheus query step. Skip a step outright (and say so) rather than guessing when the required access isn't available.
Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: mitodl
- Source: mitodl/agent-kit
- License: BSD-3-Clause
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet, be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.