Install
$ agentstack add skill-aws-samples-sample-aws-ops-skills-for-agents-aws-troubleshooting ✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ✓ Network access No
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
Verified badge
Passed review? Show it. Paste this badge into your README, it links to the public security report.
Reliability & compatibility
Declared compatibility
Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.
We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.
How agent discovery & health will work →About
AWS Troubleshooting
One skill for all AWS services. Reasoning framework is here; service-specific facts come from MCP at runtime.
When to use
Any AWS issue where the console alone is insufficient — status checks, logs, connectivity, performance, permissions, deployments, state transitions.
Investigation workflow
Step 0 — ASK FIRST (before any diagnosis)
MUST:
- Check if critical context is missing before querying or diagnosing:
- Network issues → need VPC topology, subnet type, NAT/endpoint config
- Container/task failures → need exit code, error message, logs
- Connection timeouts → need source, destination, protocol, port
- Performance issues → need instance type, workload pattern, timeline
- If missing, ASK the user. A wrong diagnosis from assumptions wastes more
time than a clarifying question.
Step 1 — Identify and collect
Determine the affected service and resource, then collect initial evidence.
MUST:
- Identify the service, resource ID, and region
- Run the service's
describe-*/get-*APIs to capture current state - Check for status checks, health checks, or equivalent (service-dependent)
- Check CloudWatch metrics for anomalies in the relevant namespace
- Read
references/stable-guardrails.mdto avoid known misdiagnosis traps
SHOULD:
- Check CloudTrail for recent API calls that may have caused the issue
- Check AWS Health Dashboard for service-level events
- Collect logs (CloudWatch Logs, system logs, console output) if available
Step 2 — Query real-time documentation
MUST:
- Use
aws-knowledgeMCPsearch_documentationto find current troubleshooting
guidance for the specific symptom. See references/mcp-query-patterns.md
- If the search returns an SOP (
sop_namefield), retrieve it with
retrieve_agent_sop for step-by-step instructions
- For ANY specific number (IOPS, limits, quotas, timeouts, cooldowns):
query MCP — NEVER rely on memorized values
SHOULD:
- Cross-reference re:Post Knowledge Center articles for the error message
- Check if the service has SSM Automation runbooks (
AWSSupport-Troubleshoot*)
that can automate diagnosis
MAY:
- Use
aws-knowledgerecommendtool on a relevant doc page to discover
related troubleshooting content
Step 3 — Diagnose
MUST:
- Read
references/hallucination-patterns.yamlbefore concluding - State the root cause with specific evidence (API response, metric value, log excerpt)
- Classify severity: CRITICAL (service down) / HIGH (degraded) / MEDIUM (suboptimal)
SHOULD:
- Check blast radius — is only one resource affected, or is it AZ/region-wide?
- Distinguish between AWS-side issues (status checks, service events) and
customer-side issues (config, permissions, application)
Step 4 — Remediate and report
MUST:
- Propose immediate mitigation with specific CLI commands
- Propose long-term prevention (alarms, auto-recovery, architecture changes)
- Output structured YAML report (see Output Format below)
SHOULD:
- Verify the fix worked (re-check status/metrics after remediation)
Output format
service: ""
resource: ""
region: ""
root_cause: " — "
evidence:
- type:
content: ""
severity: CRITICAL | HIGH | MEDIUM
blast_radius: ""
mitigation:
immediate: ""
long_term: ""
sources:
- ""
Anti-hallucination rules
- NEVER state service-specific numbers (IOPS, limits, quotas, defaults) from
memory. Always query MCP first.
- Always cite evidence: API response, metric, log excerpt, or MCP doc URL.
- Read
references/hallucination-patterns.yaml— these are patterns where
LLMs consistently get AWS behavior wrong.
- Read
references/stable-guardrails.md— these are architectural facts that
are safe to assert without querying.
- Spend no more than 2 minutes on any single hypothesis. Pivot if inconclusive.
- If MCP returns no relevant results, say so explicitly. Do not fabricate guidance.
References
| File | Purpose | |------|---------| | references/stable-guardrails.md | Architectural facts that don't change — safe to assert | | references/hallucination-patterns.yaml | Cross-service LLM mistake patterns | | references/mcp-query-patterns.md | How to query aws-knowledge MCP effectively | | references/investigation-framework.md | Detailed Phase 1/2/3 methodology for complex cases |
Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: aws-samples
- Source: aws-samples/sample-aws-ops-skills-for-agents
- License: MIT-0
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet, be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.