AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
SKILL verified MIT Self-run

Chat Failure Audit

skill-hiteshbandhu-skills-i-use-chat-failure-audit · by hiteshbandhu

>

No reviews yet
0 installs
36 views
0.0% view→install

Install

$ agentstack add skill-hiteshbandhu-skills-i-use-chat-failure-audit

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access No
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets No
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/skill-hiteshbandhu-skills-i-use-chat-failure-audit)

Reliability & compatibility

Security review passed
0 installs to date
no reviews yet
2mo ago

Declared compatibility

Claude CodeClaude Desktop

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Chat Failure Audit? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

Chat Failure Audit

Turn a real transcript into a ranked, root-caused failure list. The discipline: ground every finding in the actual session, and always push past the symptom to the cause and the layer that owns it.

Works with any agent. No vendor APIs. Reads the transcript/log you provide.

Supporting files (read when needed):

  • [failure-classes.md](failure-classes.md) — the six failure classes, symptom→cause prompts, severity rubric
  • [audit-template.md](audit-template.md) — the failure-mode report format

Output: {SKILL_OUTPUT_DIR}/chat-failure-audit/ — see [../OUTPUT.md](../OUTPUT.md)


Step 0 — Get the real session

Establish the source (ask only if missing):

  1. The transcript — a pasted log, a session file, an exported conversation, or a URL

the agent can fetch. If it's behind auth and unfetchable, ask the user to paste it.

  1. What "success" was — what the user was trying to get done. Failures only mean

something against an intended outcome.

Do not analyse from a summary or from memory — read the actual turns. If you can't see the real session, say so rather than inventing failures.

Resolve output dir per [../OUTPUT.md](../OUTPUT.md). Default ./skill-outputs/chat-failure-audit/.


Step 1 — Walk the turns, mark every breakdown

Read the session turn by turn. Mark each point where the outcome diverged from what the user wanted: a wrong answer, a stall, a re-ask, a tool that failed, a thing the user had to do twice, an abandonment. Capture the observable symptom verbatim — the actual message or behavior, not your paraphrase.


Step 2 — Symptom → root cause → layer

Read [failure-classes.md](failure-classes.md). For each breakdown, push past the symptom:

> Ask "why" until you hit a cause the system owns. "It worked on the next turn" is a > symptom; the cause is "activation happens between turns, not within the turn." "The > model gave a wrong number" may be a grounding failure (bad retrieval), not a model > failure.

Record: symptom (observed) → root cause (the system behavior that produced it) → layer (which part owns the fix: prompt, retrieval, tool/transport, state/lifetime, UX, model choice).


Step 3 — Classify and rate

Tag each with one failure class:

  • Capability gap — the system genuinely can't do the thing yet.
  • Friction — it can, but made the user do avoidable work (→ hand to friction-audit).
  • State / lifetime bug — state was kept at the wrong scope or dropped across a boundary.
  • Grounding / confabulation — answered from nothing / stale / wrong source.
  • Recovery gap — a transient/recoverable error surfaced as a dead-end.
  • UX dead-end — failed or emptied without offering a next step.

Rate frequency (one-off / recurring / every-time) and severity (annoyance / blocks-task / wrong-and-trusted). Wrong-and-trusted outranks everything.


Step 4 — Output to user

Write the report from [audit-template.md](audit-template.md) to the output dir and update index.md. Then in chat:

  1. The failure table: symptom → root cause → layer → class → freq × severity.
  2. The top failures to fix first, by frequency × severity (not by order of appearance).
  3. Explicit handoffs: friction → friction-audit; new/changed state →

state-lifetime-decision; protocol/API mismatch → spec-grounded-design; big fork → architecture-review.

  1. Do not propose code fixes here — this skill diagnoses; design/implementation is a

separate, confirmed step.


Edge cases

  • Only a summary, no raw turns — ask for the real transcript; flag that confidence is

low without it.

  • One symptom, many causes — list each cause separately; they may need different fixes.
  • The "failure" is the user's intent being unclear — that's a real finding (a clarify

gap), not a model failure; classify as UX dead-end or capability gap.

  • Everything blamed on "the model" — push harder; most "model failures" are grounding,

state, or prompt failures the system owns and can fix.

  • Sensitive data in the transcript — never copy secrets/PII into the saved report;

redact to the behavior.


Invocation examples

@chat-failure-audit analyse this chat: 
what went wrong in this session — the user gave up at the end
find the failure modes in this agent run
why did this break? push past "the model was wrong"
audit this transcript and rank what to fix first

Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.