Install
$ agentstack add skill-crewforth-crewforth-reflect ✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ✓ Network access No
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
Verified badge
Passed review? Show it. Paste this badge into your README, it links to the public security report.
Reliability & compatibility
Declared compatibility
Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.
We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.
How agent discovery & health will work →About
Reflect — step back and audit the work, not just the code
Trigger phrases: "reflect", "retro", "retrospective", "what did we miss", "step back", "introspect"
When
A nontrivial chunk of work just finished (a feature, a plan, a debugging session) and it's worth a deliberate step back before committing or moving on. This is the meta-cognitive counterpart to [[iterate]]: iterate drives a change to its exit test; reflect asks whether the exit test — and the approach behind it — was even the right one. Skip it for a one-line, unambiguous change; there's nothing to reflect on.
Measure first, then remember
Run bash .claude/hooks/session-stats.sh before answering anything below, and open the pass with what it reports. A retro built only on recall is an interview with the least reliable witness in the room: the model reconstructs a tidy story from a context that has already been summarised, and the stretches where it span on a failing approach are exactly the ones it remembers least. The script counts what actually happened — prompts, tool calls, failing loops, near-duplicate prompts, interrupts, auto-compactions.
Treat each ⚠️ as a question to answer in the pass, not a verdict: a runaway loop asks what assumption kept failing; a repeated prompt asks what context never landed; an interrupt asks where intent diverged; an auto compaction asks what state was silently dropped. If the numbers and your recollection disagree, the numbers are the record. If the script is missing (a plugin install has no .claude/hooks/), say the retro is recall-based — do not present recall as measurement.
The pass
Ask each question honestly and write the answer, not a reassurance:
- Unverified assumptions — what did I take for granted that I never checked? Name each one and whether it
was actually confirmed (read the code / ran the flow) or just assumed.
- What got skipped — an edge case, an error path, a test, a doc update, a security/privacy angle. What was
silently dropped, and was that a conscious trade-off or an oversight?
- Right approach? — with hindsight, is this the direction we'd still pick? Did scope creep in? Is there a
simpler path we walked past ([[code-review]] altitude: could 200 lines be 50)?
- Evidence gap — which claims of "done" / "works" rest on having observed behavior vs. on inference?
An unobserved "it works" is a finding (see [[iterate]] / verify: drive the real flow).
- What I'd tell the next session — the one thing a fresh context most needs to know (feeds [[handoff]]).
Guardrails
- Findings, not code. Reflect surfaces gaps; the relevant specialist fixes them (a found gap re-enters
[[iterate]] or the author agent). It changes nothing itself.
- Honest, not performative. A retro that only confirms good work is a failed retro. If nothing is found,
say specifically why you're confident (what was verified), don't just assert it.
- Bounded. One pass, concrete findings, then act or close — not an open rumination loop.
- Token discipline ([[token-budget]]): a short ranked findings list to the main thread; detail to a file.
DoD (this skill's contribution)
session-stats.shwas run and every ⚠️ it raised is answered — or its absence is stated.- Assumptions are labeled verified vs. assumed; unverified ones are flagged, not buried.
- Skipped items and any scope creep are named explicitly.
- Every "done/works" claim is traced to observed evidence or marked as inference.
- Findings are actionable (each maps to a fix, a follow-up, or an accepted trade-off).
Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: crewforth
- Source: crewforth/crewforth
- License: MIT
- Homepage: https://crewforth.com/
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet, be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.