Install
$ agentstack add skill-arasz-ai-badger-ai-raccoon-manual-checklist ✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ✓ Network access No
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
Verified badge
Passed review? Show it. Paste this badge into your README, it links to the public security report.
Reliability & compatibility
Declared compatibility
Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.
We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.
How agent discovery & health will work →About
AiRaccoon manual checklist
The hand-run pass over a live build: the things dotnet test cannot answer because they need a real install, a real server and a real bank.
A checklist is only worth the evidence behind it. The whole design below exists to keep a filled checklist distinguishable from a plausible-sounding one — read "How this checklist rots" before adding, removing or "tidying" anything here, because every failure listed there was found in a real predecessor and each is easy to reintroduce.
When NOT to use
- Anything
dotnet testalready covers. This pass is for behaviour that only appears in a real
install; duplicating unit or integration coverage by hand buys nothing and goes stale.
- Judging a diff or a PR — that is a code review, not a build verification.
- Debugging one failing symptom. Run the checklist to find out what is broken; trace why
somewhere else.
Never
These three protect the person running the checklist, not the checklist:
- Never write to the user's live bank at
~/.ai-raccoon. Read it only through?mode=ro/
PRAGMA query_only=1. The checklist's own writes go to a scratch data root you create with --data-root.
- Never bind the default port (7721) for a step that starts a server — the user's own server
is usually on it, and --restart will cycle theirs. Use --port 0 and read the bound port back.
- Never mark an item accepted from a plan. Accepted means you ran the command and read its
output.
Process
- Copy
templates/checklist-template.jsonto
docs/work/checklist/-.json, creating the directory if needed. Results live under docs/work/checklist/ — never the repo root, and never .ai-raccoon/, which is a bank directory, not a reports directory.
- Derive the facts the checklist compares against, before running anything, and write them
into derived. Run python3 scripts/derive-facts.py (the root defaults to the current directory); it prints the version, the MCP tool count and the prompt count, and it fails loudly — exit 1, with the reason and a standing instruction on stderr — when any of them comes back empty or zero, rather than handing you a confident 0. A count typed from memory or copied from a previous run is a second copy of a number that already exists, and it goes stale silently.
The checklist then asserts that the running binary matches what the tree says. That comparison is the point: it is the only step that can catch a build that installed something other than what you think you are testing.
- Work each item, filling in every field:
command— the exact command or tool call you ran.evidence— the output you read, verbatim, trimmed to the deciding lines.observed-result— what it means.status—pass,fail,skippedorsubstituted. Askippedneeds a reason; it is not
a pass. substituted means the behaviour was checked by a different instrument than the item names — an automated test standing in for a path you could not drive live, say. Folding that into pass overstates it and into skipped understates it, so it gets its own word, and the reason must name the instrument. Record any deliberate deviation from the item's stated method the same way.
accepted+acceptance-reason— whether the observed result is acceptable, and why. A
fail may still be accepted as a known, tracked defect, as long as the reason names where it is tracked.
accepted starts as null, meaning nobody has answered yet. Leave it null until you decide; a run with any null left in it is unfinished, not a run with no objections.
- Never inherit a prior verdict's reason. Where an item was
partial,fail,skippedor
substituted last time, re-derive why against the source in front of you before reusing the explanation. A previous run blamed a job's cadence for three event ids never firing; the real cause was that they log only when there is something to purge, and a fresh bank has nothing. The outcome matched, so the wrong mechanism survived a release — and it discouraged seeding the data that would have exercised them. A reason that is right about the outcome and wrong about the cause is the hardest kind of stale fact to see.
When a run finds that an earlier one was wrong, that belongs in findings-against-prior-runs, not buried in an item. A checklist that can only describe the current build cannot report that the last checklist lied, and those findings are often the most valuable thing a run produces.
- Items whose feature no longer exists get deleted from the template, not marked skipped. A
step for a removed feature is worse than no step: it either fails forever and gets waved through, or it quietly "passes" against nothing.
- Report counts by status, every
failwith its evidence line, and every finding against a prior
run. A run where some item has no command or no evidence is not complete — say so instead of reporting a total.
Items are independent, so lanes can run in parallel against one packed binary. The axis that keeps them from colliding is the data root: give every lane its own --data-root, and nothing else needs coordinating.
Making an item able to fail
Most of these steps can be written so that they pass whether or not the feature works. Three shapes account for nearly all of it:
- A filter needs a negative control. Feed clean content alongside the shapes you expect
rejected. A policy that rejects everything passes a rejection-only test, and looks healthiest exactly when it is most broken.
- A semantic-retrieval query must be one that keyword match cannot carry. Query for
"an antique navigation instrument reflecting evening light in a stargazing room" against content that says "astrolabe", "lamplight", "observatory" — no literal overlap, so a dead vector leg actually fails the item. A query sharing words with the stored text passes on BM25 alone and tells you nothing about the half you meant to test.
- An empty list is a weak check. A queue or candidate list on a fresh bank returns
[]
whether it works or is broken. Create the thing first, then assert on its content — the score, the reasons, the identity — not on the shape of the response.
Scope
Derived per run from the product, not from this list. The headings below are the stable shape of the pass; the specific steps under each come from what step 2 found and from what is actually registered in the build in front of you.
- Build and install — Release build, pack, force-update the global tool, and
--version
matches the derived version.
- Server lifecycle — the server starts and
--restartcycles the one it finds, both on a
non-default port.
- Write path — a write stores and returns a hash; a rejected write says so, with a reason,
rather than returning a fabricated entry.
- Read path — search returns the written entry, get returns its content by hash, and a
file#section anchor resolves its exact chunk.
- Noise filtering — each registered write-path policy rejects what it claims to, and the
rejected content stays retrievable from the noise store. Check which policies are registered before writing steps for them.
- Read-path query guard — the refuse and annotate tiers behave as specified, and any detector
that ships disabled is still disabled until explicitly armed.
- File watch — watch status reflects live registrations.
- Promotion queue — the promotion list reports candidates accurately.
- Full MCP surface — every derived tool and prompt is reachable.
- Observability — emitted event ids resolve against the logging event-id reference.
Anchor each item to the decision record that defines the behaviour, so a step whose ADR was superseded is easy to spot and delete.
How this checklist rots
Two predecessors were deleted after both drifted the same six ways. Each defence below is here because its absence already caused a silent failure:
- Facts pinned by hand. One asserted
--version → 1.9.1and "25 tools" while the tree was at
1.12.0 with 26 tools. The pins had been wrong for three releases and nothing noticed, because the only thing comparing them was a human reading two numbers. → step 2 derives them.
- Steps for a deleted feature. Both still tested a noise policy that had been removed by a
later ADR. A step whose subject does not exist cannot pass honestly. → step 4 deletes them.
- Results written into a bank directory. Reports landed in
.ai-raccoon/, the directory name
the product uses for banks. → step 1 fixes the destination.
- No evidence field. The template recorded a claim with no room for the command or the output
behind it, so a filled checklist and an invented one were indistinguishable afterwards. → command and evidence are required.
- Booleans that cannot express "skipped".
checked/acceptedflags collapsed three states
into two: an unrun item and a failed one both read as false/false. → status is a tri-state and accepted starts null.
- Two drifting copies. Two directories each held a copy, and one had lost its
templates/
directory entirely, so its own step 1 pointed at a file that was not there. → one copy, and the template ships beside the skill.
Retiring a checklist skill means deleting it from every root it can load from, not from the one you were looking at. The deletion that removed the two .ai-badger/skills/learned/ copies missed ~/.hermes/skills/, where a third copy was found still installed long afterwards, still carrying its templates/, still pinning an expected version three minors stale. It was the first hit for someone searching for this checklist, and they began executing it; that copy has since been removed. Enumerate the roots — the project's .ai-badger/skills/learned/, .claude/skills/, ~/.claude/skills/, ~/.hermes/skills/ — and confirm the removal in each. The stale copy wins whoever searches first, so a copy you did not delete is not dormant; it is the one in use.
Gotchas
- A count that cannot fail is not a count. Every counting primitive here has a silent zero
in it: rglob over a directory that no longer exists yields nothing, sum(()) is 0, and grep -c across several files prints one number per file so a bad sum still looks like a number. None of them raise. derive-facts.py closes that by treating zero as a failure rather than a result — it exits 1 and says which fact it could not derive — and by counting only lines that are code. Do not "simplify" either rule away: a file that documents [McpServerTool] in a /// comment already inflated the surface by one for three releases, and the number looked derived the whole time. Both properties have tests in the framework; change the script there, never by editing your copy.
--port 0binds an ephemeral port, so the port must be read back from the server's own output.
Assuming a port here is how a checklist step ends up talking to somebody else's server.
- Force-updating a global tool can silently keep the previous build if the pack step failed
earlier in the same run. The version comparison in step 2 is what catches it; do not skip it because the build "looked fine".
- A scratch
--data-rootmust be a path the running user can create. Pointing it inside a
read-only or root-owned directory produces failures that look like product bugs.
Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: Arasz
- Source: Arasz/ai-badger
- License: MIT
- Homepage: https://github.com/Arasz/ai-badger
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet, be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.