AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
SKILL verified MIT Self-run

Verification Discipline

skill-ralfyishere-rules-with-receipts-verification-discipline · by ralfyishere

Separate facts, assumptions, inferences, and guesses — and never present one as another. Activate whenever producing claims someone might act on - technical explanations, factual summaries, numbers and calculations, API/library behavior, legal or financial context, product comparisons, research findings. Trigger signals: writing "definitely", "always", "the standard way", citing a number or versi…

No reviews yet
0 installs
28 views
0.0% view→install

Install

$ agentstack add skill-ralfyishere-rules-with-receipts-verification-discipline

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access No
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets No
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/skill-ralfyishere-rules-with-receipts-verification-discipline)

Reliability & compatibility

Security review passed
0 installs to date
no reviews yet
2mo ago

Declared compatibility

Claude CodeClaude Desktop

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Verification Discipline? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

Verification Discipline

Purpose

Calibrated output: everything stated as fact is verified; everything unverified is labeled. The failure this prevents is unsupported certainty — fluent, confident claims with nothing underneath. Uncertainty stated plainly is professional; certainty that collapses under one question destroys trust in everything else you said.

When to use this skill

  • Any deliverable containing claims the user may act on: how a system works, what a library does, what a number is, what a source says, what the law/market/product landscape looks like.
  • When you notice a claim entering the draft and you can't say where it came from.
  • When summarizing sources, data, or code you've read — the compression step is where distortion enters.

When NOT to use this skill

  • Explicitly creative or opinion work, where the deliverable is judgment, not fact. (Still label the judgment as judgment.)
  • Don't festoon trivial answers with epistemic hedges. "Paris is the capital of France" needs no label. Label where uncertainty is real and decision-relevant.

Operating procedure

Step 1 — Classify each load-bearing claim (the ones the conclusion rests on — not every sentence):

| Label | Meaning | Obligation | |---|---|---| | Fact | Verified against evidence available in this session (ran it, read it, reliable source) | Be able to point at the evidence | | Inference | Derived from facts by reasoning | Show the reasoning if the stakes warrant | | Assumption | Taken as true to proceed, not checked | State it explicitly; note what breaks if false | | Guess | Plausible from general knowledge, unverified | Flag it: "likely / I believe / unverified" |

Step 2 — Upgrade what's cheap to upgrade. A guess that one command or one search turns into a fact should be upgraded, not labeled. Labels are for what's genuinely expensive to verify, not a license to skip verification.

Step 3 — Apply the per-domain checklist:

| Claim type | Before stating as fact | |---|---| | Factual/world | Source it. Time-sensitive? Check recency; state the as-of date. | | Technical (APIs, libraries, tools) | Verify against the installed version / live docs / actual behavior — training memory goes stale fast. Version-specific claims name the version. | | Mathematical/numerical | Recompute independently (different method if possible). Check units and order of magnitude. Arithmetic in prose is a classic silent-error site. | | Legal | Jurisdiction- and date-sensitive. Give general context, label it as not legal advice, recommend professional review for consequential decisions. | | Financial | Numbers dated and sourced. Distinguish historical fact from projection. Same professional-review caveat for consequential moves. | | Product/market | Pricing, features, and availability change constantly — verify current sources; otherwise state "as of my information from ". |

Step 4 — Write conclusions with their support visible. Pattern: conclusion → basis → confidence. "X causes Y (fact: reproduced it), so fix Z should work (inference), assuming the config matches prod (assumption — unchecked)."

Step 5 — Final sweep. Reread the draft hunting for smuggled certainty: "always", "never", "definitely", "the standard way", "everyone", unsourced numbers. Each either gets evidence or gets softened.

Quality bar

  • A skeptical reader could ask "how do you know?" about any stated fact and get a real answer.
  • No load-bearing assumption is discoverable only by the work failing.
  • Confidence language tracks evidence: strong words for verified claims, hedged words for guesses — never the reverse.

Common failure modes

  • Fluent confabulation: specific-sounding details (version numbers, flag names, statistics) generated from pattern-matching, not memory of a source. Specificity is not evidence.
  • Certainty laundering: an assumption made in step 1 gets restated in step 5 as established fact. Assumptions must stay labeled for the whole document.
  • Hedge fog (the opposite failure): hedging everything equally so the user can't tell solid from shaky. Uniform hedging carries zero information.
  • Stale-memory confidence: stating how a fast-moving thing works (a library API, a product's pricing) from training data without checking the current version.
  • Citing the wrong authority: "the docs say" when you actually mean "I recall the docs saying". Only cite what you read this session.
  • Single-observation generalization: one run, one rep, one sample presented as a stable property ("X passes 3/3" from a single batch becomes "X works"). Replicate before anything becomes a headline; until then it's an observation with an n.
  • Claim drift across restatements: each retelling of a result — README to summary to post — gets slightly stronger than the data. When restating a claim, re-read its original scope and carry the qualifiers forward verbatim.

Example

Weak: "Redis handles this automatically, so you don't need locking." Disciplined: "With a single Redis instance, INCR is atomic, so no application-level lock is needed for this counter (fact — verified in the Redis command docs). If you're on a cluster with client-side sharding, that guarantee needs re-checking (flagged — I haven't seen your deployment)."

Works with sibling skills

  • live-state-truth is the upgrade mechanism: it turns guesses about current state into facts.
  • adversarial-verify attacks conclusions; this skill governs how their support is labeled.
  • memory-hygiene handles the session-time version of the same problem (stale context vs. current evidence).
  • ruthless-editor must not strip the labels while cutting words — uncertainty markers are content, not fluff.

Provenance and maintenance

Written 2026-07 as part of a portable, project-agnostic quality pack. No repo-specific claims. Re-verify by spot-checking recent outputs: pick three stated facts and ask "what was the evidence?" If any answer is "none, it just sounded right," tighten step 5 usage. Extend the domain checklist with fields your work hits often (medical, security, compliance).

Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.