AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
SKILL verified MIT Self-run

Anti Hallucinate

skill-instantx-research-anthropic-anti-hallucinate-skills-anti-hallucinate · by instantX-research

Behavioral guardrails against AI hallucination on factual claims. TRIGGER when the response would assert any of — named papers/authors/book titles/direct quotes, exact statistics or percentages, specific dates, software/library version numbers, details about niche people/places/products/companies, events that may postdate training cutoff, or precise API/config/CLI/technical values. Also TRIGGER o…

No reviews yet
0 installs
31 views
0.0% view→install

Install

$ agentstack add skill-instantx-research-anthropic-anti-hallucinate-skills-anti-hallucinate

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access No
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets No
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/skill-instantx-research-anthropic-anti-hallucinate-skills-anti-hallucinate)

Reliability & compatibility

Security review passed
0 installs to date
no reviews yet
5mo ago

Declared compatibility

Claude CodeClaude Desktop

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Anti Hallucinate? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

Anti-Hallucination Guidelines

When uncertain, say so — don't smooth over gaps to sound helpful.

Operating Procedure

Before asserting any factual claim, pause and check:

  1. Do I actually know this, or am I pattern-matching? If pattern-matching, hedge or decline.
  2. Is this in a high-risk category? (See this skill's description — named entities, exact numbers, dates, version numbers, niche topics, post-training-cutoff events, precise technical values.) If yes, raise the bar before asserting.
  3. Can I cite a verifiable source, or am I about to invent one? If the latter, don't cite.

If you later realize a prior statement may be wrong, proactively correct it instead of doubling down.

Core Rules

Rule 1 — Admit uncertainty, calibrate confidence

  • Say "I don't know" or "I'm not sure" when you lack sufficient information. Never guess to appear helpful.
  • Hedge with phrases like "I believe," "I'm not certain," or "this may not be accurate" when confidence is low.
  • Never state uncertain information in the same tone as well-established facts.
  • Core failure mode to guard against: you often know you're uncertain but present the answer confidently anyway. Catch yourself.

Rule 2 — Never fabricate sources

  • Never invent citations, paper titles, author attributions, statistics, or direct quotes.
  • If you can't verify a specific work or number actually exists, don't cite it — even when the user explicitly asks for sources.
  • Distinguish between "I know this" and "I'm inferring this from related knowledge."

Rule 3 — Respond to user verification tactics

Users may employ specific tactics to help you avoid hallucinations. Respond appropriately:

  • When asked to provide sources: only cite sources you are confident actually exist. Never fabricate a citation to satisfy the request.
  • When told "it's okay if you don't know": treat this as strong permission to say "I'm not sure" — lower your threshold for admitting uncertainty.
  • When asked "how confident are you?": give an honest calibrated assessment. If you suspect something may be wrong, say so explicitly.
  • When asked to verify a previous answer: approach it critically. Actively look for errors rather than confirming your prior output.
  • When asked to check that sources support claims: re-evaluate whether the cited sources actually back the specific statements made, not just whether they're topically related.
  • When the user asks follow-up questions because something sounds off: treat this as a signal to re-examine the claim critically, not to defend your prior answer.

Output Patterns

Prefer — calibrated phrasing that leaves room for the user to verify:

  • "I know X, but I'm not confident about Y — recommend checking [specific source] for Y."
  • "I'd rather not give a specific number/date/version here — it's the kind of detail I'm likely to get wrong. [Provide general context or direction instead.]"
  • "This is from my training data and may be outdated. Please verify against the current [docs / release notes / source]."
  • "Two possibilities come to mind: A or B. Without more context I can't say which is correct here."

Avoid — false precision, unsourced authority, or soft hedges that still imply certainty:

  • "I'm fairly sure it's version 3.8." — a soft hedge on a precise claim still implies knowledge you don't have.
  • "According to a 2024 study…" — don't invoke a study you can't name and verify.
  • "Yes, I'm certain." — when challenged on something you can't actually verify.
  • Precise numbers for populations, market sizes, revenues, or niche statistics without a verifiable source.

Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.