Install
$ agentstack add skill-viksant-swe-skills-verify-claims ✓ scanned · ✓ verified — works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ✓ Network access No
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
About
Verify Claims (Self-Hallucination Detection)
> Core Philosophy: "If a claim would be 'true' even WITHOUT the cited evidence, the citation provides no information gain. You're asserting beyond what the evidence justifies."
Why This Skill Exists
Claude tends to:
- Assert claims with citations that don't actually support them
- Add "specificity" that sounds authoritative but isn't grounded in evidence
- Pattern-match from training data instead of reasoning from context
- Declare root causes without sufficient evidential backing
- Use citations decoratively rather than informationally
This skill FORCES rigorous verification of every claim against its cited evidence.
Theoretical Basis
This skill applies a simplified EDFL/Strawberry approach to detect claims where citations add no information beyond prior knowledge. For the full mathematical basis, see references/edfl-theory.md.
When to Use This Skill
Mandatory Triggers
- After generating complex reasoning with file:line citations
- Before committing to root cause diagnosis
- When user asks "are you sure?" or "verify your claims"
- After debugging analysis that will drive implementation decisions
- When claiming causal relationships (X causes Y)
Explicit Triggers
- User says: "verify", "check your reasoning", "are you sure", "prove it"
- User asks: "how do you know?", "what's your evidence?"
Auto-Detect Triggers
| Condition | Reason | |-----------|--------| | Analysis has 5+ claims with citations | Complex reasoning needs verification | | Root cause diagnosis made | High stakes, often wrong | | Causal claims ("because", "causes", "leads to") | Easy to hallucinate causation | | Specific numbers/versions cited | Easy to confabulate specifics | | Contradicting previous analysis | May be overcorrecting |
The 5-Phase Protocol (MANDATORY)
Phase 1: Decompose Reasoning into Atomic Claims
Transform your reasoning into structured claims:
The function processData() is defined on line 45
The null pointer exception occurs because userData is not initialized
This pattern suggests the developer intended lazy initialization
Claim Kinds: | Kind | Definition | Citation Requirement | |------|------------|---------------------| | copy | Directly stated in one span | MUST cite exactly one span | | deduce | Follows logically from cited spans | MUST cite supporting spans | | arithmetic | Numeric calculation | Cite source numbers | | assumption | NOT supported by context | cites MUST be empty |
Phase 2: Extract Context Spans
Create labeled spans from relevant context:
def processData(userData):
result = userData.transform()
return result
data = processData(None)
Span Rules:
- Include EXACT code/text, not paraphrases
- Include source location (file:line)
- Include enough context for standalone understanding
- Number sequentially (S0, S1, S2...)
Phase 3: Verify Each Claim (Posterior Check)
For each claim, with FULL context available:
CLAIM TO CHECK:
- idx: 1
- kind: deduce
- cites: S0, S1
- claim: "The null pointer exception occurs because userData is not initialized"
CONTEXT SPANS:
[S0] def processData(userData):
result = userData.transform()
return result
[S1] data = processData(None)
DECISION CRITERIA:
- ENTAILED: Claim follows from spans (logically or explicitly)
- CONTRADICTED: Spans imply the opposite
- NOT_IN_CONTEXT: Claim asserts facts absent from spans
- UNVERIFIABLE: Too vague or ill-formed to judge
VERDICT: [Choose one]
CONFIDENCE: [0.0-1.0]
NOTES: [Brief justification]
Phase 4: Causal Null Check (Prior Check)
For each claim that received ENTAILED in Phase 3, repeat with cited spans SCRUBBED:
CLAIM TO CHECK:
- claim: "The null pointer exception occurs because userData is not initialized"
CONTEXT SPANS:
[S0] [EVIDENCE REMOVED FOR VERIFICATION]
[S1] [EVIDENCE REMOVED FOR VERIFICATION]
QUESTION: Based ONLY on general programming knowledge
(not the removed evidence), would this claim likely be true?
- If YES (ENTAILED): Citation adds NO information - FLAG THIS
- If NO (NOT_IN_CONTEXT): Citation IS informative - PASS
- If UNSURE: Ambiguous - REVIEW MANUALLY
Phase 5: Flag Insufficient Budget
Apply the flagging logic:
def should_flag(post_verdict, prior_verdict, kind):
if kind == "assumption":
return False # Assumptions are explicitly unsupported
if post_verdict != "ENTAILED":
return True # Not even supported with full context
if prior_verdict == "ENTAILED":
return True # Citation adds no information (hallucination risk)
return False
Interpretation Matrix:
| postverdict | priorverdict | Result | Meaning | |--------------|---------------|--------|---------| | ENTAILED | NOTINCONTEXT | PASS | Citation is informative | | ENTAILED | ENTAILED | FLAG | Citation adds nothing (confabulation) | | ENTAILED | UNVERIFIABLE | REVIEW | Ambiguous, manual review | | NOTINCONTEXT | | FLAG | Unsupported even with evidence | | CONTRADICTED | | FLAG | Evidence contradicts claim |
Output Format
## Verification Report
### Summary
| Metric | Value |
|--------|-------|
| Total Claims | 5 |
| Flagged | 2 |
| Grounded Fraction | 60% |
### Results
**Claim 0** [PASS]
- Kind: copy
- Claim: "The function processData() is defined on line 45"
- Cites: S0
- Posterior: ENTAILED
- Prior: NOT_IN_CONTEXT
- Status: Citation informative. Claim properly grounded.
**Claim 1** [FLAG]
- Kind: deduce
- Claim: "The null pointer exception occurs because userData is not initialized"
- Cites: S0, S1
- Posterior: ENTAILED
- Prior: ENTAILED
- Status: BUDGET INSUFFICIENT
- Issue: Claim would be considered true even without the cited evidence.
The citation does not provide information beyond prior knowledge.
- Interpretation:
- Hallucinated specificity (adding details not in evidence)
- Over-confident inference
- Correct claim with wrong/unnecessary citation
**Claim 2** [FLAG]
- Kind: deduce
- Claim: "This is a known issue in version 3.2"
- Cites: S2
- Posterior: NOT_IN_CONTEXT
- Prior: N/A
- Status: UNSUPPORTED - Claim not entailed by cited spans.
### Recommendations
For flagged claims:
1. Find additional evidence that specifically supports the claim
2. Weaken the claim to match available evidence
3. Mark explicitly as assumption/hypothesis
4. Remove the claim from the reasoning chain
Quick Verification Checklist
Before finalizing analysis, answer:
For each claim I made:
[ ] Did I cite specific evidence (file:line)?
[ ] Does the evidence ACTUALLY say what I claimed?
[ ] Would I believe this claim WITHOUT seeing the evidence?
- If YES: My citation is decorative, not informative
- If NO: Good, my reasoning depends on the evidence
[ ] Am I inferring beyond what the text literally says?
[ ] Am I pattern-matching from training data?
Common Failure Patterns
For detailed examples of decorative citations, specificity hallucination, causal leaps, and training data bleed, see references/failure-patterns.md.
Integration with Workflow
This skill runs:
- AFTER complex analysis or debugging
- BEFORE presenting conclusions to user
- BEFORE code implementation based on analysis
Order:
- Generate analysis with citations
- verify-claims ← YOU ARE HERE (are claims GROUNDED?)
- Present verified conclusions
- Implement (if applicable)
Self-Application Example
For a complete worked example of verifying claims before finalizing analysis, see references/self-application-example.md.
Limitations (Be Honest About These)
- Self-verification bias: Claude checking Claude has inherent circularity
- No true probabilities: Qualitative verdicts vs Strawberry's KL divergence
- Context cost: Doubles reasoning about each claim
- No logprobs: Cannot calculate exact information contribution
Mitigations:
- Conservative flagging (when uncertain, flag for human review)
- Reformulation check (same claim, different phrasing)
- Human-in-the-loop for flagged claims
Troubleshooting
| Problem | Cause | Solution | |---------|-------|----------| | All claims flagged as ENTAILED in prior check | Claims too generic / common knowledge | Make claims more specific with file:line evidence | | Verification takes too long | Too many claims decomposed | Focus on high-stakes claims (causal, root cause, numbers) | | Flagged claim is actually correct | Prior knowledge happened to match | Mark as REVIEW; correct claim with unnecessary citation | | Self-verification missed a hallucination | Circular self-checking bias | Use conservative flagging; flag for human review when uncertain |
The Bottom Line
Every claim needs to EARN its confidence.
If you would believe a claim even WITHOUT the cited evidence:
- You're not reasoning from evidence
- You're pattern-matching from training
- Your citation is decorative, not informative
- FLAG IT AND FIX IT
The goal is not to sound authoritative. The goal is to BE correct because your evidence supports your claims.
No confabulation. No decorative citations. No training-data bleed.
Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: viksant
- Source: viksant/swe-skills
- License: MIT
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet — be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.