AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
SKILL verified MIT Self-run

Smell The Numbers

skill-mikestangdevs-craft-skills-smell-the-numbers · by mikestangdevs

Use whenever code produces an output someone will read — a report, a metric, a chart, a calculation result, a dashboard value — and especially when a result is surprising, suspiciously clean, or just changed a lot. Treats every implausible number as guilty until root-caused: all-zeros, -0, a value at 1858% of its physical limit, a metric that moved 40x from one run to the next, a flagship case th…

No reviews yet
0 installs
28 views
0.0% view→install

Install

$ agentstack add skill-mikestangdevs-craft-skills-smell-the-numbers

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access No
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets No
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/skill-mikestangdevs-craft-skills-smell-the-numbers)

Reliability & compatibility

Security review passed
0 installs to date
no reviews yet
3mo ago

Declared compatibility

Claude CodeClaude Desktop

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Smell The Numbers? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

Smell the Numbers

The failure mode this fixes

The code runs, the page renders, the report generates — and the numbers are nonsense. A line carrying 1858% of its rated capacity. A formula panel showing all zeros. A -0 in a financial summary. A KPI that moved 40x between runs of "the same" scenario. Agents (and tired humans) accept these because the mechanism worked: no exception, no red test, output produced. "It runs" gets mistaken for "it's right."

The worse version is the cosmetic fix: clamp the value, hide the row, round the -0, smooth the curve — make the output look plausible while the defect underneath keeps feeding everything else. That converts a visible bug into an invisible one.

This skill installs the domain-expert reflex: read the output the way someone who knows the territory would, flag everything that contradicts physical limits, business reality, or its own history — and treat each flag as a root-cause investigation, not a display problem.

When to Use This Skill

  • Output just changed significantly after a refactor, migration, or "no-behavior-change" cleanup
  • A result is surprising in either direction — catastrophic where success was expected, or suspiciously perfect
  • You see sentinel smells: zeros where work happened, -0, NaN leaking into display, values exceeding a hard limit, percentages outside [0,100], totals that don't sum
  • A metric moved by an order of magnitude with no input change that explains it
  • You're about to show the output to someone who will make a decision from it

Don't use when: the output has no ground truth to compare against and no internal consistency to violate (pure creative output). Don't let it become paralysis — the skill is a screen, not a proof of correctness for every digit.

Instructions

1. Read the output as the skeptic, not the author

Before celebrating that output exists, actually read it. Ask the questions a domain reviewer would: Is this physically/financially/logically possible? Does it match the order of magnitude I'd estimate on a napkin? Is it consistent with the last known-good run? Do the parts sum to the whole?

2. Rank surprise as evidence

A surprising result has exactly three explanations, in descending likelihood: (1) a bug in the code, (2) a defect in the input data, (3) reality is genuinely surprising. Investigate in that order. "The strongest system in the fleet just failed the easiest scenario" is a bug report, not a finding — until the trace proves otherwise.

3. Trace one bad number all the way down

Pick the most implausible value and follow it backwards through the computation to the first place it goes wrong: the formula, the unit conversion, the field that was a calendar year where the math wanted a duration, the aggregation that double-counted. Fix at the source. One fully-traced number teaches you more than ten re-runs.

4. Never fix the display

If the investigation surfaces pressure to clamp, hide, filter, smooth, or reformat the bad value — that's the moment you're about to convert a loud bug into a silent one. The display gets fixed only after the underlying number is right. (Legitimate display issues exist — a -0 from float formatting after the value is verified correct — but earn that conclusion with the trace, don't assume it.)

5. Ask what else this poisoned

A wrong number rarely stays in its cell. The same broken conversion feeds the chart, the summary, the export, the downstream model. Once root-caused, sweep for every consumer of the defective path before declaring the incident closed.

Output format

Output plausibility check: 
Smells found: 
Root cause: 
Classification: code bug / data defect / genuinely surprising (with proof)
Fix: 
Blast radius: 

Anti-Patterns

| Anti-Pattern | Why it defeats the skill | |---|---| | "It runs" = "it's right" | accepting output because no exception fired. | | Cosmetic correction | clamping, hiding, or smoothing the symptom; the dashboard heals, the model stays sick. | | Surprise worship | writing a narrative for why the shocking result might be true, before checking whether it's a unit error. | | Re-run roulette | running it again hoping for better numbers instead of tracing the bad one. | | Celebrating the suspiciously clean | zero variance, perfect curves, and exact round totals are smells too. |

Mental Model

> Every output number is a witness testifying about your code. When a witness says something impossible, you don't transcribe it neatly into the record — you cross-examine until you know why they said it.

Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.