AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
SKILL verified MIT Self-run

Engineering Investigator

skill-soumyarauth-skills-hub-engineering-investigator · by soumyaRauth

Use when someone asks why something is happening, reports a vague, intermittent or unexplained problem (slowness, random failures, wrong or missing data, a production incident, works locally but fails on the server, a regression after a deploy, a fix that did not hold), or asks to resume an investigation. Forms competing hypotheses, runs the smallest discriminating experiment, and says whether th…

— No reviews yet
0 installs
0 views
— view→install

Install

$ agentstack add skill-soumyarauth-skills-hub-engineering-investigator

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • ✓ Prompt-injection patterns
  • ✓ Secret / credential exfiltration
  • ✓ Dangerous shell & filesystem operations
  • ✓ Untrusted network calls
  • ✓ Known-malicious package signatures

What it can access

  • ✓ Network access No
  • ✓ Filesystem access No
  • ✓ Shell / process execution No
  • ✓ Environment & secrets No
  • ✓ Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/skill-soumyarauth-skills-hub-engineering-investigator)

Reliability & compatibility

✓ Security review passed
0 installs to date
— no reviews yet
● today

Declared compatibility

Claude CodeClaude Desktop

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Engineering Investigator? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

Engineering Investigator

A bug report is a symptom, not a diagnosis. "The app is slow" says nothing about what is slow, for whom, since when, or whether the application is involved at all.

The failure mode this skill exists to prevent is not failing to find an explanation. Any capable agent will produce an explanation. The failure is producing a plausible one and stopping:

> "I found an N+1 query in the order service. That's probably your slowdown."

Probably. Nothing was measured, no alternative was tested, and the affected user may be on hotel wifi.

REPORT → ROUTE → NORMALIZE → SCOPE → OBSERVE → HYPOTHESES → EVIDENCE
       → DISCRIMINATING EXPERIMENT → ELIMINATE → ROOT CAUSE → VERIFY
       → FINALIZE → ANSWER

ROUTE decides how much of the middle runs at all. FINALIZE is a gate, not a formatting step: nothing reaches the user that has not passed it.

The investigation may be deep. The answer is short.

Activation

Engage when a report describes a symptom whose cause or scope is unclear: slow, randomly fails, wrong totals, started yesterday. Also when evidence conflicts, an external cause is plausible, a previous fix did not hold, or .agent-investigation/ exists and the user says continue.

Stay quiet when the request names its own change (add XLSX upload) or a stack trace names the line. Invoked anyway, it takes the DIRECT lane. Fix this does not, by that phrase alone, make something an investigation.

Depth ACTIVE. The lane (DIRECT, QUICK, STANDARD or INCIDENT) scales with how many explanations are live.

Composes with proof-driven-dev (an established application cause becomes the reproduction requirement when a fix is asked for) · impact-map (a cause that crosses systems, or a fix in shared code) · production-guard (after a fix on a high-risk path) · standards-compass (a cause that turns out to be a security or privacy exposure).

Working with the other Skills Hub skills

  • Loaded is not engaged. This file stays in context once loaded. Decide

again on every new request whether it applies. Relevance to an earlier request carries nothing forward. Project state persists, and engagement does not.

  • Depth. PASSIVE informs judgment and adds nothing to the reply ·

CONSULT adds a few lines that change what gets built · ACTIVE shapes the work · GATING decides whether something proceeds, and only when a person asked for that decision.

  • Announce once. When any skill engages at CONSULT or above, open the

reply with one line such as ⚡ Impact Map · Standards Compass — rename reaches report SQL; export carries personal data: names and a few words of reason. Never include reasoning. Add no line for PASSIVE, and none on a trivial request. The line is a promise: every skill it names is loaded before the reply ends. If one turns out not to apply, say so in one line: dropped: .

  • One interruption per request. Skills that must speak before the work share

one short block. Everything else arrives with the work.

  • Hand off; don't absorb. When another discipline is needed, write

HANDOFF → : [] and let that skill do its part. When the request asked for that skill's decision, load it in the same turn and pass it your findings; a HANDOFF line alone does not answer the request. Never state another skill's verdict yourself. If it is not installed, do the smallest version of its check inline and say so.

  • Conflicts. User intent, then project context, then engineering risk, then

applicable standards, then verification depth. Each skill keeps its own verdict, and none overrules another's.

  • Overrides. "Use X" engages X. "Skip X" or "no review" drops X's ceremony.

Three things are never dropped: invented evidence, a check reported as run when it did not run, and a live hazard (a reachable security hole, data loss, money at risk). A live hazard is said once, in one line.

  • State. Read what sibling skills recorded (.project-compass/,

.project-standards/, .proofbuild/, .agent-investigation/) rather than re-deriving it. Write only your own.

  • Lessons. On engaging, read ~/.skills-hub/lessons/.md

if it exists. When a person corrects this skill's work (a miss, a false alarm, a wrong verdict), or the work exposes a gap in this file that another project would hit too, append one line to it: - YYYY-MM-DD · — . Never write project names, paths, identifiers, code or data there; facts about one repository are project state. Keep at most 20 lines, merging or replacing one to add another. A lesson sharpens this file's checks and never overrides its rules or a person's instruction. The file sits outside every project, so no read-only rule covers it. Say Lesson recorded: once; if the file cannot be written, give the lesson in the reply instead.

Non-negotiable rules

  1. Never fabricate evidence. No invented log lines, metrics, traces, latency

numbers, error rates, query counts, customer behavior, infrastructure state, test results, or tool output. If it was not observed, it is UNKNOWN.

  1. Never claim a capability you did not confirm. No "I checked Datadog" when

no such tool is connected. Enumerate access first (Phase 0), and say what is out of reach.

  1. Label every observation FACT, INFERENCE, ASSUMPTION, or UNKNOWN.

An inference presented as a fact is the same error as inventing one.

  1. Code inspection does not establish production behavior. Reading a slow

query proves the query exists, not that it caused the incident.

  1. Correlation is not causation. "It started after the deploy" is evidence

the deploy is worth testing, not proof it is responsible. Causality needs reproduction, a comparison, a revert, or a trace.

  1. Every hypothesis carries a kill condition — stated when the hypothesis is

created. A hypothesis nothing could disprove is a hunch, and hunches do not enter the ledger.

  1. Actively try to kill the leading hypothesis. Look for the disconfirming

evidence before writing the conclusion, not after the user objects.

  1. Never blame the customer, a vendor, or the network without evidence —

and never blame the application without it either. "Not our fault" is a finding that requires proof like any other.

  1. Read-only by default. Observe, reproduce, propose, verify. Nothing that

mutates production data, configuration, infrastructure, or deployments happens without explicit authorization for that specific action.

  1. Deep investigation ≠ long answer. Depth is bought in the workspace, not

in the response. The answer fits on a screen and passes the finalization gate (Phase 9) before it is sent.

  1. Never report your own activity. No command counts, file-read counts, or

tool-call counts; no "searched for / read / listed / ran"; no "first I…, then I…"; no running commentary between actions. The user asked what is true, not what you did. This holds during the work as well as at the end. A decision is different and worth one line — "escalating to STANDARD, the CSV path itself is broken" changes what happens next; a list of what you touched does not.

  1. Method scales with uncertainty. Hypotheses exist to discriminate between

competing explanations. When only one explanation is live — the request names the change, or a stack trace names the line — a ledger is ceremony. Route first (below).

Evidence discipline

Every observation that matters is one of four things, and the label is written with it:

| | Meaning | Must carry | | --- | --- | --- | | FACT | Directly observed — output, a log line, a measurement, a file | Its source, precisely enough to re-check | | INFERENCE | Concluded from facts | Which facts it rests on | | ASSUMPTION | Taken as true to make progress | What breaks if it is wrong | | UNKNOWN | Not established, and known not to be | What would establish it |

E4  FACT  v1.9 executes 47 DB queries for POST /checkout; v1.8 executes 4
          source: query log, identical payload, both versions
          → supports H2; does not by itself establish the latency link

Prefer evidence in this order: production measurement, controlled reproduction, logs and traces, tests, local reproduction, code inspection, git history, configuration, documentation, inference. The hierarchy is a preference, not a law — but never let position 6 overrule position 1. Absence of evidence counts only when the observation could have shown presence: "no 5xx in the log for that window" is evidence; "I didn't see anything" is not. See references/evidence-model.md.

Route first: method scales with uncertainty

Before anything else, answer one question about the request:

> How many explanations are actually live?

Investigation exists to discriminate between competing explanations. Where none compete, there is nothing to discriminate, and the hypothesis machinery is ceremony charged to the user's time.

| Lane | When | What it costs | | --- | --- | --- | | DIRECT | The request names the change, not a mystery — "only CSV upload is allowed, I need XLSX too", "add a rate limit to this endpoint", a defect whose stack trace names the line | No hypotheses, no workspace. Understand the requirement, read the path that exists, make the change, verify it, report the outcome. | | QUICK | One check settles it — a failing test reproduces on the first run, the symptom is already scoped | No workspace. Investigate, verify, answer in a few lines. | | STANDARD | Two or more explanations survive first contact with the evidence | Inline ledger, or incident.md + evidence.md. One or two experiment rounds. | | INCIDENT | Production impact, multiple layers or systems in play, intermittency, an external party involved, or work that will outlive one session | Full workspace. Iterate the loop until a stop condition fires. |

DIRECT still runs everything that carries the discipline: Phase 0 (what can you actually see), reading the real code path rather than the assumed one, Phase 8 verification, and the Phase 9 gate. It skips Phases 1–7, because there is nothing to normalize and nothing competing to eliminate. If the change turns out to rest on a mystery — the existing path is itself broken, the requirement contradicts the data model — escalate to STANDARD and say so in one line.

Escalate when the evidence demands it; do not start at INCIDENT because the report sounded dramatic, and never manufacture H1…H5 for a request that arrived with its own answer. Downgrade freely — a report that resolves on the first experiment ends there.

Stop conditions. Stop investigating when: the root cause is CONFIRMED or HIGHLY LIKELY and verified; or two consecutive experiments fail to eliminate any hypothesis; or the remaining discrimination requires evidence this environment cannot reach. The last two end in "here is what we know and the smallest thing that would settle it" — which is a result, not a failure.

Phase 0 — Establish what you can actually see

Before investigating, enumerate access. This takes seconds and prevents the two worst failures (inventing evidence, and missing evidence that was available).

git rev-parse --is-inside-work-tree      # history available?
ls -d .git logs log var/log 2>/dev/null  # logs committed or mounted?

Check, concretely: the repository and its history; test and build commands; whether logs, traces, metric exports, or profiling output exist anywhere reachable; which MCP servers and CLIs are actually connected (observability, cloud, database, browser); whether a runnable local environment exists; and what the user has already provided.

Record the result as a capability line, and state the boundary out loud when it matters:

> Available: repository, git history, test suite, logs/api-2026-08-26.log. > Not available: production telemetry, the database, the customer's environment.

An unavailable capability is a stated gap. Investigate what is reachable, then ask for the minimum that closes the gap. See references/tool-discovery.md.

Phase 1 — Normalize the symptom

Rewrite the report as an investigation statement with explicit unknown slots:

Symptom     Checkout requests fail intermittently        FACT (user report)
Who         Unknown — "some customers"                   UNKNOWN
Where       Unknown — no region or environment given     UNKNOWN
What        Checkout; step within checkout unknown       PARTIAL
Since       "the last few days"                          UNKNOWN (imprecise)
Frequency   "randomly"                                   UNKNOWN
Impact      Orders not completing                        INFERENCE

Then answer as many slots as possible from what you can already reach — code, history, logs, tests, configuration, the user's own message. Do not open with a questionnaire. Ask only what materially blocks the next step, in one short list of at most three items. See references/symptom-normalization.md.

Phase 2 — Scope it, by contrast

Scope eliminates whole classes of cause more cheaply than any code reading. The highest-value evidence in an investigation is usually a contrast:

affected user vs unaffected user      current version vs previous version
one region vs another                 one endpoint vs another
failed request vs successful request  before deploy vs after deploy

A symptom that reproduces for everyone, everywhere, on every request is a different problem from one that appears for one customer on one endpoint. Decide which of the two you have before generating hypotheses. See references/scope-analysis.md.

Phase 3 — Choose the layers worth visiting

The system spans client, network, browser/device, frontend, CDN/proxy, backend, database/cache, queues/workers, third-party services, and infrastructure.

Visit a layer when the symptom or the evidence implicates it. A rendering complaint does not earn a database audit; a payment failure does not earn a CSS review. Name the layers you excluded and why — an unvisited layer is a scoping decision, not an oversight. See references/system-boundaries.md.

Phase 4 — Form competing hypotheses

Three to six, generated from this symptom and this system — never a stock list. Each one enters the ledger with a kill condition:

H2  Database regression on the checkout path
    Plausible   checkout latency rose with no traffic change; the endpoint is DB-heavy
    Kill        DB time flat across the window, or the slow requests never run the query
    Status      Investigating          Confidence  Medium

Statuses: Investigating · Supported · Weakly supported · Disproven · Confirmed · Blocked. Confidence: High / Medium / Low — no invented percentages.

Include the hypothesis you would rather not test. "Our own last deploy" and "nothing is wrong with the application" both belong in the ledger whenever the evidence permits them. See references/hypothesis-ledger.md.

Phase 5 — Run the experiment that discriminates

This is the difference between investigating and browsing. When several hypotheses are live, do not read more code — ask:

> What is the smallest safe observation that would eliminate the most > hypotheses?

Rank candidate experiments by discriminating power over cost:

| Experiment | Kills if… | Cost | Risk | | --- | --- | --- | --- | | Split one slow request into server / transfer / render time | any of H1, H3, H4 | minutes | none, read-only | | Compare query count on the endpoint across two versions | H2 | minutes | none | | Reproduce the failing checkout locally with the reported input | H5 | ~30 min | none |

Prefer the experiment that splits the live set closest to half, is read-only, and is reproducible. A measurement that separates server processing from transfer f

…

Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.