AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
SKILL verified MIT Self-run

Litmus

skill-hacka0wi-litmus-litmus · by hacka0wi

>

No reviews yet
0 installs
32 views
0.0% view→install

Install

$ agentstack add skill-hacka0wi-litmus-litmus

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access No
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets Used
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/skill-hacka0wi-litmus-litmus)

Reliability & compatibility

Security review passed
0 installs to date
no reviews yet
2mo ago

Declared compatibility

Claude CodeClaude Desktop

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Litmus? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

Litmus — the live test that a fix is real

Hard rule: never mark a ticket done / ready-to-test on deploy, unit-test, or DB-only evidence. Reproduce the ticket's steps in the real UI as the specified user, observe the defect, then observe it gone after the fix. If you can't reproduce (wrong user / no access / missing data), say so and DO NOT close it.

> Environment specifics — app host, tracker API + ids, DB hosts, credentials, relay pods — belong in a > private/local config (e.g. agent memory or a .env), never in this public skill. The steps > below reference them as placeholders.

0. Read the ticket

Pull the user, test steps, actual, expected, and the screen URL. If the tracker redacts/escapes the raw description, read the rendered issue page instead.

1. Open the app as the ticket's user

  • Drive the browser via an MCP browser tool; create/restore a tab.
  • Switch account through the app's own "sign in as another user" — **fill the email, then have the human

type the password and submit. Never type/submit a password yourself.**

  • Select the correct sub-system/tenant if the app has several (wrong one shows wrong menus).
  • Confirm identity from the app's whoami/userinfo endpoint.
  • Access check: if that user lacks the menu/screen, it's an access gap — note it on the ticket and

do NOT guess-fix. (Don't modify permissions yourself.)

2. Reproduce — follow the steps exactly

Navigate the exact screen; click through the steps; screenshot the failing state; confirm it matches the ticket's actual. If the app embeds modules in an iframe, reach them via the iframe's contentWindow/contentDocument.

3. Instrument the running app (pick what fits)

  • Network hook — hook BOTH fetch AND XHR. Many SPA calls use fetch; an XHR-only hook misses

them. Capture request body + response; unwrap nested payloads to read the real fields/action.

  • Front-end scope: for AngularJS, walk angular.element(el).scope() to read the actual data array —

check counts/dups/values from state, don't trust the rendered (paginated) page.

  • Exercise the real function with dialogs stubbed: when the UI flow is blocked, override the confirm

dialog to auto-accept, call the page's real handler, and read the captured request/result. Verify the deployed code actually contains the fix: fn.toString().includes('...').

  • Unique variable names per eval — a shared eval scope throws "already declared" if you reuse names.

4. Verify in the database (source of truth)

  • Reach the DB through a relay (e.g. a socat pod + kubectl port-forward) and query with a real driver

(e.g. python oracledb, disable_oob=True); lightweight CLIs often fail where the driver works.

  • Trust the live DB definition of procs/functions as canonical, not the repo .sql if your team

edits the DB directly — read the live source and diff environments.

  • Always tear the relay/port-forward down when done.

5. Prove the fix, then update the tracker

  • Show concrete before/after: payload now carries the right value / list now de-duped / the exact DB row

is now correct. Capture both the SUCCESS response AND the persisted state.

  • If you created or changed test data to verify, revert it; before any destructive DB change, **back

up first** (CREATE TABLE AS SELECT ...) and confirm the rows are safe (e.g. childless) to touch.

  • Comment on the ticket with the evidence (captured payload, before/after counts, DB values) and set the

status. Get a human's confirmation before destructive shared-data changes (blast radius).

Field gotchas

  • Browser tab groups can reset and silently drop the session — re-create and re-login.
  • "Browser not connected" is usually transient — wait ~10s and retry.
  • Search boxes that don't filter on type → verify counts from app state, not the rendered list.
  • Shared tables (one row used by many modules): identify the keeper, back up, and confirm before deleting.

Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.