AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
SKILL verified MIT Self-run

Testing

skill-crewforth-crewforth-testing · by crewforth

|

— No reviews yet
0 installs
0 views
— view→install

Install

$ agentstack add skill-crewforth-crewforth-testing

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • ✓ Prompt-injection patterns
  • ✓ Secret / credential exfiltration
  • ✓ Dangerous shell & filesystem operations
  • ✓ Untrusted network calls
  • ✓ Known-malicious package signatures

What it can access

  • ✓ Network access No
  • ✓ Filesystem access No
  • ✓ Shell / process execution No
  • ✓ Environment & secrets No
  • ✓ Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/skill-crewforth-crewforth-testing)

Reliability & compatibility

✓ Security review passed
0 installs to date
— no reviews yet
● today

Declared compatibility

Claude CodeClaude Desktop

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Testing? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

Testing Discipline

Trigger phrases: "write a test", "run the tests", "coverage", "are the tests green", "unit test", "integration test", "tests pass locally", "flaky test", "test is flaky"

Goal: behavior correctness — test real behavior without breaking product code just to make a test pass.

Principles

  • Pyramid: many unit, fewer integration, few end-to-end (e2e). Limit e2e to critical flows.
  • AAA: Arrange-Act-Assert; one test = one behavior.
  • Isolation & determinism: external dependencies are mocked/faked; time and randomness are fixed; test order is independent.
  • Risk-coverage: risk, not metrics. Critical path + boundary + negative + authorization (IDOR/404) scenarios.
  • Naming: what_it_tests_under_which_condition_what_it_expects — on failure it is clear what broke.
  • Red-green: first a failing test, then the implementation.

Watch out

  • Flaky = bug: an occasionally failing test is not tolerated; it is fixed at the root.
  • In snapshot/golden-file tests, avoid needless brittleness (assert only the meaningful output).

Tests that cannot fail

Green is evidence only if the test could have gone red. Shapes that check nothing:

  • Assertion restates the code: it recomputes the implementation's own expression, so it moves with every change.
  • Setup guarantees the result: arrange plants the value the assert reads back; it holds even if the act never ran.
  • Expected value derived the way the code derives it: same formula, query or parser — the same wrong assumption on

both sides still matches.

  • A double asserted against its own stub: the fake is told to return X and the test asserts X; only the fake runs.
  • No assertion: it passes because nothing threw. If not throwing is the behavior under test, assert that.

Prove it can fail: break what the test covers — invert a condition, return a wrong constant — confirm it goes red, then restore. Still green means the test is documentation, not a gate: fix it or delete it, never count it as coverage.

Flaky triage

An intermittent failure gets exactly one outcome, decided the day it is seen: fix it, quarantine it (out of the gating lane, never out of the suite, with an owner and a date), or escalate it as the product defect it signals. Infrastructure flakes may get a bounded, logged retry; product flakes never do. Classifying them, the retry rule, and what a quarantine entry must carry: references/flaky-triage.md.

DoD (this skill's contribution)

  • The project's own test command is green — read it off the manifest/CI (dotnet test · npm test · pytest ·

go test ./...), don't assume one; critical paths are covered; no empty/meaningless tests.

Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.