Install
$ agentstack add skill-jpbaking-agentic-tests-agentic-test-update ✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ✓ Network access No
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
Verified badge
Passed review? Show it. Paste this badge into your README, it links to the public security report.
Reliability & compatibility
Declared compatibility
Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.
We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.
How agent discovery & health will work →About
Agentic Test Update
Agent tests lock old behavior; the user changed behavior on purpose. This skill re-locks — but only where the user confirms the change was intended. Unconfirmed failures are potential regressions and must stay failing.
RULES
- NEVER edit main code or user tests. Only agent tests and plan files.
- Never update a test without the user's explicit per-diff confirmation. No batch "update all".
- An updated test must assert the NEW actual behavior (test-quality rules from
agentic-unit-testapply). Never weaken a test to pass both old and new behavior.
Step 1 — Collect failures
- Run the full agent-test suite.
- All green? Report "nothing to update" and STOP.
- For each failing test, capture: test name, file, expected (old locked behavior), actual (new behavior). Group by source file.
Step 2 — Confirm with the user, per behavior change
For each group, show a compact diff and ask ONE question:
src/pricing.ts — 3 agent tests failing
old locked: priceWithTax(0, 0.2) → 0
new actual: priceWithTax(0, 0.2) → throws RangeError
Intended change? (yes = update tests to lock new behavior / no = keep failing as regression)
Record every answer in agentic-test-plan.md under a ## Behavior updates section: confirmed or regression.
Step 3 — Update confirmed tests
For each confirmed group:
- Rewrite the failing assertions to lock the NEW behavior. Keep test names honest (rename if the name describes old behavior).
- Run the test file 3× (flakiness check). 3 attempts; after the 3rd failure revert the test and mark it
FAILED to update:in the plan. - Lint if configured.
Leave every regression test untouched and failing.
Step 4 — Report
- Updated: tests re-locked to new behavior (per file).
- Regressions: failing tests the user did NOT confirm — listed loudly; the suite is intentionally left red until the user fixes main code or re-runs this skill.
- Failed to update: 3-attempt casualties with reasons (or "none").
- Final suite status: green, or red-with-known-regressions.
Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: jpbaking
- Source: jpbaking/agentic-tests
- License: MIT
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet, be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.