AgentStack
SKILL verified MIT Self-run

Prompt Injection Tester

skill-novacode37-claude-security-skills-prompt-injection-tester · by NovaCode37

>-

No reviews yet
0 installs
14 views
0.0% view→install

Install

$ agentstack add skill-novacode37-claude-security-skills-prompt-injection-tester

✓ scanned · ✓ verified — works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access No
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets No
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

Are you the author of Prompt Injection Tester? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

Prompt Injection Tester

A defensive red-team harness for evaluating the prompt-injection resistance of LLM applications you own or are authorized to test. It ships a library of well-documented public attack techniques and a canary-based detection engine that decides whether each attack succeeded — then scores overall resilience.

> ⚠️ Use only against systems you own or have permission to test. The payloads > are public hardening techniques, intended to strengthen guardrails.

When to use this skill

  • "Is my chatbot vulnerable to prompt injection / jailbreaks?"
  • "Red-team / pentest my LLM app's system prompt."
  • "Score how well my guardrails resist instruction-override attacks."
  • Regression-testing guardrails in CI after a prompt change.

Attack categories covered

instruction-override · system-prompt-leak · role-play (DAN-style) · delimiter-escape · encoding (base64/leetspeak) · data-exfiltration (indirect injection) · refusal-suppression.

How it works

  1. A unique canary secret is embedded into a guarded system prompt.
  2. Each payload is sent as the user turn.
  3. The response is scored: it's vulnerable if it hits an attack

success-marker or leaks the canary; resisted if it refuses.

  1. You get a resilience score (0–100) and a per-category breakdown.

How to run it

List the payload library (no model calls):

python skills/prompt-injection-tester/attacker.py --list
python skills/prompt-injection-tester/attacker.py --list --category role-play

See the harness detect a weak vs. a hardened model:

python skills/prompt-injection-tester/attacker.py --demo            # weak model
python skills/prompt-injection-tester/attacker.py --demo --hardened # safe model
python skills/prompt-injection-tester/attacker.py --demo --json

Testing a real model (recommended workflow for Claude)

Wrap the user's backend in a model(system_prompt, user_prompt) -> str callable and run the suite. Example skeleton the user can adapt:

import sys
sys.path.insert(0, "skills/prompt-injection-tester")
import attacker

def model(system_prompt, user_prompt):
    # call the user's real API/app here and return the text response
    return my_llm_app(system=system_prompt, user=user_prompt)

results = attacker.run_suite(model)
print(attacker.summarize(results))

Then:

  1. Report the resilience score and list every failed_payload.
  2. For each failure, show the technique and recommend a mitigation —

e.g. stronger system-prompt instructions, input/output canary filtering, privilege separation for tool calls, or a dedicated injection classifier.

  1. Re-run after fixes to confirm the regression is closed.

Offline scoring

If you already captured a model response, score it without any network call:

import attacker
payload = attacker.PAYLOADS[0]
verdict = attacker.score_response(payload, response_text, canary="CANARY-abc")
print(verdict.vulnerable, verdict.reason)

Limitations

Marker/canary detection is a heuristic: a refusal-shaped response that still complies, or a creative leak, may need human review (score_response flags ambiguous cases as "review manually"). Treat the score as a guardrail signal, not a guarantee.

Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet — be the first.

Versions

  • v0.1.0 Imported from the upstream source.