Install
$ agentstack add skill-ssrjkk-claude-skills-ai-safety ✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ✓ Network access No
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
Verified badge
Passed review? Show it. Paste this badge into your README, it links to the public security report.
Reliability & compatibility
Declared compatibility
Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.
We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.
How agent discovery & health will work →About
AI Safety
> Implement responsible AI practices including guardrails, monitoring, and ethical guidelines.
Quick Start
from guardrails import Guard
from guardrails.validators import Validator
class NoPIIValidator(Validator):
def validate(self, value: str, metadata: dict) -> dict:
import re
# Check for emails, SSNs, credit cards
patterns = {
"email": r'\b[\w.+-]+@[\w-]+\.[\w.]+\b',
"ssn": r'\b\d{3}-\d{2}-\d{4}\b',
"credit_card": r'\b\d{4}[- ]?\d{4}[- ]?\d{4}[- ]?\d{4}\b'
}
found = {name: re.findall(pat, value)
for name, pat in patterns.items()
if re.search(pat, value)}
if found:
return {"valid": False, "error": f"PII detected: {found}"}
return {"valid": True}
# Content moderation guard
content_guard = Guard().use(NoPIIValidator())
# Usage
result = content_guard.validate("My email is user@example.com")
print(result.error) # "PII detected: {'email': ['user@example.com']}"
Key Concepts
AI safety spans: prompt injection prevention, PII/redaction, content moderation, output validation, rate limiting, audit logging, and bias monitoring. Defense in depth — multiple layers of protection.
When to Use
- Any production LLM deployment
- Applications handling user data or PII
- Systems where AI outputs affect real-world decisions
- Regulated industries (healthcare, finance, legal)
Validation
- Prompt injection attempts are blocked or sanitized
- PII is detected and redacted in both inputs and outputs
- Audit logs capture all LLM interactions for review
- Rate limits prevent abuse and cost overruns
Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: ssrjkk
- Source: ssrjkk/claude-skills
- License: MIT
- Homepage: https://claude.ai
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet, be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.