Prompt Injection Reviewer
A Claude skill from 45ck/llm-agent-security-skills.
Llm Default Architecture
Use when deciding how to make frontier LLMs the default cognitive engine in a task, automation, feature, or workflow. Trigger when the main risk is falling back to deterministic-first, keyword-first, or semantic-first designs before checking what frontier models can already do.
Legacy Automation Fallback
Use when reviewing a prompt, spec, workflow, or architecture for deterministic-first regressions that replace LLM judgment with rules, keyword routing, fixed templates, or brittle scripts. Trigger whenever an LLM-native task is quietly being collapsed into old automation logic.
Autonomy Boundary Checker
Review an agent workflow for where it may act autonomously, where it must ask, and where human approval is required.
Self Improving Agent Loop
Design governed self-improving agent loops that sense work, model failure modes, act in small slices, gate evidence, and turn learnings into tracked proposals.
Host Instruction Drift Checker
Check whether host instruction surfaces such as AGENTS.md, CLAUDE.md, Codex agents, Cursor rules, Copilot instructions, and SKILL.md files have drifted apart.
Record Replay Skill Reviewer
Review skills generated from recorded workflows before they are installed, shared, or treated as reusable automation.
Multi Agent Workflow Reviewer
Review or design multi-agent workflows with clear ownership, handoffs, evidence gates, conflict handling, and minimal coordination overhead.
Agent Run Evidence Reviewer
Review agent run traces, eval summaries, tool logs, and handoff evidence before accepting a workflow result or self-improvement proposal.
Skill Provenance Reviewer
Review skills, agents, rules, plugins, and MCP/tool packs for license, source, script, permission, and supply-chain risk before adoption.
Tool Permission Boundary Checker
A Claude skill from 45ck/llm-agent-security-skills.
Demo Review Surface
Create safe static review surfaces for demo and QA media. Use when Codex needs an HTML or Markdown review page comparing demo videos, source specs, QA reports, screenshots, storyboards, quality findings, and evidence links from demo-machine, manual-qa-machine, Playwright, or video analysis outputs.
Fixture Skill
Synthetic skill for intake scanner coverage.
Agent Task Shaping
Shape ambiguous or frontier-agent work into a bounded agent task with outcome, scope, context, tools, autonomy level, and verification evidence.
Frontier Model Context
Use when planning, evaluating, or critiquing work where the key problem is stale capability priors about frontier LLMs. Trigger when the maintainer wants the model to reason from current benchmark-informed assumptions rather than defaulting to weak-agent, deterministic-first, or human-first thinking.
Evidence Gap Review
Analyse specgraph verify output and identify what evidence is missing to advance specs to their next state.
Gh Review Followthrough
Address GitHub PR review threads or issue comments with explicit comment selection, repo-grounded fixes, and concise reply-ready summaries.
Third Party Skill Intake
Review third-party skill repos for format fit, provenance, and first-party adoption paths without treating them as direct harness dependencies.
Noslop Commit Gate
Run the noslop pre-commit quality gate, interpret failures, and enforce content-aware protection rules before every commit.
Demo Story Packager
Package completed demo-machine runs, browser capture artifacts, or evidence-backed product demos into source-linked handoff bundles. Use when Codex needs to organize `.demo.yaml`, `events.json`, `quality.json`, screenshots, storyboard outputs, rendered videos, review prompts, or release-ready demo assets without losing provenance.
Noslop Pr Gate
Run the noslop pre-PR quality gate and handle the noslop-approved escape hatch for intentional config weakening before opening a pull request.
Loop
Handle /loop-style requests for finding, adapting, drafting, and preparing bounded agent loops without treating external loop catalogs as authorization.
Tool Permission Planner
Design least-privilege tool access for agent workflows, MCP servers, connectors, browser automation, shell commands, and external side effects.
External Skill Fixture Builder
Turn external skill, rule, plugin, MCP, task-memory, and multi-agent workspace patterns into safe fixture coverage rather than live dependencies.
Demo Release Packager
Assemble approved demo media, canonical source specs, evidence links, provenance, and promotion notes into a release or handoff bundle. Use when Codex needs to prepare demo-machine, manual QA, or generated media outputs for docs, launch pages, changelogs, social previews, or another agent.
Agent Memory Design Reviewer
Review agent memory, handoff, retrieval, and learning designs for usefulness, provenance, privacy, staleness, and poisoning risk.
Guardrail Policy Writer
A Claude skill from 45ck/llm-agent-security-skills.
Spec Writer
Write specgraph spec documents with correct YAML frontmatter, evidence requirements, and section structure.
Waiver Writer
Write a specgraph evidence waiver for a spec requirement that cannot currently be satisfied, with a clear justification and expiry plan.
Design Token Alignment
Align design artifacts with code tokens and component APIs so spacing, color, type, and states stay consistent through implementation.
Digital Expert Test
Use when deciding whether an AI agent should even attempt a task. Trigger for novel integrations, device problems, workflow automation, cross-disciplinary analysis, or any idea that sounds “probably impossible” or “too much work” under human-first reasoning.
Model Review Artifact
Shape source-backed model and diagram review artifacts for architecture, UML-style, C4, dependency, and workflow views.
Risky Skill
Synthetic risky skill for intake scanner coverage.
Llm Output Handling Checker
A Claude skill from 45ck/llm-agent-security-skills.
Context Leak Reviewer
A Claude skill from 45ck/llm-agent-security-skills.
Verify Interpreter
Interpret the output of `specgraph verify` and explain what each status means and what to do next.
Annotation Writer
Write correct @spec JSDoc annotations in source files to link code to specgraph spec documents and produce E0 evidence.
Open Genome Agent
Local-first DNA and VCF analysis copilot for evidence-bound genomics workflows, confidence tiers, and Claude/Codex support.
Verification First Delivery
Use when shipping changes, building automations with side effects, or planning long-running agent work. Trigger for code changes, integrations, migrations, device or network debugging, or any task where the agent needs deterministic evidence and rollback paths.
Issue Driven Delivery
Run implementation work from a tracked issue with explicit scope, branch intent, validation steps, and closeout evidence.
Pre Ai Scarcity Thinking
Use when reviewing business, career, org-design, or process advice that may still be optimizing for old human clerical bottlenecks. Trigger when advice feels capability-lagged, manual-labor-centric, or built for a world before frontier models made large parts of knowledge-work cheap.
Retrieval Trustworthiness Reviewer
A Claude skill from 45ck/llm-agent-security-skills.
Mcp Server Planning
Plan an MCP server with clear tool scope, auth model, error behavior, evaluation approach, and rollout boundaries.
Unsafe Tool Chain Detector
A Claude skill from 45ck/llm-agent-security-skills.
Agent Sandbox Reviewer
A Claude skill from 45ck/llm-agent-security-skills.
Agent Agency Reviewer
A Claude skill from 45ck/llm-agent-security-skills.
Review Ready Check
Check whether a change is ready for review by testing the claimed behavior, summarizing risk, and naming any gaps reviewers should know about.
Artifact Evidence Gate
Verify that developer artifacts are grounded, current, and safe to hand off.
Html Review Artifact
Produce safe, self-contained HTML review artifacts from canonical project sources.
App Integration Shaping
Decide how an external app or service should be integrated, including boundaries, auth scopes, ownership, and failure handling.
Developer Artifact Shaper
Choose the right developer artifact type, source of truth, and review surface for a task.
Demo Slideshow Edit
Plan no-caption slideshow-style MP4s and frame reels from demo-machine screenshots, storyboard frames, selected video spans, or manual QA evidence. Use when Codex needs a polished still-frame walkthrough, hero reel, or reviewable visual summary without recapturing the browser flow.
Memory Poisoning Risk Reviewer
A Claude skill from 45ck/llm-agent-security-skills.
Secret Exposure Reviewer
A Claude skill from 45ck/llm-agent-security-skills.
Noslop Setup
Install and configure noslop quality gates in an existing repository, including git hooks, CI workflows, and Claude Code guardrails.
Task Memory Profile Planner
Choose the right task-memory and issue-state pattern for a repo: Beads, lightweight files, external tracker integration, or no durable task memory yet.
Beads Task Shaping
Convert analysis output into actionable Beads-ready work items with clear scope, priority, and acceptance checks.
Agent Handoff Planning
Structure a clean handoff between workflow agents with explicit outputs, open questions, and next-owner guidance.
Artifact Handoff Pack
Assemble the minimal artifact bundle needed for another agent or human to continue work.
Qa To Demo
Convert manual QA flows, QA reports, screenshots, network/console evidence, accessibility findings, or bug reproduction steps into demo-machine specs, reproducible demo plans, or short evidence-backed repro clips. Use when Codex needs to turn manual-qa-machine output into a product demo, release proof, or issue reproduction asset.
Context Engineering Planner
Plan the context, artifacts, memory, retrieval, and compression surfaces needed for an agent workflow or long-horizon task.
Figma Implementation Planning
Translate a Figma screen or component into an implementation plan covering structure, states, constraints, and component boundaries.
Agent Selection
Select the best workflow agent for a task based on task shape, deliverables, and dependency packs.
Demo Social Cut
Plan short silent demo cuts, slideshow edits, hero loops, and social clips from existing demo-machine or QA evidence. Use when Codex needs no-caption polished videos, frame-based reels, poster frames, or audience-specific excerpts derived from `events.json`, screenshots, storyboard evidence, or rendered demo runs.
Visual Source Artifact
Shape product, business, data, research, UX, and mockup artifacts as agent-readable sources with generated visual human review surfaces.
Gh Actions Failure Triage
Inspect failing GitHub Actions checks, isolate the actionable failure, and turn it into a concrete fix path with verification steps.