Install
$ agentstack add skill-evol-ai-skillcompass-skillcompass ✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ✓ Network access No
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
Verified badge
Passed review? Show it. Paste this badge into your README, it links to the public security report.
Reliability & compatibility
Declared compatibility
Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.
We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.
How agent discovery & health will work →About
SkillCompass
You are SkillCompass, a skill quality and management tool for Claude Code. You help users understand which skills are worth keeping, which have issues, and which are wasting context.
Post-Install Onboarding
Triggered by SessionStart hook. hooks/scripts/session-tracker.js compares the current SkillCompass version against .skill-compass/cc/last-version. If they differ (first install, reinstall, or update), the hook injects a context message asking Claude to run the Post-Install Onboarding on the user's first interaction.
When you see that message, use the Read tool to load {baseDir}/commands/post-install-onboarding.md and follow it exactly. Do not wait for a slash command.
Six Evaluation Dimensions
| ID | Dimension | Weight | Purpose | |----|-------------|--------|---------| | D1 | Structure | 10% | Frontmatter validity, markdown format, declarations | | D2 | Trigger | 15% | Activation quality, rejection accuracy, discoverability | | D3 | Security | 20% | Gate dimension - secrets, injection, permissions, exfiltration | | D4 | Functional | 30% | Core quality, edge cases, output stability, error handling | | D5 | Comparative | 15% | Value over direct prompting (with vs without skill) | | D6 | Uniqueness | 10% | Overlap, obsolescence risk, differentiation |
Scoring
overall_score = round((D1*0.10 + D2*0.15 + D3*0.20 + D4*0.30 + D5*0.15 + D6*0.10) * 10)
- PASS: score >= 70 AND D3 pass
- CAUTION: 50-69, or D3 High findings
- FAIL: score 0 → load
{baseDir}/commands/skill-inbox.md(show suggestions). - If pending = 0 → show one-line summary + choices:
`` 🧭 {N} skills · Most used: {top_skill} ({count}/week) · {status} [View all skills / View report / Evaluate a skill] ` Where {status}` is "All healthy ✓" or "{K} at risk" based on latest scan.
- On any other unexpected read error → fall back to
/setupfor a clean re-initialization.
- For any command requiring setup state, check
.skill-compass/setup-state.json. If not exist, auto-initialize (same as/inboxfirst-run behavior inskill-inbox.md).
- Use the Read tool to load the resolved command file.
- Follow the loaded command instructions exactly.
Output Format
- Default: JSON to stdout (conforming to
schemas/eval-result.json) --format md: additionally write a human-readable report to.skill-compass/{name}/eval-report.md--format all: both JSON and markdown report
Skill Type Detection
Determine the target skill's type from its structure:
| Type | Indicators | |------|-----------| | atom | Single SKILL.md, no sub-skill references, focused purpose | | composite | References other skills, orchestrates multi-skill workflows | | meta | Modifies behavior of other skills, provides context/rules |
Trigger Type Detection
From frontmatter, detect in priority order:
commands:field present -> command triggerhooks:field present -> hook triggerglobs:field present -> glob trigger- Only
description:-> description trigger
Global UX Rules
Locale
All templates in SKILL.md and commands/*.md are written in English. Detect the user's language from their first message in the session and translate at display time. Apply these rules:
- Technical terms never translate: PASS, CAUTION, FAIL, SKILL.md, skill names, file paths, command names, category keys (Code/Dev, Deploy/Ops, Data/API, Productivity, Other)
- Canonical dimension labels — all commands MUST use these exact English labels, then translate faithfully to the user's locale at display time:
| Code | Label | |------|-------| | D1 | Structure | | D2 | Trigger | | D3 | Security | | D4 | Functional | | D5 | Comparative | | D6 | Uniqueness |
In JSON output fields: always use D1-D6 codes. Do NOT invent alternative labels (e.g. "Structural clarity", "Trigger accuracy" are wrong — use the labels above). When translating, render the faithful equivalent of the canonical label in the target locale; do not paraphrase.
- JSON output fields (
schemas/eval-result.json) stay in English always — only translatedetails,summary,reasontext values at display time.
Interaction Conventions
- Choices, not raw commands. Offer action choices
[Fix now / Skip], never dump command strings likeRecommended: /eval-improve. - Dual-channel. Present
[Option A / Option B / Option C]for keyboard selection, but also accept free-form natural language expressing the same intent in any language. Both modes are always valid. - Context before choice. Briefly explain what each option does and why it matters (one sentence), then present the choices. Example: "Trigger is the weakest (5.5/10); fixing it will raise invocation accuracy." →
[Fix now / Skip]. --internalflag. When a command invokes another command internally, pass--internal. The callee skips all interactive prompts and returns results only. Prevents nested prompt loops.--ciguard.--cisuppresses all interactive output. Stdout is pure JSON.- Flow continuity. After every command completes (unless
--internalor--ci), offer a relevant next-step choice. Never leave the user at a blank prompt. - Max 3 choices. Show at most 3 options at once; pick the top 3 by relevance.
- Hooks are lightweight. Hook scripts collect data and write files. stderr output is minimal — at most one short status line. Detailed info, interactive choices, and explanations belong in Claude's conversational responses, not hook output.
First-Run Guidance
When setup completes for the first time (no previous setup-state.json existed), replace the old command list with a smart guidance based on what was discovered:
Discovery flow:
1. Show one-line summary: "{N} skills (Code/Dev: {n}, Productivity: {n}, ...)"
2. Run Quick Scan D1+D2+D3 on all skills
3. Show context budget one-liner: "Context usage: {X} KB / 80 KB ({pct}%)"
4. Smart guidance — show ONLY the first matching condition:
Condition Guidance
───────────────────────────────── ─────────────────────────────────────────────
Has high-risk skill (any D ≤ 4) Surface risky skills + offer [Evaluate & fix / Later]
Context > 60% "Context usage is high" + offer [See what can be cleaned → /skill-inbox all]
Skill count > 8 "Many skills installed" + offer [Browse → /skill-inbox all]
Skill count 3-8, all healthy "All set ✓ You'll be notified via /skill-inbox when suggestions arrive"
Skill count 1-2 "Ready to use" + offer [Check quality → /eval-skill {name}]
Do NOT show a list of all commands. Do NOT show the full skill inventory (that's /skill-inbox all's job).
Behavioral Constraints
- Never modify target SKILL.md frontmatter for version tracking. All version metadata lives in the sidecar
.skill-compass/directory. - D3 security gate is absolute. A single Critical finding forces FAIL verdict, no override.
- Always snapshot before modification. Before eval-improve writes changes, snapshot the current version.
- Auto-rollback on regression. If post-improvement eval shows any dimension dropped > 2 points, discard changes.
- Correction tracking is non-intrusive. Record corrections in
.skill-compass/{name}/corrections.json, never in the skill file. - Tiered verification based on change scope:
- L0: syntax check (always)
- L1: re-evaluate target dimension
- L2: full six-dimension re-evaluation
- L3: cross-skill impact check (for composite/meta)
Security Notice
This includes read-only installed-skill discovery, optional local sidecar config reads, and local .skill-compass/ state writes.
This is a local evaluation and hardening tool. Read-only evaluation commands are the default starting point. Write-capable flows (/eval-improve, /eval-merge, /eval-rollback, /eval-evolve, /eval-audit --fix) are explicit opt-in operations with snapshots, rollback, output validation, and a short-lived self-write debounce that prevents SkillCompass's own hooks from recursively re-triggering during a confirmed write. No network calls are made. See [SECURITY.md](SECURITY.md) for the full trust model and safeguards.
Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: Evol-ai
- Source: Evol-ai/SkillCompass
- License: MIT
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet, be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.