Install
$ agentstack add skill-andregusman-raiz-a-gusman-claude-ag-criar-cenario-benchmark ✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ✓ Network access No
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
Verified badge
Passed review? Show it. Paste this badge into your README, it links to the public security report.
Reliability & compatibility
Declared compatibility
Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.
We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.
How agent discovery & health will work →About
ag-criar-cenario-benchmark — QAT-Benchmark Scenario Designer (INTERNAL)
> Internal skill — invocada automaticamente por ag-4-teste-final benchmark via Agent tool. Não use diretamente via /. Para criar cenário benchmark use /ag-4-teste-final benchmark [url].
Papel
O designer de cenarios de benchmark: cria cenarios de alta qualidade para o QAT-Benchmark (ag-benchmark-qualidade), seguindo metodologia de 8 dimensoes, anti-contaminacao e criterios por camada (L1-L4).
Diferenca de ag-criar-cenario-qat: ag-criar-cenario-qat cria cenarios de qualidade ABSOLUTA (QAT). ag-criar-cenario-benchmark cria cenarios de benchmark COMPARATIVO (app vs baseline). Diferenca de ag-criar-cenario-ux-qat: ag-criar-cenario-ux-qat cria cenarios visuais/UI. ag-criar-cenario-benchmark cria cenarios de conteudo/AI. Diferenca de ag-benchmark-qualidade: ag-benchmark-qualidade EXECUTA cenarios. ag-criar-cenario-benchmark CRIA cenarios para ag-benchmark-qualidade executar.
Invocacao
/ag-criar-cenario-benchmark capability="tool use" # 5 rotatable scenarios
/ag-criar-cenario-benchmark capability="reasoning" count=10 # 10 scenarios
/ag-criar-cenario-benchmark capability="teaching" category=fixed # Fixed scenarios
/ag-criar-cenario-benchmark capability="safety" domain="matematica 8o ano" # Domain-specific
Pre-requisitos
- Estrutura
tests/qat-benchmark/scenarios/no projeto - Cenarios existentes para analise de cobertura
Output
- Arquivos TypeScript em
scenarios/fixed/ouscenarios/rotatable/ - Cada cenario com: ID, prompt, dimensoes-alvo, criterios L1-L4, functionalChecks
- Fixed: BM-XX (sequencial, baseline tracking)
- Rotatable: BM-RXXX (pool grande, anti-contaminacao)
Anti-contaminacao
- Fixed (30%): pool pequeno (12-15), NUNCA modificar existentes
- Rotatable (70%): pool grande (50+), variar complexidade/dominio/formato
Interacao com outros agentes
- ag-benchmark-qualidade: Complementar (ag-criar-cenario-benchmark cria, ag-benchmark-qualidade executa)
- ag-criar-cenario-qat: Paralelo (ag-criar-cenario-qat cria QAT, ag-criar-cenario-benchmark cria benchmark)
Referencia
- Agent completo:
~/.claude/agents/ag-criar-cenario-benchmark.md - Patterns:
~/.claude/shared/patterns/qat-benchmark.md - Templates:
~/.claude/shared/templates/qat-benchmark/scenarios/
Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: andregusman-raiz
- Source: andregusman-raiz/a-gusman-claude
- License: MIT
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet, be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.