# Ag Criar Cenario Benchmark

> [INTERNAL — invocada via ag-4-teste-final benchmark] QAT-Benchmark Scenario Designer — cria cenarios de benchmark com dual-run, 8 dimensoes, anti-contaminacao e criterios L1-L4 por dimensao.

- **Type:** Skill
- **Install:** `agentstack add skill-andregusman-raiz-a-gusman-claude-ag-criar-cenario-benchmark`
- **Verified:** Yes — security-reviewed for prompt injection and unsafe behavior
- **Seller:** [andregusman-raiz](https://agentstack.voostack.com/s/andregusman-raiz)
- **Installs:** 0
- **Category:** [Agent Skills](https://agentstack.voostack.com/c/agent-skills)
- **Latest version:** 0.1.0
- **License:** MIT
- **Upstream author:** [andregusman-raiz](https://github.com/andregusman-raiz)
- **Source:** https://github.com/andregusman-raiz/a-gusman-claude/tree/main/skills/ag-criar-cenario-benchmark

## Install

```sh
agentstack add skill-andregusman-raiz-a-gusman-claude-ag-criar-cenario-benchmark
```

Requires the [AgentStack CLI](https://agentstack.voostack.com/docs/cli). Works with Claude Code, Cursor, and any MCP-compatible agent.

## About

# ag-criar-cenario-benchmark — QAT-Benchmark Scenario Designer (INTERNAL)

> **Internal skill** — invocada automaticamente por `ag-4-teste-final benchmark` via Agent tool. Não use diretamente via `/`. Para criar cenário benchmark use `/ag-4-teste-final benchmark [url]`.

## Papel

O designer de cenarios de benchmark: cria cenarios de alta qualidade para o QAT-Benchmark (ag-benchmark-qualidade), seguindo metodologia de 8 dimensoes, anti-contaminacao e criterios por camada (L1-L4).

Diferenca de ag-criar-cenario-qat: ag-criar-cenario-qat cria cenarios de qualidade ABSOLUTA (QAT). ag-criar-cenario-benchmark cria cenarios de benchmark COMPARATIVO (app vs baseline).
Diferenca de ag-criar-cenario-ux-qat: ag-criar-cenario-ux-qat cria cenarios visuais/UI. ag-criar-cenario-benchmark cria cenarios de conteudo/AI.
Diferenca de ag-benchmark-qualidade: ag-benchmark-qualidade EXECUTA cenarios. ag-criar-cenario-benchmark CRIA cenarios para ag-benchmark-qualidade executar.

## Invocacao

```
/ag-criar-cenario-benchmark capability="tool use"                          # 5 rotatable scenarios
/ag-criar-cenario-benchmark capability="reasoning" count=10                # 10 scenarios
/ag-criar-cenario-benchmark capability="teaching" category=fixed           # Fixed scenarios
/ag-criar-cenario-benchmark capability="safety" domain="matematica 8o ano" # Domain-specific
```

## Pre-requisitos

1. Estrutura `tests/qat-benchmark/scenarios/` no projeto
2. Cenarios existentes para analise de cobertura

## Output

- Arquivos TypeScript em `scenarios/fixed/` ou `scenarios/rotatable/`
- Cada cenario com: ID, prompt, dimensoes-alvo, criterios L1-L4, functionalChecks
- Fixed: BM-XX (sequencial, baseline tracking)
- Rotatable: BM-RXXX (pool grande, anti-contaminacao)

## Anti-contaminacao

- Fixed (30%): pool pequeno (12-15), NUNCA modificar existentes
- Rotatable (70%): pool grande (50+), variar complexidade/dominio/formato

## Interacao com outros agentes

- ag-benchmark-qualidade: Complementar (ag-criar-cenario-benchmark cria, ag-benchmark-qualidade executa)
- ag-criar-cenario-qat: Paralelo (ag-criar-cenario-qat cria QAT, ag-criar-cenario-benchmark cria benchmark)

## Referencia

- Agent completo: `~/.claude/agents/ag-criar-cenario-benchmark.md`
- Patterns: `~/.claude/shared/patterns/qat-benchmark.md`
- Templates: `~/.claude/shared/templates/qat-benchmark/scenarios/`

## Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

- **Author:** [andregusman-raiz](https://github.com/andregusman-raiz)
- **Source:** [andregusman-raiz/a-gusman-claude](https://github.com/andregusman-raiz/a-gusman-claude)
- **License:** MIT

Install and usage instructions live in the source repository linked above.

## Pricing

- **Free** — Free

## Security capabilities

Automated source analysis of v0.1.0 — what this tool can access:

- **Network access:** no
- **Filesystem access:** no
- **Shell / process execution:** no
- **Environment & secrets:** no
- **Dynamic code execution:** no

*"Yes" means the capability is present in the source — more access means more to trust, not that it is unsafe.*


## Versions

- **0.1.0** — security scan: passed — Imported from the upstream source.

## Links

- Listing page: https://agentstack.voostack.com/l/skill-andregusman-raiz-a-gusman-claude-ag-criar-cenario-benchmark
- Seller: https://agentstack.voostack.com/s/andregusman-raiz
- Browse the marketplace: https://agentstack.voostack.com/browse

---
Listed on AgentStack — the marketplace for AI agent skills and MCP servers. Every listing is security-reviewed. Creators keep 70%.
