# Context Benchmark

> Measure context preparation across fixed cases. Use when token savings, hard-budget compliance, protected-fact recall, evidence anchors, determinism, runtime, or downstream answer quality must be demonstrated with reproducible evidence rather than marketing claims.

- **Type:** Skill
- **Install:** `agentstack add skill-tikazi-tikaz-codex-context-economy-context-benchmark`
- **Verified:** Yes — security-reviewed for prompt injection and unsafe behavior
- **Seller:** [TIKAZI](https://agentstack.voostack.com/s/tikazi)
- **Installs:** 0
- **Category:** [Content & Media](https://agentstack.voostack.com/c/content-and-media)
- **Latest version:** 0.1.0
- **License:** MIT
- **Upstream author:** [TIKAZI](https://github.com/TIKAZI)
- **Source:** https://github.com/TIKAZI/TIKAZ-Codex-Context-Economy/tree/main/context-benchmark
- **Website:** https://tikazi.github.io/TIKAZ-AI-Skills/skills/context-economy/

## Install

```sh
agentstack add skill-tikazi-tikaz-codex-context-economy-context-benchmark
```

Requires the [AgentStack CLI](https://agentstack.voostack.com/docs/cli). Works with Claude Code, Cursor, and any MCP-compatible agent.

## About

# Context Benchmark

Designed, integrated, independently refactored, and continuously maintained by **TIKAZ**.

## Inputs

Accept a versioned benchmark manifest, fixed source fixtures, declared budgets, protected facts, expected anchors, optional route labels, and an optional externally scored downstream-answer rubric. Use identical inputs and settings when comparing systems.

## Workflow

Run a versioned manifest of independent cases and keep raw per-case results. Report efficiency and quality separately:

- source and packed tokens;
- final-budget compliance;
- protected-fact recall;
- evidence-anchor correctness;
- deterministic repeatability;
- preparation runtime;
- optional externally supplied answer score.
- document-route correctness, informative-visual recall, decorative/duplicate skip accuracy, and complex-table fidelity warnings for multimodal fixtures.

Do not hide failures inside averages. A smaller pack with lower fidelity is a regression, not a win. Do not claim superiority until the same files, questions, model/detail settings, budgets, and blind answer rubric are used. Use the shared CLI `benchmark` command and read `../references/benchmark-method.md` when publishing results.

## Output contract

Publish `summary.json`, `metrics.json`, raw `cases.json`, and the generated evidence card together. Keep context efficiency, exact-repeat prompt efficiency, literal fact and anchor fidelity, multimodal routing, and pending provider, vision, or downstream evidence separate; never replace them with one composite fidelity score.

## Validation and fallback

Keep failed cases visible and verify manifest version, fixture identity, budgets, settings, and denominators. Estimated tokens must be labeled estimates. If provider telemetry or blind downstream scoring is unavailable, mark it `Pending`; do not infer superiority from local fixtures.

## Example

```text
Benchmark this context workflow against the fixed manifest. Report efficiency and protected-fact recall separately, retain failed cases, and label provider-token measurements Pending.
```

From the suite directory, run `python scripts/tikaz_context.py benchmark --manifest benchmarks/manifest.json --output `. Inspect both `summary.json` and `cases.json`; a passing case is not evidence of positive savings or semantic equivalence.

## Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

- **Author:** [TIKAZI](https://github.com/TIKAZI)
- **Source:** [TIKAZI/TIKAZ-Codex-Context-Economy](https://github.com/TIKAZI/TIKAZ-Codex-Context-Economy)
- **License:** MIT
- **Homepage:** https://tikazi.github.io/TIKAZ-AI-Skills/skills/context-economy/

Install and usage instructions live in the source repository linked above.

## Pricing

- **Free** — Free

## Security capabilities

Automated source analysis of v0.1.0 — what this tool can access:

- **Network access:** no
- **Filesystem access:** no
- **Shell / process execution:** no
- **Environment & secrets:** no
- **Dynamic code execution:** no

*"Yes" means the capability is present in the source — more access means more to trust, not that it is unsafe.*


## Versions

- **0.1.0** — security scan: passed — Imported from the upstream source.

## Links

- Listing page: https://agentstack.voostack.com/l/skill-tikazi-tikaz-codex-context-economy-context-benchmark
- Seller: https://agentstack.voostack.com/s/tikazi
- Browse the marketplace: https://agentstack.voostack.com/browse

---
Listed on AgentStack — the marketplace for AI agent skills and MCP servers. Every listing is security-reviewed. Creators keep 70%.
