AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
MCP verified MIT Self-run

Jev Harness

mcp-ismaelsoilet-jev-harness · by ismaelsoilet

Zero-dependency System One decision harness: 5 semantic gates saving frontier AI agent tokens on trivial errors & doom loops. Python + TypeScript + Rust. MCP-compatible.

— No reviews yet
0 installs
2 views
0.0% view→install

Install

$ agentstack add mcp-ismaelsoilet-jev-harness

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • ✓ Prompt-injection patterns
  • ✓ Secret / credential exfiltration
  • ✓ Dangerous shell & filesystem operations
  • ✓ Untrusted network calls
  • ✓ Known-malicious package signatures

What it can access

  • ● Network access Used
  • ✓ Filesystem access No
  • ✓ Shell / process execution No
  • ● Environment & secrets Used
  • ✓ Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/mcp-ismaelsoilet-jev-harness)

Reliability & compatibility

✓ Security review passed
0 installs to date
— no reviews yet
● today

Declared compatibility

Claude CodeClaude DesktopCursorWindsurf

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Jev Harness? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

⚡ Jev Harness: The Token Optimizer & Decision Gate for AI Coding Agents

🇬🇧 English | 🇧🇷 Português

> Stop burning 50,000 frontier tokens on missing packages, network flakiness, and circular doom loops. > > jev-harness is an ultra-fast, deterministic-first token optimizer, test failure triage gate, and semantic guardrail for AI coding agents (OpenCode, Command Code, Claude Code, Cursor, Antigravity IDE, Windsurf, Zed, and Pi). Combines instant zero-token offline deterministic heuristics ( > 🤖 Operating as an AI Agent? If you are working on this repository, read [AGENTS.md](AGENTS.md) ([Português](AGENTS.pt-BR.md)) first. If you are integrating the tool into your own project, you do not need it — use the [Universal AI Agent Integration Guide](docs/AGENTINTEGRATIONGUIDE.md) (2-minute setup, no API key required to start). > > What it is: a deterministic-first System 1.5 decision layer for the code-quality loop — test-failure triage, doom-loop breaking, completion veto, reasoning-effort governance and model routing — callable as a CLI, MCP server, git/CI hook and typed SDKs across three runtimes (Python, TypeScript, Rust). > What it is not: a coding agent; a runtime supervisor (that is Foreman); a tool-call guardrail (jev-guard); a context sieve (Winnow); or a capability router (JevRouter). See [Where it fits](#-where-it-fits-the-system-15-decision-layer).

📑 Table of contents

  • [The Problem](#-the-problem) · [How It Works](#-how-it-works-system-1--system-15--system-2) · [Where It Fits](#-where-it-fits-the-system-15-decision-layer) · [Features](#-features) · [Quickstart](#-quickstart) · [CLI](#-cli-usage) · [Astra-Jev effort governance](#-astra-jev-dynamic-reasoning-effort-governance-2026-frontier-models) · [Integrations](#-universal-agent--ide-integrations) · [SDKs](#-python-sdk) ([TS](#-typescript--javascript-sdk--cli), [Rust](#-rust-crate--standalone-cli)) · [Git & CI guardrails](#-git--cicd-guardrails) · [Economics & benchmarks](#-economics--benchmarks-september-2026-frontier) · [Architecture & roadmap](#-architecture--roadmap) · [Contributing](#-contributing--submissions) · [Release](#-multi-registry-release--synchronization-pypi-npm-cratesio) · [Changelog](#-whats-new-in-v020)

⚡ 15-Second Quickstart

Get running in seconds across any stack — zero external dependencies, zero API keys required for local deterministic mode:

# 1. Try it instantly without installing (zero token cost, zero API keys needed)
npx @ismaelsoilet/jev-harness triage --sample "ModuleNotFoundError: No module named 'pytest'"

# 2. Or install in your favorite runtime:
pip install jev-harness                 # Python (CLI + SDK)
npm install @ismaelsoilet/jev-harness   # Node.js (CLI + SDK)
cargo install jev-harness               # Rust (Standalone CLI 'jev')

# 3. Verify health and active engine mode:
jev-harness status

🎯 The Problem

When an autonomous coding agent encounters a test failure or compiler error, the standard reaction is to dump 500 lines of raw traceback into an expensive frontier reasoning model (GPT-6 Astra, Claude Fable 5.1).

| Failure Scenario | Without Jev Harness | With Jev Harness | | :--- | :--- | :--- | | Missing dependency (ModuleNotFoundError, Cannot find module, TS2307, E0463) | 💸 ~50,000 LLM tokens (estimate) (~$0.50 - $2.50) + 15s delay to output pip/npm install ... | ⚡ Jev triage (measured: 📅 Provider verification date: 2026-09-22. Model IDs, free tiers and prices change weekly — agents and engineers should re-verify them (and record their own date) if more than 30 days have passed. Step-by-step key acquisition for every provider: [Universal AI Agent Integration Guide](docs/AGENTINTEGRATIONGUIDE.md#-provider-access--api-keys). > > 🔒 Privacy: offline mode (--mock, or no credentials) makes zero network calls. Live mode transmits the typed questions and the failure log (head 2,000 + tail 4,000 characters) to the provider endpoint. Since v0.2.0 the state itself is redacted in all three runtimes: credential-shaped material (API keys, JWTs, GitHub/AWS tokens, DB URLs, private keys) is masked before it leaves the process — in every gate and over MCP/SDK alike. That is shape-based hygiene, not a data-loss-prevention layer: structured customer data that does not look like a credential is still transmitted, so use --mock for repositories with regulated data.

Credential resolution priority:

  1. Environment variables (TYPESAFE_API_KEY, CMD_API_KEY, COMMAND_CODE_API_KEY, OPENCODE_API_KEY, OPENROUTER_API_KEY, or AI_GATEWAY_API_KEY)
  2. Local repository .jev.json, .env, or ~/.commandcode/auth.json
  3. Global configuration ~/.config/jev/credentials.env
  4. Resilient Fallback: active when no credentials are configured and on retryable provider failures (429/5xx/timeouts/network) or 401/403 from any provider. Retries use capped backoff + Retry-After; the engine prints [JEV WARNING] to stderr and every degraded result is flagged is_mock=true with a degraded_reason. Use --fail-closed to surface those errors instead (CLI exit 2, no traceback).
# Check current connection & provider status anytime
jev-harness status

Repository Configuration (.jev.json)

jev-harness init scaffolds a repository-local .jev.json. Honored keys:

| Key | Type | Default | Effect | | :--- | :--- | :--- | :--- | | model | string | provider default | Overrides the model sent to the provider. The scaffold placeholder jev-latest means "use the provider-optimized default", so it never clobbers provider model IDs. | | skip_llm_threshold | float 0-1 | 0.65 | Minimum confidence for test-gate to set skip_llm=true (a deep_logic verdict is never bypassed). | | abort_threshold | float 0-1 | 0.70 | Minimum dead-end probability for abort-check to abort a trajectory. | | shadow | bool | false | Decide and report, but never change the exit code (see [Shadow Mode](#-features)). | | api_key / provider | string | — | Optional credentials. Environment variables take precedence. |

Values are clamped to [0, 1], and a corrupted file degrades to defaults instead of breaking CI. The same keys work identically in Python, TypeScript and Rust.

Pinning the model. The effective model resolves as explicit argument → JEV_MODEL env var → model in .jev.json → provider default (run jev-harness status to see both the model and where it came from). The alias jev-latest moves: the provider may change what it points to. Once your skip_llm_threshold is calibrated against a model version, pin it — for example "model": "jev-1.13.0" — so a provider release cannot silently change your decisions.


🛠️ CLI Usage

1. Test Failure Triage (test-gate)

Pipe error logs directly or pass a file:

# Pipe directly from your test runner
npm test | jev-harness test-gate
pytest | jev-harness test-gate

# Or analyze a saved log file
jev-harness test-gate --log error.log

# Or get machine-readable JSON
jev-harness test-gate --log error.log --json

# Observe decisions without blocking a pipeline (always exits 0)
jev-harness test-gate --shadow --log error.log

# Fail hard on provider outages instead of degrading to the offline engine
jev-harness test-gate --fail-closed --log error.log

# Bound how long retryable provider failures are retried (default: 3 attempts)
jev-harness test-gate --retries 1 --log error.log

Output Example:

--- JEV TEST TRIAGE VERDICT ---
Category:        ENV_MISSING
Confidence:      92.0%
Skip LLM Call:   YES (Save Tokens!)
Skip Probability: 96.0%
Severity Score:  1.0 / 4.0
Recommendation:  AUTO-ACTION: Install missing dependency or check environment configuration (Do NOT call LLM).
--------------------------------

2. Guard Against Doom Loops & Dead-Ends (abort-check)

Verify that a proposed plan isn't repeating a failed path:

jev-harness abort-check \
  --plan "Retry rewriting the entire database schema without backup" \
  --history "Attempt 1 failed with timeout. Attempt 2 failed with circular foreign key error."

Returns exit code 1 if abort is recommended, enabling automated CI stops.

3. Model Tier Routing (route)

Pick the cheapest model capable of solving the task:

jev-harness route --task "Fix typo in docstring and reformat with black"
# -> TIER: DETERMINISTIC | Model: Direct Python/Bash Script (0 LLM Tokens)

jev-harness route --task "Refactor distributed actor supervision tree across 14 modules"
# -> TIER: HEAVY_SYSTEM2 | Model: Claude Fable 5.1 / GPT-6 Astra (~$10.00 in / $50.00 out)

4. Step Completion Verification (verify)

Verify evidence against criteria with calibrated confidence:

jev-harness verify \
  --criteria "Must export format_date function and pass all 10 unit tests" \
  --output "All 10 unit tests passed in 0.02s. format_date exported in index.ts."

5. Dynamic Reasoning Effort Governance (reasoning-effort / astra-jev)

Dynamically modulate reasoning effort per-generation (inspired by Vechen @miu21590) to eliminate latency and save thousands of tokens on mechanical tool steps:

# Evaluate immediate step for DeepSeek (e.g. DeepSeek V4.1-Flash / V4-Pro)
jev-harness reasoning-effort \
  --context "git status e verificar arquivos alterados no commit recente" \
  --target-provider deepseek

# Output:
# Effort: LOW | Dialect: {"extra_body": {"thinking": {"type": "enabled"}}, "reasoning_effort": "low"}
# Latency eliminated: ~200s internal CoT reduced to 1.5s!

# Evaluate architectural task for Anthropic (Claude Fable 5.1 / Claude Opus 5)
jev-harness reasoning-effort \
  --context "Architect distributed actor supervision tree with raft consensus" \
  --target-provider anthropic --json

# Safeguard check for direct models (returns empty params and warnings for non-reasoning models)
jev-harness reasoning-effort \
  --context "Run bash command" \
  --target-provider openai \
  --model gpt-5.6-luna

6. Continuation Nudge Gate (nudge-gate / nudge)

Inspired by CommandCodeAI/cmd-mod-jev-nudge, nudge-gate combines gated workflow phases (research, ask, plan, execute, verify, complete) with calibrated Noul probabilities (nudge, waiting, progress) to evaluate whether an autonomous agent paused prematurely with unfinished work or unverified changes (should_nudge = true, exit code 0), while automatically vetoing nudges when waiting on user input (waiting >= 0.5 or phase == "ask"), when the previous nudge produced no progress (`progress shouldnudge: true | workflow_phase: "verify" | exit code 0

Evaluate when waiting on user choice (vetoed automatically)

jev-harness nudge-gate \ --transcript "Assistant: Which AWS region should I deploy to? Would you like me to proceed?"

-> shouldnudge: false | workflowphase: "ask" | exit code 1


### 7. ROI & Token Savings Telemetry (`metrics`)
Inspect cumulative tokens saved, dollars saved, and doom loops intercepted:

```bash
# View active telemetry
jev-harness metrics

# Reset session telemetry counters
jev-harness metrics --reset

Output Example:

============================================================
              JEV HARNESS TELEMETRY & ROI
============================================================
Total Triage Interceptions:      14 calls
LLM Frontier Calls Skipped:      11 calls (78.6%)
Abort Guard Stops Triggered:     2 doom loops killed
Deterministic Routes:            6 tasks
Reasoning Effort Modulations:    8 steps (6 low, 2 high)
Estimated Tokens Saved:          422,200 tokens (heuristic estimate)
Estimated Frontier Dollars Saved: $6.12 USD (heuristic estimate)
Assumption Model:                26,200 tokens/$0.31 per intercepted triage; 80,000 tokens/$1.20 per aborted doom loop
============================================================

> 📊 These figures are a planning estimate, not metered usage. The per-event assumptions are fixed constants (26,200 tokens/$0.31 per intercepted triage, 80,000 tokens/$1.20 per aborted loop). --json exposes estimates_are_heuristic: true so downstream tooling can label them correctly.

8. Self-diagnosis, audit trail and calibration (doctor / receipts / replay)

# Is my installation healthy? (never prints secrets; --live spends ONE request)
jev-harness doctor
jev-harness doctor --live --json

# What did this repository decide? (append-only, hashes + metadata only)
jev-harness receipts --tail 10
jev-harness receipts --json

# How accurate are the gates? (confusion matrix, P/R/F1, ECE per gate; fails on regression)
jev-harness replay --corpus tests/corpus

7. One-Command Agent Setup (init)

Automatically scaffold MCP configurations for your active agent or IDE:

# Setup for Cursor
jev-harness init --cursor

# Setup for Antigravity IDE
jev-harness init --antigravity

# Setup git pre-commit hook (detects npm/pytest/cargo; never overwrites an existing hook)
jev-harness init --git

# Setup all supported tools at once
jev-harness init --all

⚡ Astra-Jev: Dynamic Reasoning Effort Governance (2026 Frontier Models)

Inspired by Vechen's (@miu21590) groundbreaking work on Astra-Codex and the Astra-Ares framework, Astra-Jev introduces autonomous, per-generation reasoning effort modulation governed by TypeSafe Jev System One.

Instead of locking an entire multi-turn coding session into heavy, slow reasoning (or risking bugs by running exclusively in low reasoning), Astra-Jev evaluates the cognitive demand of the immediate next generation in **(GPT-6 Astra, Claude Fable 5.1) | Dollar Cost($10/1M in, $50/1M out) | Agent burns ~8,000 reasoning tokens ($0.40 - $1.20) just to inspect git status or read a file | Injects effort="low", burning only ~300 tokens. Saves up to $1.15 per mechanical generation. | | Chinese Frontier(DeepSeek V4.1-Flash, Qwen 3.8 Max, Kimi-k3, MiMo) | Latency & GPU Starvation(Tokens are cheap, but internal CoT takes 3–5 minutes) | Agent enters 200–300 second internal thinking loop before running a trivial bash command | Disables thinking CoT or sets effort="low". Response delivered in 1.5s instead of 240s. |

🛡️ Critical Safeguards Built into Astra-Jev

  1. Direct Single-Pass Model Safeguard: Models that do not support internal reasoning (e.g. gpt-5.6-luna, gemini-3.8-live, claude-3.5-haiku) will return fatal HTTP 400 Bad Request if reasoning parameters are injected. Astra-Jev automatically detects non-reasoning targets, sets is_reasoning_supported = False, and returns clean empty payloads {}.
  2. Preservation of reasoning_content (DeepSeek multi-turn): In DeepSeek V4.1-Flash/Pro APIs, stripping reasoning_content across multi-turn tool calling can corrupt tool execution. Astra-Jev enforces dialect compliance to preserve thinking structures across turn transitions.
  3. Prompt Cache (KV Cache) Trade-off Advisory: Toggling reasoning parameters back-and-forth mid-session can invalidate prefix cache on long contexts (>100k tokens). Astra-Jev provides cache_safe_recommendation advisories:
  • For pure mechanical actions, use Jev's skip_llm=true to execute directly without calling the LLM at all.
  • Keep reasoning effort stable across related sub-steps of a single complex implementation.

🤖 Universal Agent & IDE Integrations

> 📖 Looking for a turnkey setup for any project? Read the [Universal AI Agent Integration Guide](docs/AGENTINTEGRATIONGUIDE.md) (

…

Source & license

This open-source MCP server is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.