# Deepseek Harness

> Harness for DeepSeek V4-Pro / V4-Flash. Python lib (pip install deepseek-harness) + dsh CLI + MCP server (npx @deepseek-harness/mcp) + Anthropic SKILL.md. 16 documented protocol quirks, 12 probes, 270+ trials.

- **Type:** MCP server
- **Install:** `agentstack add mcp-henryz838978-deepseek-harness`
- **Verified:** Yes — security-reviewed for prompt injection and unsafe behavior
- **Seller:** [HenryZ838978](https://agentstack.voostack.com/s/henryz838978)
- **Installs:** 0
- **Category:** [AI & ML](https://agentstack.voostack.com/c/ai-and-ml)
- **Latest version:** 0.1.0
- **License:** MIT
- **Upstream author:** [HenryZ838978](https://github.com/HenryZ838978)
- **Source:** https://github.com/HenryZ838978/deepseek-harness
- **Website:** https://github.com/HenryZ838978/deepseek-harness

## Install

```sh
agentstack add mcp-henryz838978-deepseek-harness
```

Requires the [AgentStack CLI](https://agentstack.voostack.com/docs/cli). Works with Claude Code, Cursor, and any MCP-compatible agent.

## About

# `deepseek-harness`

### Protocol-aware adapters for DeepSeek V4-Pro and V4-Flash

[](https://pypi.org/project/deepseek-harness/)
[](https://npmjs.com/package/@deepseek-harness/mcp)
[](packages/skill/SKILL.md)
[](reports/probes/)
[](reports/REPORT_2026-05-09.md)
[](spec/06_context_limits.md)
[](spec/04_cache_hit.md)
[](LICENSE)

A single protocol contract distributed in four wrapper formats. Designed to meet the integration requirements of any OpenAI-compatible client.

---

## Status

| Form | Status | Distribution channel |
|---|---|---|
| Python library `deepseek-harness` | published `0.2.0` | https://pypi.org/project/deepseek-harness/ |
| Command-line tool `deepseek-harness-cli` | published `0.2.0` | https://pypi.org/project/deepseek-harness-cli/ |
| MCP server `@deepseek-harness/mcp` | published `0.2.0` | https://www.npmjs.com/package/@deepseek-harness/mcp |
| Anthropic Skill | source ready | (see [`packages/skill/SKILL.md`](packages/skill/SKILL.md)) |

## Installation

```bash
pip install deepseek-harness                  # Python library
pip install deepseek-harness-cli              # `dsh` command-line tool
npx -y @deepseek-harness/mcp                  # MCP server (stdio transport)
```

For zero-dependency integration:

```bash
curl -sL https://raw.githubusercontent.com/HenryZ838978/deepseek-harness/main/packages/skill/scripts/safe_init.py -o safe_init.py
```

For Anthropic Skill-aware agents:

```bash
git clone https://github.com/HenryZ838978/deepseek-harness && \
cp -r deepseek-harness/packages/skill ~/.claude/skills/deepseek-harness
```

All five paths derive from the same `spec/` source of truth. Behaviour is identical across forms.

---

## Architecture

```mermaid
flowchart LR
    classDef spec fill:#fef3c7,stroke:#f59e0b,color:#78350f
    classDef pkg  fill:#e0e7ff,stroke:#6366f1,color:#312e81
    classDef out  fill:#d1fae5,stroke:#10b981,color:#064e3b

    SPEC["spec/10 contract rulesRFC 2119 normative"]:::spec

    CORE["packages/coreDeepSeekHarness"]:::pkg
    CLI["packages/clidsh"]:::pkg
    MCP["packages/mcpTypeScript stdio"]:::pkg
    SKILL["packages/skillSKILL.md + scripts/"]:::pkg

    PIP["pip installdeepseek-harness"]:::out
    PIPCLI["pip installdeepseek-harness-cli"]:::out
    NPM["npx -y@deepseek-harness/mcp"]:::out
    DROP["~/.claude/skills/drop-in"]:::out

    SPEC --> CORE
    SPEC --> CLI
    SPEC --> MCP
    SPEC --> SKILL

    CORE  --> PIP
    CLI   --> PIPCLI
    MCP   --> NPM
    SKILL --> DROP
```

---

## Compatibility matrix

| Environment                                                | Recommended form                               | Verification command          |
|------------------------------------------------------------|------------------------------------------------|-------------------------------|
| Python projects (LangChain, LlamaIndex, custom agents)     | `pip install deepseek-harness`                 | `python -c "import deepseek_harness"` |
| Command-line / debugging / CI                              | `pip install deepseek-harness-cli`             | `dsh doctor`                  |
| MCP-aware desktop clients (Claude Desktop, Cline, Roo Code, ChatWise, Cherry Studio) | `npx -y @deepseek-harness/mcp` | configure `mcpServers` in client |
| Anthropic Skill-aware agents (Claude Code)                 | drop `packages/skill/` into `~/.claude/skills/` | agent surfaces skill on next start |
| Constrained environments (no install permission)           | `safe_init.py` zero-dependency snippet         | `python safe_init.py`         |

---

## Background

DeepSeek V4-Pro and V4-Flash expose an OpenAI-compatible HTTP API. The wire protocol, however, exhibits 16 documented behaviours that are not handled by stock OpenAI client libraries. These include:

- Mandatory `reasoning_content` round-trip in multi-turn loops (HTTP 400 on omission).
- Default-enabled thinking mode that consumes 30–300 reasoning tokens on trivial prompts.
- Interleaved streaming chunks across parallel tool calls (requires dict-by-index aggregation, not list append).
- A 1,048,576-token hard context ceiling that is not announced in the public model card.
- A prefix cache that grants a 50× cost discount on hits but invalidates on prefix mutation.

The goal of this repository is to characterize each behaviour with a reproducible probe, codify the resulting contract in `spec/`, and ship reference implementations of that contract in the four most common distribution formats.

---

## Without and with the harness

```mermaid
sequenceDiagram
    autonumber
    participant App as Agent application
    participant SDK as openai SDK
    participant DS as DeepSeek V4

    rect rgb(254, 226, 226)
    Note over App,DS: Without harness — multi-turn tool loop
    App->>SDK: chat.completions.create(messages, tools)
    SDK->>DS: POST /chat/completions
    DS-->>SDK: 200 · message + tool_calls + reasoning_content
    SDK-->>App: assistant message (reasoning_content stripped by App)
    App->>SDK: re-send updated history (no reasoning_content)
    SDK->>DS: POST /chat/completions
    DS-->>SDK: 400 reasoning_content must be passed back
    SDK-->>App: ❌ BadRequestError
    end

    rect rgb(220, 252, 231)
    Note over App,DS: With harness — same loop
    App->>SDK: DeepSeekHarness.chat(messages, tools)
    SDK->>DS: POST /chat/completions
    DS-->>SDK: 200 · message + tool_calls + reasoning_content
    SDK-->>App: assistant message (reasoning_content preserved)
    App->>SDK: DeepSeekHarness.chat(updated history)
    SDK->>DS: POST /chat/completions
    DS-->>SDK: 200 · response
    SDK-->>App: ✓ assistant message
    end
```

---

## Cache discount in practice

Cache hit progression observed across a five-turn conversation (probe_10 / S1 on V4-Pro). Each turn appends to the same prefix; cache miss decreases monotonically until the prefix exceeds the 1,024-token activation threshold and the 256-token block boundaries align.

```mermaid
xychart-beta
    title "Cache hit ratio over five conversation turns (probe_10/S1, V4-Pro)"
    x-axis "Turn" [0, 1, 2, 3, 4]
    y-axis "Cache hit ratio" 0 --> 1
    bar [0, 0.56, 0.72, 0.78, 0.95]
```

At the equilibrium hit ratio of 95%, input cost is reduced by a factor of approximately 50 relative to a cache-miss workload (`$0.0028/M` vs `$0.14/M` for V4-Flash input pricing).

---

## Findings summary

A summary of all 16 documented findings, each linked to the probe that produced it.

Click to collapse

| # | Finding | Reference (community) | Empirical result (V4-Pro / V4-Flash) | Probe |
|---|---|---|---|---|
| 1  | `thinking=enabled` is the default on V4-Pro/Flash | undocumented | reproduced (~30 reasoning tokens on trivial prompts) | smoke |
| 2  | Each streamed response contains ~3 chunks with empty `choices` | cline #1594 | reproduced / reproduced | probe_1 |
| 3  | Multi-turn assistant→tool messages must echo `reasoning_content` | agent-framework #5538 | 3/3 reproduce 400 / 3/3 reproduce 400 | probe_2 |
| 4  | Parallel `tool_call` deltas are interleaved across `tc.index` | none | 3/3 (Pro 30 chunks · Flash 38 chunks) | probe_7 |
| 5  | `length` cut on a thinking-on tool call returns empty content and empty `tool_calls` | none | reproduced | probe_8 |
| 6  | Tool-call payload leaked into `content` (community: ~11% on V3) | DeepSeek-V3 #1244 | 0/50 / 0/50 (apparent fix in V4) | probe_3 / 3b |
| 7  | `strict: true` produced corrupt JSON (community status: WONTFIX) | DeepSeek-V3 #1069 | 0/32 / 0/32 (apparent fix in V4) | probe_4 |
| 8  | `/beta` endpoint silently remaps `v4-pro` to `deepseek-reasoner` | none | reproduced | probe_4 |
| 9  | Cache-hit token field uses both DeepSeek-native and OpenAI-shape names | pi-mono #3880 | both fields populated | probe_5 |
| 10 | Mid-prefix character mutation preserves the first 512 cached tokens | none | reproduced (256-token block alignment) | probe_5 |
| 11 | Cache eviction is observable across otherwise identical requests | none | reproduced (S1#3 returned 0% hit) | probe_5 |
| 12 | Hard context ceiling = 2²⁰ = 1,048,576 tokens | none | reproduced (verbatim 400 includes byte count) | probe_6b |
| 13 | Long `reasoning_content` may exceed downstream V8 string limit | community screenshots | partial (V4 reasoning bounded ≤ 26 KB) | probe_9 |
| 14 | SSE chunk granularity is 1–3 characters, producing thousands of chunks per response | none | reproduced (7,941 chunks on 26 KB response) | probe_9 |
| 15 | Five-turn agentic loop succeeds when contract rules are followed | (refutes broad community claim) | 15/15 turns successful | probe_10 |
| 16 | V4-Flash protocol contract is identical to V4-Pro | none | confirmed across all probes | probe_11 |

The full numerical detail is in [`reports/REPORT_2026-05-09.md`](reports/REPORT_2026-05-09.md). The paper-style write-up with reproducibility instructions is in [`docs/technical_report.md`](docs/technical_report.md).

---

## Contract specification

Ten normative rules derived from the findings above. Each rule maps to a section of [`spec/`](spec/).

| ID  | Rule                                                                                | Spec section                                |
|-----|-------------------------------------------------------------------------------------|---------------------------------------------|
| C1  | Disable thinking by default; enable explicitly when reasoning is required           | [§1](spec/01_reasoning_content.md)          |
| C2  | Preserve `reasoning_content` on assistant messages within a tool-use loop           | [§1](spec/01_reasoning_content.md)          |
| C3  | Set `max_tokens` on every request; default to 4096                                  | [§5/§6](spec/05_streaming_finish_reason.md) |
| C4  | Aggregate parallel `tool_call` deltas by `tc.index`, not by list position           | [§5](spec/05_streaming_finish_reason.md)    |
| C5  | Use list buffers and `"".join()` for streaming content; avoid string concatenation  | [§5](spec/05_streaming_finish_reason.md)    |
| C6  | Tolerate stream chunks where `choices` is empty                                     | [§5](spec/05_streaming_finish_reason.md)    |
| C7  | Validate `prompt_tokens + max_tokens ≤ 1,048,576` before sending                    | [§6](spec/06_context_limits.md)             |
| C8  | Avoid injecting volatile content into the cached prefix                             | [§4](spec/04_cache_hit.md)                  |
| C9  | Do not route to the `/beta` endpoint when tool calls are involved                   | [§3](spec/03_strict_mode.md)                |
| C10 | `strict: true` is empirically valid on V4; continue to perform schema validation post-hoc | [§3](spec/03_strict_mode.md)            |

The harness enforces all ten rules by default. Each rule may be disabled individually for diagnostic purposes via constructor flags on `DeepSeekHarness`.

---

## Repository layout

```
deepseek-harness/
├── packages/
│   ├── core/         Python library                · pip install deepseek-harness
│   ├── cli/          dsh command-line tool         · pip install deepseek-harness-cli
│   ├── mcp/          TypeScript MCP server         · npx @deepseek-harness/mcp
│   └── skill/        Anthropic SKILL.md            · drop into ~/.claude/skills/
├── spec/             Six chapters of normative protocol contract
├── reports/          12 probes, 270+ trial JSONL fixtures, 16 finding summaries
└── docs/             Paper-style technical report and machine-readable trust ledger
```

---

## Five-year wrapper-protocol timeline

The same underlying contract — Markdown documentation, executable scripts, and structured configuration — has been repackaged under successive wrapper protocols over the past five years. This repository ships all four currently active formats from a single source.

```mermaid
timeline
    2022-2023 : Prompt templates
              : LangChain · DSPy
    2023-2024 : Command-line tools
              : OpenAI Functions · openai-python
    2024-2025 : Model Context Protocol
              : MCP servers · Claude Desktop
    2025-2026 : Anthropic Skills
              : SKILL.md · Claude Code
    2026-     : Harness (this work)
              : All four formats from one spec
```

Across all four generations the durable asset is the contract specification, not the wrapper format.

---

## Quick reference per form

Python library

```python
from deepseek_harness import DeepSeekHarness, estimate_cache_hit

client = DeepSeekHarness(disable_thinking_by_default=True)
response = client.chat(
    model="deepseek-v4-pro",
    messages=[{"role": "user", "content": "Hello"}],
    max_tokens=4096,
)
print(response["message"]["content"])
print(f"cost: ${response['usage']['estimated_cost_usd']:.6f}")
print(f"cache hit ratio: {response['usage']['cache_hit_rate']:.0%}")
```

MCP server (Claude Desktop, Cline, Roo Code, ChatWise, Cherry Studio)

Add the following to the client's MCP configuration:

```json
{
  "mcpServers": {
    "deepseek-harness": {
      "command": "npx",
      "args": ["-y", "@deepseek-harness/mcp"],
      "env": { "DEEPSEEK_API_KEY": "sk-..." }
    }
  }
}
```

The server exposes four tools: `deepseek_chat`, `deepseek_chat_stream`, `validate_message_history`, `estimate_cache_hit`. The latter two perform contract validation without consuming API quota.

Anthropic Skill (Claude Code and SKILL.md-aware agents)

```bash
git clone https://github.com/HenryZ838978/deepseek-harness
cp -r deepseek-harness/packages/skill ~/.claude/skills/deepseek-harness
```

The skill is automatically surfaced when the conversation references DeepSeek. It includes the ten contract rules, the bundled `safe_init.py`, and a compact reference card of the 16 findings.

Command-line tool

```bash
pip install deepseek-harness-cli
export DEEPSEEK_API_KEY=sk-...

dsh doctor                       # verify environment, single-token live call
dsh chat                         # interactive REPL with all guards enabled
dsh chat -r                      # enable thinking mode
dsh validate path/to/msgs.json   # offline contract audit
dsh estimate path/to/msgs.json   # offline cache-hit estimate
dsh probe probe_2 --n 3          # run a probe by name
```

Zero-dependency snippet

```bash
curl -sL https://raw.githubusercontent.com/HenryZ838978/deepseek-harness/main/packages/skill/scripts/safe_init.py -o safe_init.py
```

```python
from safe_init import safe_deepseek_call

response = safe_deepseek_call(
    messages=[{"role": "user", "content": "hello"}],
    model="deepseek-v4-flash",
    max_tokens=2048,
)
print(response["content"])
```

Single Python file, ~200 lines, depending only on the `openai` SDK. Implements all ten contract rules.

---

## Acid test

Two complementary commands establish that the harness performs a non-trivial transformation:

```bash
# 1. Reproduce the underlying protocol error using a stock OpenAI client.
python reports/probes/probe_2_reasoning_lifecycle.py --n 3
# Expected: 3 of 3 phase-B trials return BadRequestError with the message
#   "The reasoning_content in the thinking mode must be passed back to the API."

# 2. Submit the same scenario through the harness.
dsh doctor
# Expected: green status table; live call cost ≈ $0.000002 USD.
```

A regression in the first command would indicate that DeepSeek has revised the contract; the spec should be updated accordingly.

---

## Further reading

- [`reports/REPORT_2026-05-09.md`](reports/REPORT_2026-05-09.md) — full audit report (Chinese, 270+ trials)
- [`docs/technical_report.md`](docs/technical_report.md) — paper-style technical report (English)
- [`spec/00_overview.md`](spec/00_overview.md) — RFC 2119 protocol contract index
- [`docs/trust_ledger.yaml`](docs/trust_ledger.yaml) — machine-readable repository metadata
- [`docs/blog/2026-05-09-deepseek-v4-bug-tour.md`](docs/blog/2026-05-09-deepseek-v4-bug-tour.md) — narrative companion piece

---

## Naming and visual identity

The project name combines two visual references. _Harness_

…

## Source & license

This open-source MCP server is cataloged on AgentStack and links to its original source — we do not rehost the code.

- **Author:** [HenryZ838978](https://github.com/HenryZ838978)
- **Source:** [HenryZ838978/deepseek-harness](https://github.com/HenryZ838978/deepseek-harness)
- **License:** MIT
- **Homepage:** https://github.com/HenryZ838978/deepseek-harness

Install and usage instructions live in the source repository linked above.

## Pricing

- **Free** — Free

## Security capabilities

Automated source analysis of v0.1.0 — what this tool can access:

- **Network access:** yes
- **Filesystem access:** no
- **Shell / process execution:** no
- **Environment & secrets:** no
- **Dynamic code execution:** no

*"Yes" means the capability is present in the source — more access means more to trust, not that it is unsafe.*


## Versions

- **0.1.0** — security scan: passed — Imported from the upstream source.

## Links

- Listing page: https://agentstack.voostack.com/l/mcp-henryz838978-deepseek-harness
- Seller: https://agentstack.voostack.com/s/henryz838978
- Browse the marketplace: https://agentstack.voostack.com/browse

---
Listed on AgentStack — the marketplace for AI agent skills and MCP servers. Every listing is security-reviewed. Creators keep 70%.
