Install
$ agentstack add mcp-microsoft-thinkingbox ✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ✓ Network access No
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
Verified badge
Passed review? Show it. Paste this badge into your README, it links to the public security report.
Reliability & compatibility
Declared compatibility
Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.
We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.
How agent discovery & health will work →About
ThinkingBox
Introduction
ThinkingBox is a framework designed to:
- Define tool mocks as MCP servers.
- Create scenarios and test cases.
- Run an LLM agent and enable interaction with tools.
- Evaluate the outcomes of agent execution.
It can be used to:
- Generate conversations for "offline" LLM training or evaluation.
- Train LLMs with reinforcement learning, using the entire system in the training loop.
It supports:
- Spawning and initializing multiple isolated tool execution environments.
- LLM Agent loop with tool use.
- Interaction with a simulated (LLM) User that responds based on a prompt with added context.
Repositories
ThinkingBox is split across two repositories:
- thinkingbox (this repo) — the
framework: the tb CLI, the MCP Session Proxy, the agent/user/judge loop, and the evaluation harness. Ships a single bundled scenario (cloud_drive) and its matching MCP server (mcp_cloud_drive.py) only as an offline smoke-test for an install — see [Verify install](#verify-install).
- thinkingbox-data — the
curated datasets, the MCP tool server packages (under servers/, e.g. thinkingbox_tools, ms_toloka_servers), and supporting data files (embeddings, knowledge bases, etc.) under support/. This is where real scenarios, test cases, and tools live; clone it for any non-trivial work.
For tutorial-style worked examples (running scenarios, batch evaluation, interactive chat against real datasets), see the thinkingbox-data README. This README focuses on the framework itself — install, architecture, and dataset format reference.
Setup
ThinkingBox is only tested on Linux. Most of it might work on other systems but we only target Linux (including WSL) at the moment.
Clone
git clone https://github.com/microsoft/thinkingbox.git
git clone https://github.com/microsoft/thinkingbox-data.git
# OR using GitHub CLI
gh repo clone microsoft/thinkingbox
gh repo clone microsoft/thinkingbox-data
Development Setup
We recommend using uv for python management, and python version 3.12.
# Install thinkingbox in editable mode with dev dependencies
uv venv --python 3.12
uv sync --group dev
# (for contributors) Install pre-commit hooks
uv run pre-commit install
Note that this creates a virtual environment in .venv which uv uses by default when running from this directory (e.g. uv run). Just manually activate this virtual environment in order to use it elsewhere.
source .venv/bin/activate
If use of uv is restricted in your environment, you can use pip as an alternative.
python -m venv .venv
source .venv/bin/activate
pip install -U pip setuptools
pip install --config-settings editable-mode=compat -e '.[dev]'
pre-commit install
Pre-commit hooks
Pre-commit hooks are configured to automatically format code on commit (check previous section).
This project uses Black for code formatting. A pre-commit hook automatically re-formats changed files on commit. A PR pipeline also enforces that code is correctly formatted before merging into main.
To use pre-commit manually (including the black re-formatter):
# run pre-commit hooks once on modified files
uv run pre-commit run
# run pre-commit hooks once on all files
uv run pre-commit run --all-files
CLI overview
All actions in ThinkingBox are accessed through one command: tb.
> uv run tb --help
Usage: tb [OPTIONS] COMMAND [ARGS]...
Options:
--help Show this message and exit.
Commands:
agg Aggregate metrics from a JSONL file.
dump-tests Dump test cases for a given agent and dataset.
infer Execute inference for a single test or set of tests.
mcp-start Start the MCP Session Proxy.
pp Pretty-print a decoded result from a file or stdin.
run-test Process decoder results and optionally update or write new...
sbs Compare candidate vs baseline JSONL results and report lift...
tui Launch the ThinkingBox TUI for a single test case or a...
Verify install
The framework ships a single offline scenario, cloud_drive, so you can sanity-check the install without cloning thinkingbox-data first.
In one terminal, start the Session Proxy with no --servers flag — auto-discovery picks up the bundled mcp_cloud_drive server:
uv run tb mcp-start
In another terminal, run a single bundled test:
uv run tb infer -c config/config_o4mini.yaml --dataset ./dataset --agent think \
--name cloud_drive.py:test_append_some_more_text --output output.yaml
uv run tb pp output.yaml
If tb pp shows a conversation and the assertions pass, the framework and your LLM endpoint are wired up. For real scenarios, datasets, and tool servers, see the thinkingbox-data README.
ThinkingBox unit tests
# run the session proxy with the test MCP servers configuration
uv run tb mcp-start --servers tests/servers.yaml
# run the tests
uv run pytest -v tests
Architecture
This section describes the framework's runtime pieces and their CLI-level controls. For tutorial-style worked examples, see the thinkingbox-data README.
MCP Session Proxy
ThinkingBox interacts with tools through the MCP Session Proxy — a long-running HTTP server that fronts a fleet of MCP tool processes.
It works as follows:
- TB sends a "scenario" initialization to the Session Proxy, which spawns and
initializes MCP servers as needed, creating a new isolated "session" for the current conversation.
- TB requests tool schemas from the Session Proxy.
- TB interacts with tools by sending requests to the Session Proxy.
- TB retrieves side effects from the Session Proxy for judging. This could be
any change occurring within the session resulting from tool execution.
- TB sends a destroy request to the Session Proxy, which terminates the
related MCP servers and releases the memory.
┌─────────────────────────────────────────────────────────────────────────────┐
│ tb infer (CLI Process) │
│ ───────────────────── │
│ 1. Load config, test case, scenario │
│ 2. Create LLM sessions (agent, user, judge) │
│ 3. Connect to session_proxy, create session │
│ 4. Run agent loop (decode_turn_iter) │
│ 5. Retrieve effects, run test assertions │
└──────┬─────────────────────────────────────┬────────────────────────────────┘
│ │
│ LLM API calls │ HTTP to session_proxy
│ (agent reasoning, │ (tool calls, effects)
│ user simulation, │
│ judge evaluation) │
▼ ▼
┌──────────────────┐ ┌─────────────────────────────────────────┐
│ Azure OpenAI │ │ session_proxy (:7111) │
│ or Anthropic │ │ ───────────────────── │
│ │ │ POST /session_create → spawn servers │
│ - Agent LLM │ │ POST /list_tools → get schemas │
│ - User LLM │ │ POST /call_tool → execute tool │
│ - Judge LLM │ │ POST /get_effects → retrieve state │
└──────────────────┘ │ POST /session_destroy → cleanup │
└──────────────────┬──────────────────────┘
│
│ stdio (JSON-RPC)
│ one process per server
▼
┌─────────────────────────────────────────┐
│ MCP Server Processes │
│ ─────────────────── │
│ mcp_cloud_drive.py → file storage │
│ mcp_online_banking.py → account state │
│ mcp_email_system.py → sent emails │
│ mcp_ms_store.py → store FAQ │
│ ... (more in thinkingbox-data) │
│ │
│ Each server has: │
│ - __reserved__init (setup state) │
│ - tool functions (get_accounts, etc) │
│ - __reserved__geteffects (for testing) │
└─────────────────────────────────────────┘
Start the Session Proxy with tb mcp-start. The choice of --servers controls which tool servers are loaded:
# Auto-discover the bundled servers under thinkingbox/tools/mcp_*.py
# (only mcp_cloud_drive — useful for the smoke test, nothing else)
uv run tb mcp-start
# Real workloads: point at thinkingbox-data's master servers config
uv run tb mcp-start --servers ../thinkingbox-data/servers/servers.yaml
Some tools require additional setup (running services, the THINKINGBOX_DATA environment variable for support files). See [Tools with additional setup](docs/toolswithadditional_setup.md).
Data flow: a single tool call
Agent LLM returns: ToolCall(name="get_accounts", args={})
│
▼
decode_turn_iter() calls mcp_proxy.call_tool("get_accounts", {})
│
▼
MCPProxyClient POST /call_tool ──► session_proxy
│ │
│ ▼
│ ToolDispatcher routes to server
│ │
│ ▼
│ mcp_online_banking (JSON-RPC)
│ │
│ ▼
│ get_accounts() executes
│ │
◄─────────────────────────────────┘
│ result: '{"accounts": [...]}'
▼
ToolResponse added to conversation, yielded
│
▼
Agent LLM sees tool result, continues reasoning
LLM Configuration
See [LLM Endpoint Configuration](docs/llmendpointconfig.md) for all the options.
If using Azure OpenAI endpoint, log in with azure-cli (az login) and configure some endpoints you have access to in the main configuration file. Check the example in config/config_o4mini.yaml.
If using OpenAI-Compatible deployments, check the example in config/config_vllm.yaml.
Interactive TUI
tb tui launches an interactive session to chat with a scenario or a test case. See the thinkingbox-data README for invocation examples; this section covers the TUI's UX details.
IMPORTANT: Use ESC then ENTER to submit a message, or just ENTER for newline. This is necessary for multiline input.
Note: check the prompt_toolkit documentation for more information, our instructions are Linux-specific and other platforms have different key bindings.
When prompted with [user::text], provide a user response, or one of the special commands starting with /:
# run a test from file
/test dataset/test_case/.py:
# or if chatting with a test case (--name), execute its associated test
/test
# show tool definition
/tool get_text_content
# show conversation in raw format
/conversation
# get effects/state from the server
/effects
# exit
/quit
Inspecting and aggregating results
Pretty-print individual conversations from the JSONL or YAML output of tb infer:
uv run tb pp input_file.yaml
# or (first example in a JSONL)
head -n1 input_file.jsonl | uv run tb pp
Aggregate results and statistics into a table summary from the JSONL output of tb infer:
uv run tb agg input_file.jsonl
# or (for a subset of results)
cat input_file.jsonl | grep "" | uv run tb agg
Troubleshooting
| Symptom | Cause | Fix | |---|---|---| | Port 7111 already in use | Stale proxy process | lsof -ti:7111 \| xargs kill | | ModuleNotFoundError: thinkingbox | Venv not activated | uv sync or source .venv/bin/activate | | Scenario not found | Wrong dataset path | Check -d points to ../thinkingbox-data/dataset | | 401 Unauthorized / timeout | Azure auth expired | Run az login | | FileNotFoundError: support/... | Missing data files | Set THINKINGBOX_DATA env var | | Connection refused localhost:7111 | Proxy not running | Start uv run tb mcp-start in another terminal | | test_case not found | Typo in test name | Format is filename.py:function_name | | TUI: can't submit message | Wrong key combo | Press ESC then Enter (not just Enter) | | Pre-commit fails | Formatting issues | Run uv run pre-commit run --all-files |
Documentation
Deeper references for specific topics live under [docs/](docs/):
Tutorials and authoring
- [
tutorial.md](docs/tutorial.md) — End-to-end walkthrough: create a server, a scenario, and a test case, then progressively add assertions, state, the LLM judge, the simulated user, and debugging. - [
adding_tools.md](docs/adding_tools.md) — Production-grade pattern for new MCP tools (custom exception class, success/error helpers, unit-test fixture). - [
test_case_format.md](docs/testcaseformat.md) — Python and YAML test-case formats; full field reference. - [
writing_effective_tests.md](docs/writingeffectivetests.md) — How to write tests that produce useful signal for evaluation and RL training. - [
test_cases_deep_dive.md](docs/testcasesdeep_dive.md) — Deeper examples and patterns for test-case authoring. - [
history_and_metadata.md](docs/historyandmetadata.md) — Multi-turn test cases with prior conversation history loaded from a companion.meta.yaml. - [
debugging_tests.md](docs/debugging_tests.md) — How to debug a failing test (VSCode launch configs and friends).
Fixtures and judges
- [
fixtures.md](docs/fixtures.md) — How fixtures are wired up (dependency injection viaconftest.yamland scenario overrides). - [
rubrics_judge.md](docs/rubrics_judge.md) — Rubric Judge: design, scoring, and how rewards are calculated. - [
generated_answer_evaluator.md](docs/generatedanswerevaluator.md) —GeneratedAnswerEvaluatorfixture for knowledge-QA / RAG test cases.
Configuration
- [
llm_endpoint_config.md](docs/llmendpointconfig.md) — Configuring LLM endpoints (Azure OpenAI, OpenAI-compatible, Anthropic). - [
session_proxy_config.md](docs/sessionproxyconfig.md) — Session Proxy configuration file (servers.yaml, auth, GC). - [
scenario_tools_config.md](docs/scenariotoolsconfig.md) — Tools list and per-tool overrides inside a scenario YAML. - [
prompts.md](docs/prompts.md) — How system, user-LLM, and judge prompts are constructed.
Operations
- [
tools_with_additional_setup.md](docs/toolswithadditional_setup.md) — Tools that require extra setup (Typesense, embeddings server, theTHINKINGBOX_DATAenv var).
Configuration and dataset
Configuration
The configuration file (--config, config_types.py:ConfigFile) contains:
- MCP session proxy address
- LLM service configurations
See examples i
…
Source & license
This open-source MCP server is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: microsoft
- Source: microsoft/thinkingbox
- License: MIT
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet, be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.