# Ffvoice Engine

> Offline speech-to-text & speaker diarization for AI agents — Whisper ASR + MCP server + CLI + Python bindings, fully on-device, no cloud | 离线语音识别与说话人分离,Agent 开箱即用

- **Type:** MCP server
- **Install:** `agentstack add mcp-chicogong-ffvoice-engine`
- **Verified:** Yes — security-reviewed for prompt injection and unsafe behavior
- **Seller:** [chicogong](https://agentstack.voostack.com/s/chicogong)
- **Installs:** 0
- **Category:** [Content & Media](https://agentstack.voostack.com/c/content-and-media)
- **Latest version:** 0.1.0
- **License:** MIT
- **Upstream author:** [chicogong](https://github.com/chicogong)
- **Source:** https://github.com/chicogong/ffvoice-engine

## Install

```sh
agentstack add mcp-chicogong-ffvoice-engine
```

Requires the [AgentStack CLI](https://agentstack.voostack.com/docs/cli). Works with Claude Code, Cursor, and any MCP-compatible agent.

## About

# ffvoice-engine

[](https://github.com/chicogong/ffvoice-engine/actions/workflows/ci.yml)
[](https://github.com/chicogong/ffvoice-engine/actions/workflows/release.yml)

[](https://opensource.org/licenses/MIT)
[](https://en.cppreference.com/w/cpp/20)
[](https://cmake.org/)

[]()
[](https://github.com/chicogong/ffvoice-engine/actions)
[](https://github.com/chicogong/ffvoice-engine/actions)
[](https://github.com/chicogong/ffvoice-engine/actions)

[](https://pypi.org/project/ffvoice/)
[](https://pypi.org/project/ffvoice/)
[](https://github.com/chicogong/ffvoice-engine/releases)
[](https://github.com/chicogong/ffvoice-engine/stargazers)
[](https://github.com/chicogong/ffvoice-engine/network/members)

[](https://ffmpeg.org/)
[](http://www.portaudio.com/)
[](https://xiph.org/flac/)
[](https://github.com/ggerganov/whisper.cpp)

[](https://google.github.io/styleguide/cppguide.html)
[](CONTRIBUTING.md)

> 🎙️ **Offline speech-to-text & speaker diarization for AI agents** — Whisper ASR, live captioning, an MCP server, a CLI and Python bindings. Fully on-device, no cloud API.
>
> 🎙️ **离线语音识别 + 说话人分离,AI Agent 开箱即用** —— Whisper 实时转写 · 实时字幕 · MCP server · CLI · Python 绑定 · 100% 本地运行,音频不上云。

---

## Why ffvoice? / 为什么用 ffvoice？

**The honest pitch**: ffvoice is an **integration layer**, not a new ASR engine. It embeds [whisper.cpp](https://github.com/ggerganov/whisper.cpp) as-is and makes no changes to its accuracy or inference speed. What ffvoice adds is a **batteries-included, pre-wired pipeline** — microphone capture → RNNoise denoising → VAD segmentation → Whisper ASR → speaker diarization → live captions / WAV / FLAC / subtitles — delivered as a single C++ SDK with Python bindings, a CLI, and an **MCP server** that lets AI agents (Claude and others) transcribe audio out of the box — all in one `pip install` or `cmake` build.

**诚实定位**: ffvoice 是一个**集成层**，而非新的 ASR 引擎。它内嵌 whisper.cpp，不修改其识别精度或推理速度。ffvoice 带来的是一条**开箱即用、预连接的完整管道**——麦克风采集 → RNNoise 降噪 → VAD 分段 → Whisper ASR → 说话人分离 → 实时字幕 / WAV / FLAC / 字幕输出——打包成 C++ SDK + Python 绑定 + CLI，外加一个 **MCP server**,让 AI agent(Claude 等)开箱即用地转写音频——一条 `pip install` 或 `cmake` 即可完成。

### Pain points it addresses / 解决的痛点

| Pain point | ffvoice approach |
|------------|-----------------|
| **Privacy / 隐私合规** — audio must not leave the device (GDPR, HIPAA, enterprise policy) | 100% offline; audio never transmitted |
| **Cloud cost / 云端费用** — commercial APIs charge per minute ($0.01–0.024/min at scale) | Zero per-minute cost; runs on your own hardware |
| **Glue code / 胶水代码** — wiring PortAudio + RNNoise + VAD + whisper.cpp + FLAC yourself takes days | All wired together and tested; one SDK |
| **Offline / 断网场景** — embedded systems, air-gapped environments, poor connectivity | Fully offline; no network dependency |
| **Low latency / 低延迟** — cloud round-trips add 200–800ms per request | Local inference;  **注意**:
> - **Linux/macOS**: 使用 `./build/ffvoice`
> - **Windows**: 使用 `.\build\Release\ffvoice.exe`

```bash
# 查看帮助
./build/ffvoice --help

# 生成测试 WAV 文件（440Hz A4 音符，3秒）
./build/ffvoice --test-wav test.wav

# 列出可用音频设备
./build/ffvoice --list-devices

# 录制 10 秒 WAV 音频（默认格式）
./build/ffvoice --record -o recording.wav -t 10

# 录制 30 秒 FLAC 音频（无损压缩）
./build/ffvoice --record -o recording.flac -t 30

# 使用最大压缩级别录制 FLAC
./build/ffvoice --record -o recording.flac --compression 8 -t 60

# 选择特定设备录制立体声
./build/ffvoice --record -d 1 -o stereo.wav --channels 2 -t 20

# 启用音频处理（音量归一化 + 高通滤波）
./build/ffvoice --record -o clean.wav --enable-processing -t 10

# 仅启用音量归一化
./build/ffvoice --record -o normalized.wav --normalize -t 10

# 自定义高通滤波频率（去除 100Hz 以下噪声）
./build/ffvoice --record -o filtered.flac --highpass 100 -t 20

# 组合：FLAC + 音频处理
./build/ffvoice --record -o studio.flac --normalize --highpass 80 -t 30

# RNNoise 深度学习降噪（推荐用于语音录制，仅 Linux/macOS）
./build/ffvoice --record -o clean.wav --rnnoise -t 10

# 完整处理链（高通 + RNNoise + 归一化，仅 Linux/macOS）
./build/ffvoice --record -o studio.flac --highpass 80 --rnnoise --normalize -t 30

# RNNoise + VAD (实验性，仅 Linux/macOS)
./build/ffvoice --record -o vad.wav --rnnoise-vad -t 20

# 播放录音
afplay recording.wav   # 或 recording.flac

# ==================== 语音识别（需启用 ENABLE_WHISPER） ====================

# 转写音频文件为纯文本
./build/ffvoice --transcribe recording.wav -o transcript.txt

# 生成 SRT 字幕文件
./build/ffvoice --transcribe recording.wav --format srt -o subtitles.srt

# 生成 VTT 字幕文件
./build/ffvoice --transcribe recording.wav --format vtt -o subtitles.vtt

# 生成 JSON 转写文件（含分段级与词级时间戳）
./build/ffvoice --transcribe recording.wav --format json -o transcript.json

# 指定语言（中文）
./build/ffvoice --transcribe recording.wav --language zh -o transcript_zh.txt

# 转写 FLAC 文件
./build/ffvoice --transcribe recording.flac --format srt -o subtitles.srt

# 完整工作流：录制 + 音频处理 + 转写
./build/ffvoice --record -o speech.flac --highpass 80 --rnnoise --normalize -t 30
./build/ffvoice --transcribe speech.flac --format srt -o speech.srt

# ==================== 实时语音识别（需启用 ENABLE_RNNOISE 和 ENABLE_WHISPER） ====================

# 边录边转写（实时模式）
./build/ffvoice --record -o speech.wav --rnnoise-vad --transcribe-live -t 60

# 实时转写 + 音频处理
./build/ffvoice --record -o speech.flac --rnnoise-vad --transcribe-live --highpass 80 --normalize -t 120

# ==================== 实时字幕流（需启用 ENABLE_WHISPER；LiveCaptioner） ====================

# 边说边出字幕：partial 字幕实时刷新，final 字幕断句落定
./build/ffvoice --record -o talk.wav --live-captions -t 60

# 调整 partial 字幕刷新间隔（毫秒）
./build/ffvoice --record -o talk.wav --live-captions --partial-interval 300 -t 60

# ==================== 说话人分离（需 -DENABLE_DIARIZATION=ON 构建） ====================

# 转写并标注说话人：每段带 speaker_id（JSON 输出可见）
./build/ffvoice --transcribe meeting.wav --diarize --format json -o meeting.json

# 指定说话人数量（默认自动判定）
./build/ffvoice --transcribe meeting.wav --diarize --num-speakers 2 --format srt -o meeting.srt
```

> ⚠️ **说话人分离适用边界**:diarization 的 embedding 模型默认**英文调优**。中文或混合语言音频请改用内置的中英双语模型 —— 在 `transcribe_file_with_diarization` MCP 工具(或 `make_diarizer`)传 `embedding="multilingual"`,首次使用自动下载。已知说话人数时请传 `--num-speakers N`,自动判定数量在多说话人场景不够稳。

## 🐍 Python Bindings

ffvoice 提供高性能的 Python 绑定，让您在 Python 中轻松使用所有功能。

### 安装

**从 PyPI 安装** (推荐):
```bash
pip install ffvoice                    # 核心:转写 / VAD / 采集 / 混音
pip install 'ffvoice[mcp]'             # + MCP server(供 AI agent 调用)
pip install 'ffvoice[diarization]'     # + 说话人分离(零编译、全平台)
```

> 💡 **说话人分离免编译**:`[diarization]` extra 经 `sherpa-onnx` 提供 diarization,无需从源码构建 ONNX Runtime;Whisper 与 diarization 模型在首次使用时自动下载到 `~/.cache/ffvoice/`,无需手动配置。

**从源码安装**:
```bash
git clone https://github.com/chicogong/ffvoice-engine.git
cd ffvoice-engine
pip install .
```

### 平台兼容性

| 平台 | PyPI Wheel | 安装方式 | 状态 |
|------|-----------|---------|------|
| **🍎 Apple Silicon (M1/M2/M3)** | ✅ ARM64 | `pip install ffvoice` | ✅ 原生支持 |
| **🍎 Intel Mac** | ❌ 不兼容 | 从源码编译 | ⚠️ 需手动构建 |
| **🐧 Linux x86_64** | ✅ x86_64 | `pip install ffvoice` | ✅ 原生支持 |
| **🪟 Windows x86_64** | ✅ x86_64 | `pip install ffvoice` | ✅ 原生支持 |

**重要说明**:
- **Apple Silicon 用户**: 直接使用 `pip install ffvoice` 即可，性能最佳
- **Windows 用户**: 现已支持 Windows x86_64 预编译 wheels，直接使用 `pip install ffvoice` 即可
  - 支持 Python 3.10-3.14
  - 自动包含所有必需的依赖（无需手动安装 FFmpeg 等）
  - **注意**: Windows 版本禁用了 RNNoise 降噪（MSVC 不支持 VLA），其他功能完全可用
- **Intel Mac 用户**: PyPI wheel 不兼容，需要从源码编译:
  ```bash
  # 确保已安装依赖
  brew install cmake ffmpeg portaudio flac

  # 从源码安装
  git clone https://github.com/chicogong/ffvoice-engine.git
  cd ffvoice-engine
  pip install .
  ```
- **Rosetta 2 用户**: ARM64 wheel 在 Rosetta 环境下不工作，请使用 ARM64 原生 Python:
  ```bash
  # 检查 Python 架构
  python -c "import platform; print(platform.machine())"
  # 应该输出 'arm64'，如果是 'x86_64' 则需要重新安装 ARM64 Python

  # 强制使用 ARM64 Python
  arch -arm64 python3 -m pip install ffvoice
  ```

### 快速示例

```python
import ffvoice
import numpy as np

# 1. 语音识别
config = ffvoice.WhisperConfig()
config.model_type = ffvoice.WhisperModelType.TINY
asr = ffvoice.WhisperASR(config)
asr.initialize()

# 从文件转写
segments = asr.transcribe_file("audio.wav")
for seg in segments:
    print(f"[{seg.start_ms}ms - {seg.end_ms}ms] {seg.text}")

# 从 NumPy 数组转写
audio = np.zeros(48000, dtype=np.int16)  # 1秒音频
segments = asr.transcribe_buffer(audio)

# 2. 噪声抑制
rnnoise = ffvoice.RNNoise(ffvoice.RNNoiseConfig())
rnnoise.initialize(sample_rate=48000, channels=1)

audio = np.random.randint(-1000, 1000, 256, dtype=np.int16)
rnnoise.process(audio)  # 原地处理
vad_prob = rnnoise.get_vad_probability()

# 3. 实时音频采集
def audio_callback(audio_array):
    print(f"收到 {len(audio_array)} 个采样")

ffvoice.AudioCapture.initialize()
capture = ffvoice.AudioCapture()
capture.open(sample_rate=48000, channels=1, frames_per_buffer=256)
capture.start(audio_callback)
# ... 录制中 ...
capture.stop()
capture.close()
ffvoice.AudioCapture.terminate()

# 4. 多音轨混音
mixer = ffvoice.AudioMixer()
mixer.initialize(sample_rate=48000, channels=2)
track = mixer.add_track(gain=1.0, pan=0.0)
mixed = mixer.mix_block({track: np.zeros(480, dtype=np.int16)})

# 5. 无锁环形缓冲区（实时音频路径的线程间交接）
ring = ffvoice.RingBuffer(capacity=4096)
ring.push_bulk(np.zeros(1024, dtype=np.int16))
chunk = ring.pop_bulk(512)

# 6. 词级时间戳
config.word_timestamps = True  # 转写结果的每个分段附带 words 数组
for seg in asr.transcribe_file("audio.wav"):
    for word in seg.words:
        print(f"  [{word.start_ms}-{word.end_ms}ms] {word.text}")
```

### 完整文档

详细文档和示例请查看 [`python/README.md`](python/README.md):
- 📖 完整 API 参考（含 `AudioMixer` 多音轨混音、`RingBuffer` 无锁环形缓冲区、词级时间戳 `Word` / `TranscriptionSegment.words`）
- 🎯 16+ 代码示例
- 🚀 Quick Start 指南
- 📓 Jupyter Notebook 教程

**性能优势**:
- ⚡ **3-10x 更快** - C++ 核心 vs 纯 Python 实现
- 💾 **零拷贝** - NumPy 数组直接传递
- 🔒 **100% 离线** - 无需网络，隐私安全
- 🎙️ **完整工作流** - 采集 → 降噪 → VAD → 识别

## 🤖 AI Agent 集成 — MCP Server + Agent Skill

ffvoice 内置了一个 [MCP (Model Context Protocol)](https://modelcontextprotocol.io/) 服务器,让 AI agent(如 Claude Desktop)能够直接调用本地离线语音识别能力,**全程无需联网,音频数据绝不离开本机**。

### 安装

```bash
pip install 'ffvoice[mcp]'                # MCP server
pip install 'ffvoice[mcp,diarization]'    # + 说话人分离工具
```

### 提供的工具

| 工具 | 说明 |
|------|------|
| `transcribe_file` | 转写本地音频文件（WAV/FLAC 等），支持语言选择、模型大小、词级时间戳 |
| `transcribe_file_with_diarization` | 转写并标注说话人 —— 每段带 `speaker_id`，回答"谁在何时说什么"（需 `pip install 'ffvoice[diarization]'`,免编译） |
| `capture_and_transcribe` | 录制指定时长的麦克风音频并实时转写（内置 VAD 分段 + 可选 RNNoise 降噪） |
| `capture_and_caption` | 录制麦克风音频并产出实时字幕流（LiveCaptioner，partial/final 事件） |
| `list_audio_devices` | 列出所有可用的音频输入/输出设备及默认设备 ID |

### 接入 Claude Desktop

将以下配置粘贴到 Claude Desktop 的 `claude_desktop_config.json`：

```json
{"mcpServers": {"ffvoice": {"command": "ffvoice-mcp", "args": []}}}
```

重启 Claude Desktop 后，即可在对话中直接请求转写本地音频或录制语音。

### Claude Agent Skill

仓库内置一个 **Claude Agent Skill**([`.claude/skills/ffvoice-transcription/`](.claude/skills/ffvoice-transcription/))—— 用 Claude Code 打开本仓库即被**自动发现**,无需任何配置。它教 agent 何时、如何用 ffvoice 的 MCP 工具与 CLI 完成转写、说话人分离、实时字幕。配合 MCP server,ffvoice 对 AI agent 真正做到开箱即用。

## 📁 项目结构

```
ffvoice-engine/
├── CMakeLists.txt          # 主构建文件
├── include/ffvoice/        # 公共头文件
│   └── types.h             # 核心类型定义
├── src/                    # 源代码
│   ├── audio/              # 音频采集与处理模块
│   │   ├── audio_capture_device.* # ✅ PortAudio 采集器
│   │   ├── audio_mixer.*          # ✅ 多音轨混音器
│   │   ├── audio_processor.*      # ✅ 音频处理框架
│   │   ├── rnnoise_processor.*    # ✅ RNNoise 深度学习降噪 (可选)
│   │   ├── vad_segmenter.*        # ✅ VAD 音频分段器
│   │   ├── whisper_processor.*    # ✅ Whisper ASR 语音识别 (可选)
│   │   ├── live_captioner.*       # ✅ 实时字幕流 LiveCaptioner (可选)
│   │   └── diarizer.*             # ✅ 说话人分离 Diarizer (可选)
│   ├── media/              # 媒体编码/封装
│   │   ├── wav_writer.*    # ✅ WAV 文件写入器
│   │   └── flac_writer.*   # ✅ FLAC 无损压缩
│   └── utils/              # 工具类
│       ├── signal_generator.* # ✅ 音频信号生成
│       ├── ring_buffer.*   # ✅ 环形缓冲区
│       ├── audio_converter.*  # ✅ 音频格式转换
│       ├── subtitle_generator.* # ✅ 字幕生成（SRT/VTT）
│       └── logger.*        # ✅ 日志工具
├── apps/cli/               # CLI 应用
│   └── main.cpp            # ✅ 完整录音功能
├── tests/                  # 单元测试
│   ├── unit/               # ✅ 311 个测试用例（全部通过）
│   ├── mocks/              # Mock 对象
│   └── fixtures/           # 测试夹具
├── models/                 # AI 模型文件
└── scripts/                # 辅助脚本
```

## 🛣️ 路线图

**Milestone 1–6** —— 基础录制 → 音频增强(RNNoise) → 离线 ASR(Whisper) → 实时 ASR → 性能优化 → AudioMixer / RingBuffer / 词级时间戳 —— ✅ 全部完成。

**集成层路线图（v0.8.3）—— 四个阶段全部交付：**

- ✅ **Phase 1 Agent 集成** — CLI 硬化 + MCP server，让 AI agent 把 ffvoice 当本地离线语音工具调用；v0.7.0 已发布到 PyPI
- ✅ **Phase 2 实时字幕流** — LiveCaptioner，partial/final 字幕事件，边说边出字
- ✅ **Phase 3 说话人分离** — Diarizer（sherpa-onnx 离线 diarization），CLI `--diarize` + MCP `transcribe_file_with_diarization`

完整历史见 [CHANGELOG.md](CHANGELOG.md)。

## 📝 开发说明

主分支：`master`

### 代码规范

- C++20 标准
- Google C++ Style Guide（部分）
- 使用 clang-format 格式化

### 测试

```bash
# 配置并编译测试
cmake .. -DBUILD_TESTS=ON -DCMAKE_BUILD_TYPE=Debug
make -j4

# 运行所有测试
make test

# 运行单个测试（详细输出）
./build/tests/ffvoice_tests --gtest_filter=WavWriter*
```

### 已实现功能

#### AudioCaptureDevice - 音频采集器
- 基于 PortAudio 的跨平台音频捕获
- 实时流式采集（回调模式）
- 设备枚举和自动选择
- 低延迟配置（256 帧缓冲）
- 支持 mono/stereo
- 可配置采样率（默认 48kHz）

#### WavWriter - WAV 文件写入器
- 手写 RIFF/WAV 格式实现
- 支持 PCM 16-bit 音频
- 支持 mono/stereo
- 可调采样率
- 实时写入支持

#### FlacWriter - FLAC 无损压缩
- 基于 libFLAC 1.5.0
- 实时流式编码
- 可配置压缩级别（0-8，默认 5）
- 压缩比 1.5-3x（取决于音频内容）
- 支持 16/24-bit PCM
- 自动压缩比统计

#### SignalGenerator - 音频信号生成器
- 正弦波生成（可调频率、时长、振幅）
- 静音生成
- 白噪声生成
- 用于测试和调试

#### AudioProcessor - 音频处理框架
**架构设计**：
- 抽象接口 `AudioProcessor` 支持模块化扩展
- `AudioProcessorChain` 处理器链（串联多个处理器）
- 实时处理（在采集回调中）
- 就地处理（in-place）提高效率

**VolumeNormalizer - 音量归一化**：
- 基于 RMS 的自动增益控制
- 平滑增益调整（exponential moving average）
  - Attack time: 0.1s（增益提升速度）
  - Release time: 0.3s（增益下降速度）
- 目标电平：0.3（可配置 0.0-1.0）
- 增益范围：0.1x - 10.0x
- 防止削波和保持一致响度

**HighPassFilter - 高通滤波器**：
- 一阶 IIR 滤波器实现
- 去除低频噪声（呼吸声、麦克风碰撞、环境噪音）
- 默认截止频率：80Hz（可配置）
- 每通道独立状态（支持立体声）
- 滤波器公式：`y[n] = α(y[n-1] + x[n] - x[n-1])`

**RNNoiseProcessor - RNNoise 深度学习降噪** (可选)：
- 基于 Xiph RNNoise 的 RNN 深度学习模型
- 专为语音优化的降噪算法
- 帧大小：480 samples (10ms @48kHz)
- 支持采样率：48kHz, 44.1kHz, 24kHz
- 多声道支持：每通道独立 DenoiseState
- 格式转换：自动处理 int16 ↔ float
- 帧缓冲管理：256 samples → 480 samples
- VAD 选项：可选语音活动检测（实验性）
- CPU 开销：~5-10%（显著低于 WebRTC APM）
- 降噪效果：~20dB（语音场景）

**性能**：
- 实时处理（
  Made with ❤️ by the ffvoice-engine team

  ⬆️ Back to Top

## Source & license

This open-source MCP server is cataloged on AgentStack and links to its original source — we do not rehost the code.

- **Author:** [chicogong](https://github.com/chicogong)
- **Source:** [chicogong/ffvoice-engine](https://github.com/chicogong/ffvoice-engine)
- **License:** MIT

Install and usage instructions live in the source repository linked above.

## Pricing

- **Free** — Free

## Security capabilities

Automated source analysis of v0.1.0 — what this tool can access:

- **Network access:** no
- **Filesystem access:** yes
- **Shell / process execution:** no
- **Environment & secrets:** no
- **Dynamic code execution:** no

*"Yes" means the capability is present in the source — more access means more to trust, not that it is unsafe.*


## Versions

- **0.1.0** — security scan: passed — Imported from the upstream source.

## Links

- Listing page: https://agentstack.voostack.com/l/mcp-chicogong-ffvoice-engine
- Seller: https://agentstack.voostack.com/s/chicogong
- Browse the marketplace: https://agentstack.voostack.com/browse

---
Listed on AgentStack — the marketplace for AI agent skills and MCP servers. Every listing is security-reviewed. Creators keep 70%.
