# Neuralforge

> Local AI workstation dashboard Neuralforge— manage LLMs, agents, RAG, Telegram AI bot with 14 meme personas & voice cloning, image/video/3D generation, and LoRA fine-tuning and SMM module from a single web UI. Runs entirely on your hardware.

- **Type:** MCP server
- **Install:** `agentstack add mcp-definitelyn0tme-neuralforge`
- **Verified:** Pending review
- **Seller:** [DefinitelyN0tMe](https://agentstack.voostack.com/s/definitelyn0tme)
- **Installs:** 0
- **Category:** [AI & ML](https://agentstack.voostack.com/c/ai-and-ml)
- **Latest version:** 0.1.0
- **License:** MIT
- **Upstream author:** [DefinitelyN0tMe](https://github.com/DefinitelyN0tMe)
- **Source:** https://github.com/DefinitelyN0tMe/neuralforge

## Install

```sh
agentstack add mcp-definitelyn0tme-neuralforge
```

Requires the [AgentStack CLI](https://agentstack.voostack.com/docs/cli). Works with Claude Code, Cursor, and any MCP-compatible agent.

## About

🧠 NeuralForge
  
    Self-hosted AI command center. 11 services. 69 APIs. Zero cloud.
  
  
    LLM agents, SMM autopilot for 7 platforms, image/video/3D/music generation,RAG, LoRA fine-tuning, voice cloning, Telegram bot with vision — all from localhost:9000
  
  
    
    
    
    
  
  
    
    
    
    
    
    
    
    
    
  

---

> **Why another AI dashboard?** Because no other open-source project gives you LLM orchestration, automated SMM for 7 platforms, image/video/3D generation, RAG, fine-tuning, Telegram bot with 14 personas, voice cloning, and MCP integration for Claude — all in a single self-hosted panel with zero cloud dependencies.

---

### Highlights

- **🤖 11 AI Services** managed from one UI — Ollama, ComfyUI, Whisper, Qdrant, SearXNG, and more
- **📱 SMM AI Department** — discover trends → generate posts → create images → auto-publish to Telegram, Twitter, Facebook, Instagram, Threads, LinkedIn, Discord simultaneously
- **🧠 Multi-Agent System** — 13 roles, 3 modes (Solo/Team/Orchestrator), 9 tools including web search, code execution, RAG
- **🎨 Full Generation Pipeline** — Image (FLUX) → Video (Wan2.2) → 3D (Hunyuan3D) with smart VRAM management
- **📊 69 API Endpoints** — everything is programmable, extensible, and automatable
- **🔒 100% Local** — your data never leaves your machine. No API keys required for core features

---

### Table of Contents

[What is this?](#what-is-this) · [AI Model Stack](#ai-model-stack) · [Features](#features-at-a-glance) · [Dashboard](#dashboard) · [Agents](#ai-agents) · [RAG](#rag-retrieval-augmented-generation) · [LoRA](#lora-fine-tuning) · [Pipeline](#generation-pipeline) · [Telegram Bot](#telegram-ai-bot) · [SMM](#smm-ai-department) · [MCP Server](#mcp-server--claude-code-integration) · [Quick Start](#quick-start) · [Requirements](#requirements) · [FAQ](#faq)

---

## What is this?

A self-hosted web panel (`localhost:9000`) that unifies your **entire local AI infrastructure** into one powerful dashboard. No subscriptions, no cloud APIs, no data leaving your machine.

```
┌──────────────────────────────── NeuralForge ──────────────────────────────────┐
│                                                                                │
│  Dashboard     Agents       RAG        Telegram     LoRA        SMM           │
│  ┌────────┐   ┌────────┐  ┌────────┐  ┌────────┐  ┌────────┐  ┌────────┐    │
│  │GPU/VRAM│   │13 Roles│  │Qdrant +│  │14 Meme │  │Unsloth │  │7 Socials│   │
│  │Services│   │Solo    │  │ONNX GPU│  │Personas│  │LoRA    │  │Trend AI │   │
│  │Metrics │   │Team    │  │1800/sec│  │Voice   │  │16 base │  │Post Gen │   │
│  │Alerts  │   │Orchestr│  │Multi-DB│  │Cloning │  │models  │  │Analytics│   │
│  └────────┘   └────────┘  └────────┘  └────────┘  └────────┘  └────────┘    │
│                                                                                │
│  Pipeline: Image ──→ Video ──→ 3D   │   MCP Server: 24 tools for Claude      │
│  (ComfyUI)  (Wan2GP)  (Hunyuan3D)   │   + Music, TTS, STT, Search...        │
└────────────────────────────────────────────────────────────────────────────────┘
```

## AI Model Stack

Every model runs **locally via Ollama** — no API keys, no cloud, no subscriptions.

### LLMs (Text Generation & Reasoning)

| Model | Size | VRAM | Used for |
|-------|------|------|----------|
| **Qwen 3.5** | 35B (A3B MoE) | ~20 GB | Primary workhorse — agents, SMM posts, trend analysis |
| **Nemotron 3 Nano** | 30B | ~18 GB | RAG answers, balanced quality/speed |
| **Mistral Small** | 24B | ~14 GB | Summarization, translation, email |
| **Qwen 3.5** | 9B | ~6 GB | Telegram bot — fast persona responses |
| **Gemma 3** | 27B | ~16 GB | Alternative general-purpose |
| **DeepSeek R1** | 14B | ~9 GB | Math, reasoning, code |
| + 9 more | 1B–35B | 1–20 GB | User-selectable per task |

### Vision (Image Understanding)

| Model | Size | VRAM | Used for |
|-------|------|------|----------|
| **MiniCPM-V** | 8B | ~5 GB | **Telegram bot photo analysis** — describes images, answers questions about photos sent to your account |
| **Qwen2.5-VL** | 27B | ~16 GB | Agent image analysis tool — detailed visual Q&A |

### Embeddings (RAG Search)

| Model | Size | Speed | Used for |
|-------|------|-------|----------|
| **bge-m3** (ONNX) | 560M | **1,800 docs/sec** | GPU-accelerated document indexing |
| **bge-m3** (Ollama) | 560M | 10 docs/sec | Fallback CPU embedding |

### Audio (Speech & Music)

| Model | Size | VRAM | Used for |
|-------|------|------|----------|
| **Whisper** (faster-whisper) | base/large | 2-10 GB | Speech-to-text, 99 languages, diarization |
| **Qwen3-TTS** | ~4 GB | ~4 GB | Text-to-speech, **3-second voice cloning** |
| **ACE-Step 1.5** | ~4 GB | 4-6 GB | AI music generation — lyrics + style → full song |

### Image / Video / 3D Generation

| Model | Engine | VRAM | Used for |
|-------|--------|------|----------|
| **FLUX Klein** | ComfyUI | 8-12 GB | Image generation (SMM posts, pipeline) |
| **Wan 2.2** | Wan2GP | 12-24 GB | Video generation from image + prompt |
| **Hunyuan3D v2** | Gradio | 13-20 GB | 3D model generation from image |

### LoRA Fine-Tuning (16 base models)

| Model | Size | Training time |
|-------|------|---------------|
| NVIDIA Nemotron 3 Nano | 4B | ~30 min |
| Llama 3.1 / 3.2 | 1B–8B | 30 min – 2h |
| Qwen 2.5 | 7B / 32B | 1–4h |
| Gemma 2 | 2B–27B | 30 min – 3h |
| Mistral v0.3 | 7B | ~1h |
| Phi 3.5 | 3.8B | ~40 min |

> **Total unique AI models available: 30+** — all running locally, swappable per task, with automatic VRAM management.

## Features at a Glance

| | Feature | Description |
|:---:|---|---|
| **GPU** | Smart VRAM Management | Exclusive groups auto-stop conflicting services. Never OOM again |
| **Dashboard** | Real-time Monitoring | GPU temp, VRAM, RAM, CPU, disk — live metrics with health alerts |
| **Agents** | Multi-Agent Orchestration | 13 roles, 3 modes (Solo/Team/Orchestrator), shared memory, RAG tools |
| **RAG** | Vector Search at GPU Speed | ONNX embeddings at 1,800 docs/sec, Qdrant DB, multi-collection search |
| **Bot** | 14 Telegram Personas | Each with unique personality — from Philosopher to Crypto Maniac |
| **Voice** | Real-time Voice Cloning | Send voice → get reply in *your own voice* with AI-generated text |
| **LoRA** | Fine-Tuning UI | 16 base models, dataset upload, live training output, adapter export |
| **Gen** | Image→Video→3D Pipeline | Automated chain with smart VRAM switching between steps |
| **MCP** | Claude Code Integration | 24 tools — let Claude manage your entire AI stack |
| **SMM** | 7-Platform Social Media | Trend Scout → AI Post Writer → Image Gen → Auto-Publish to all platforms |
| **Ext** | YAML Module System | Add any new service in 10 lines of YAML |

## Dashboard

The main hub. Everything starts here.

  
  

**Live Metrics:**
- GPU VRAM usage with free memory indicator
- GPU temperature and power draw
- RAM usage with available memory
- CPU load across all threads
- Disk usage with free space alerts

**Service Management:**
- Start/stop any service with one click
- **Exclusive GPU groups** — when you start ComfyUI, Wan2GP auto-stops (and vice versa). No more VRAM crashes
- Service health indicators (running/stopped/starting)
- Quick actions: "Start basics", "Stop heavy", "Free VRAM"

**Monitoring:**
- Active Ollama models with per-model VRAM usage
- GPU process list (what's eating your VRAM right now)
- Qdrant RAG collections with vector counts
- Storage breakdown by service (ComfyUI outputs, Wan2GP videos, etc.)
- Health alerts: GPU overheating, low disk, service down — all visible at a glance

**YAML Module System — add any service:**
```yaml
name: My New Service
category: generation
start_cmd: "python3 app.py --port 7777"
port: 7777
vram_estimate: "4-8 GB"
exclusive_group: heavy_gpu    # auto-stops conflicting services
```
Drop it in `modules/` → restart panel → it appears. That's it.

## AI Agents

A full multi-agent framework built into the panel.

  

**13 Role Presets:**

| Role | What it does | Default model |
|------|-------------|---------------|
| Researcher | Web search, source analysis, fact compilation | Qwen 3.5 35B |
| Analyst | Data analysis, pattern recognition, insights | Qwen 3.5 35B |
| Coder | Write, debug, refactor code in any language | Qwen 3.5 35B |
| Writer | Articles, reports, creative writing | Qwen 3.5 35B |
| Critic | Quality review, scoring, improvement suggestions | Qwen 3.5 35B |
| Summarizer | Condense long texts into key points | Mistral Small 24B |
| Translator | Multi-language translation with context | Mistral Small 24B |
| Email Writer | Professional emails from brief instructions | Mistral Small 24B |
| Tester | Generate test cases, find edge cases | Qwen 3.5 35B |
| Trade Analyst | Market analysis, trend identification | Qwen 3.5 35B |
| Tutor | Explain concepts at adjustable complexity | Qwen 3.5 35B |
| Security Auditor | Code/config security review, vulnerability scan | Qwen 3.5 35B |
| Image Analyst | Describe and analyze images | Qwen Vision 27B |

**3 Execution Modes:**

| Mode | How it works | Best for |
|------|-------------|----------|
| **Solo** | Single agent with tools | Quick tasks, Q&A |
| **Team** | Agent chain — each passes context to next | Complex multi-step tasks |
| **Orchestrator** | AI creates plan → delegates to agents → reviews result (retries if score 

## RAG (Retrieval-Augmented Generation)

Ask questions about your documents. The AI retrieves relevant passages and answers with citations.

  

**Performance:**
| Method | Speed | GPU VRAM |
|--------|-------|----------|
| ONNX GPU (bge-m3) | **1,800 texts/sec** | ~2 GB |
| Ollama embeddings | 10 texts/sec | ~4 GB |

That's **180x faster** indexing with ONNX.

**Capabilities:**
- **Multi-format indexing** — PDF, TXT, MD, DOCX, CSV, HTML
- **Batch processing** — index entire directories recursively
- **Multi-collection** — separate databases for different topics (e.g., "laws", "docs", "codebase")
- **Smart search** — auto-detects which collection to search based on query
- **Context memory** — remembers previous Q&A in the same chat session
- **Embedding cache** — repeat queries are instant

**Built-in chat interface:**
- Markdown rendering with syntax highlighting
- Copy button on every response
- Export conversation to Markdown file
- Collection selector and search settings
- localStorage persistence — your chat survives page reload

**Example use case:**
> Indexed all 390 Estonian laws (52,314 vectors) — now ask legal questions in any language and get answers with article references.

## LoRA Fine-Tuning

Train custom model adapters directly from the panel UI.

  

**16 Base Models Ready to Fine-Tune:**

| Model | Size | Notes |
|-------|------|-------|
| Llama 3.1 | 8B | Great all-rounder |
| Llama 3.2 | 1B / 3B | Lightweight, fast |
| Mistral v0.3 | 7B | Strong reasoning |
| Qwen 2.5 | 7B / 32B | Multilingual |
| Gemma 2 | 2B / 9B / 27B | Google's latest |
| Phi 3.5 | 3.8B | Microsoft, compact |
| + custom | any | Enter any Unsloth-compatible model ID |

**Training UI Features:**
- Dataset upload (JSON, JSONL, CSV) or HuggingFace dataset ID
- Auto-format detection (instruction/output, messages, or raw text)
- Configurable: LoRA rank, alpha, epochs, batch size, learning rate, max sequence length
- **Live training output** — see loss, progress, ETA in real-time
- Timer showing elapsed training time
- Stop button to cancel mid-training
- Trained adapters listed with size and date

**Powered by [Unsloth](https://github.com/unslothai/unsloth)** — 2x faster training, 60% less memory than standard LoRA.

## Generation Pipeline

Automated **Image → Video → 3D** chain with smart VRAM management between steps.

| Step | Engine | VRAM | Automation |
|------|--------|------|------------|
| Image | ComfyUI (FLUX Klein 4B) | 8-12 GB | Fully automated API |
| Video | Wan2GP (Wan 2.2 / LTX) | 12-24 GB | Gradio API + manual fallback |
| 3D | Hunyuan3D | 13-20 GB | Gradio API + manual fallback |

VRAM is automatically freed between steps — only one heavy service runs at a time.

```bash
# 5 built-in examples
python3 pipeline.py --example robot     # chibi robot → animate → 3D model
python3 pipeline.py --example dragon    # crystal dragon → animate → 3D
python3 pipeline.py --example car       # cyberpunk car → animate → 3D
python3 pipeline.py --example cat       # cat astronaut → animate → 3D
python3 pipeline.py --example sword     # magic sword → 3D (skip video)

# Custom prompt
python3 pipeline.py "a golden crown with gems" --steps image,3d
python3 pipeline.py "a phoenix" --video-prompt "spreads wings and flies"
```

## Telegram AI Bot

Not a Telegram bot — responds **from your own account** via Telethon User API.

  

**14 Unique Personas:**

| | Persona | Style |
|:---:|---|---|
| 🧘 | **Philosopher** | *"You wrote 'hi', but what is a greeting if not a scream of loneliness into the void?"* |
| 🧢 | **Street Philosopher** | *"bro, your argument is logically inconsistent, purely by Kant"* |
| 👾 | **IT Demon** | *"segfault in your logic, recompile that thought"* |
| 👵 | **Granny from 2077** | *"sweetie, browsing without a firewall again? you'll catch a virus!"* |
| 🕵️ | **Noir Detective** | *"The message came at 3am. Like all bad news in this city"* |
| 🏴‍☠️ | **Nerd Pirate** | *"arrr, your meme is a true treasure!"* |
| 🐱 | **Cat Overlord** | *"I'd help, but I need to lie down for 14 more hours"* |
| 🔺 | **Conspiracy Nut** | *"Telegram was created by Masons to track memes"* |
| 🎭 | **Budget Shakespeare** | *"To be online or not to be — that is the question!"* |
| 🧟 | **Polite Zombie** | *"good evening, could you... share some brains?"* |
| 📋 | **Corporate Robot** | *"let's sync on this in the next sprint"* |
| 🫎 | **Capybara** | *"why stress when you can just... not"* + random capybara photo |
| 🚀 | **Crypto Maniac** | *"RED CANDLE, I'M BANKRUPT, wait... GREEN! I'M RICH!"* |
| 🛠️ | **Custom** | Write your own character |

**Voice Clone Pipeline:**
```
🎤 Voice in → ffmpeg (OGG→WAV) → Whisper STT → LLM response
  → unload LLM → Qwen3-TTS (clone voice) → ffmpeg (WAV→OGG) → 🔊 Voice out
```

**Vision — Photo Analysis (MiniCPM-V 8B):**
```
📸 Photo in → MiniCPM-V (image description) → LLM (persona-styled response) → 💬 Reply
```
Send a photo to your account → the bot describes it through the vision model → responds in character. Toggle on/off from panel UI.

**Features:**
- **Vision mode** — understands photos via MiniCPM-V 8B (auto VRAM swap: unload LLM → load vision → analyze → unload → reload LLM)
- Auto-detects language → responds in same language
- Conversation memory (5 exchanges per user)
- Session-based logs grouped by contact
- Voice clone toggle from panel UI
- Capybara persona sends random capybara photos via [capy.lol](https://capy.lol) API

## SMM AI Department

Fully automated social media management system — from trend discovery to publishing across 7 platforms.

  

**Complete Workflow:**
```
Trend Scout → Post Writer → Image Gen → Content Queue → Auto-Publish
    │              │             │              │              │
    ▼              ▼             ▼              ▼              ▼
 6 Sources     2-Pass LLM    ComfyUI FLUX   Schedule +     7 Platforms
 (Reddit,HN,  (Scrape→       + ffmpeg      Calendar       simultaneously
  GitHub,RSS,  Summary→       resize        view           with retry
  SearXNG,     Platform
  GoogTrends)  posts)
```

**7 Connected Platforms:**
| Platform | Auth Method | Features |
|---|---|---|
| Telegram | Bot API | Text + Photo, channel posting |
| Discord | Webhook | Text + File upload |
| Twitter/X | OAuth 1.0a | Text + Media upload (Pay-Per-Use) |
| Facebook | Page Token (permanent) | Text + Photo, Page posting |
| Instagram | Graph API via FB | Photo + Caption (via imgur) |
| Threads | Threads API | Text + Image |
| LinkedIn | OAuth 2.0 | Text + Image (3-step upload) |

**Key Features:**
- **Trend Scout v2** — Multi-source intelligence with nich

…

## Source & license

This open-source MCP server is cataloged on AgentStack and links to its original source — we do not rehost the code.

- **Author:** [DefinitelyN0tMe](https://github.com/DefinitelyN0tMe)
- **Source:** [DefinitelyN0tMe/neuralforge](https://github.com/DefinitelyN0tMe/neuralforge)
- **License:** MIT

Install and usage instructions live in the source repository linked above.

## Pricing

- **Free** — Free

## Security capabilities

Automated source analysis of v0.1.0 — what this tool can access:

- **Network access:** yes
- **Filesystem access:** no
- **Shell / process execution:** no
- **Environment & secrets:** no
- **Dynamic code execution:** no

*"Yes" means the capability is present in the source — more access means more to trust, not that it is unsafe.*


## Versions

- **0.1.0** — security scan: flagged — Imported from the upstream source.

## Links

- Listing page: https://agentstack.voostack.com/l/mcp-definitelyn0tme-neuralforge
- Seller: https://agentstack.voostack.com/s/definitelyn0tme
- Browse the marketplace: https://agentstack.voostack.com/browse

---
Listed on AgentStack — the marketplace for AI agent skills and MCP servers. Every listing is security-reviewed. Creators keep 70%.
