Install
$ agentstack add skill-zedarvates-botte-secrete-llm-backends ✓ scanned · ✓ verified — works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ✓ Network access No
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
About
llm_backends — Local LLM discovery, audit & routing
Turn idle local hardware into a token-saving tier. Every task served by a local model is a task not billed to the cloud.
When to use
- The user mentions LM Studio, Ollama, LocalAI, vLLM, llama.cpp, "local model",
"run it locally", or "use my GPU".
- You need to know what models are reachable (this machine or the network).
- A task is cheap/local-suitable (classification, extraction, short summary,
routing, spell-check, simple Q&A) and could skip the cloud entirely.
- The user has no local model yet and wants step-by-step, hardware-aware setup.
Quick commands
# Discover + register backends (writes configs/llm-endpoints.json)
python -m skills.llm_backends.cli scan # localhost only
python -m skills.llm_backends.cli scan --subnet # sweep local /24
python -m skills.llm_backends.cli scan 192.168.1.47 # specific host(s)
# What's registered?
python -m skills.llm_backends.cli list
# Audit: are local models used? what can this machine run? next steps?
python -m skills.llm_backends.cli audit --fresh
# Run a prompt locally (0 cloud tokens)
python -m skills.llm_backends.cli chat "classify: bug or feature?" --max-tokens 128
# Suggest a local model for a project (adaptive per project type)
python -m skills.llm_backends.cli profile ~/my-project
Programmatic use
from skills.llm_backends import registry, quick_chat, audit
registry.refresh() # discover + persist
best = registry.best_chat_backend() # lowest-latency chat backend
model = registry.preferred_model(best) # coder/instruct over voice/reasoning
res = quick_chat("summarize in 1 line: ...", max_tokens=200)
print(res.text, res.total_tokens) # all local — no cloud cost
Supported backends
| Backend | Default port | API | |---------|-------------|-----| | LM Studio | 1234 | OpenAI /v1 | | Ollama | 11434 | native /api/tags + OpenAI /v1 | | LocalAI | 8080 | OpenAI /v1 | | vLLM | 8000 | OpenAI /v1 | | Jan / KoboldCpp / text-gen-webui | 1337 / 5001 / 5000 | OpenAI /v1 | | ComfyUI | 8188 | image gen | | Qdrant | 6333 | vector search |
How it saves tokens
- Audit finds reachable local backends and profiles RAM/VRAM/GPU.
- Route sends local-suitable tasks (see
tiered_routerL0/L1) to a local model. - Call runs them via the OpenAI-compatible client — zero cloud tokens.
- Onboard users with no local model, recommending the largest model their
hardware can run and the right server to install.
Related: [[llm_mcp]] (MCP tools for agents), tiered_router (cost tiers), local_router (task→backend mapping), response_cache (skip repeated calls).
## Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- **Author:** [zedarvates](https://github.com/zedarvates)
- **Source:** [zedarvates/botte-secrete](https://github.com/zedarvates/botte-secrete)
- **License:** MIT
- **Homepage:** https://github.com/zedarvates/botte-secrete
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet — be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.