Install
$ agentstack add skill-timfewi-tentaflake-hermes-provider-setup ✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ✓ Network access No
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ● Environment & secrets Used
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
Verified badge
Passed review? Show it. Paste this badge into your README, it links to the public security report.
Reliability & compatibility
Declared compatibility
Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.
We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.
How agent discovery & health will work →About
Hermes Provider Setup
> Tentaflake context: Inside a tentaflake agent container, you can run > hermes config set interactively — but if your agent was defined with a > settings attrset in my-agents.nix, the config.yaml inside the > container is read-only (mounted :ro). Interactive changes are lost > on restart. For a persistent configuration, put settings in my-agents.nix > instead and rebuild. API keys always go in /run/secrets/hermes-.env > (or agenix), not via hermes config set.
When to Use
- Configure LLM provider for Hermes (first-time setup)
- Switch between providers
- Add multi-provider fallback
- Set up custom/self-hosted endpoint (Ollama, vLLM, SGLang)
- Configure credential pool for multi-key rotation
- Troubleshoot provider connectivity
- Switch models mid-session
Procedure
1. Provider Selection Guide
| Provider | Auth | Best for | | -------------------- | ------------------- | ----------------------------------------------------------------- | | Nous Portal | OAuth (browser) | All-in-one: 300+ models + Tool Gateway (web, image, TTS, browser) | | OpenAI Codex | OAuth (device code) | Codex models, ChatGPT subscribers | | Anthropic | OAuth or API key | Claude models (Max plan + credits) | | OpenRouter | API key | Multi-provider routing, 200+ models | | Google AI Studio | API key | Gemini models | | Custom Endpoint | API key + base URL | Self-hosted (Ollama, vLLM, SGLang, LM Studio) | | DeepSeek | API key | DeepSeek models direct | | AWS Bedrock | IAM role | Claude, Nova within AWS | | GitHub Copilot | OAuth | Copilot subscription models | | xAI | API key or OAuth | Grok models | | z.ai / ZhipuAI | API key | GLM models | | Kimi / Moonshot | API key | Kimi models | | MiniMax | API key or OAuth | MiniMax models |
Minimum context: 64K tokens. Models with smaller windows rejected at startup.
Rule of thumb: If Hermes can't complete a normal chat, don't add features yet. Get one clean conversation working.
2. Fastest Path — Nous Portal
hermes setup --portal
This one command:
- Opens browser for OAuth login (portal.nousresearch.com)
- Stores refresh token at
~/.hermes/auth.json - Lets you pick a model
- Sets Nous as inference provider
- Turns on Tool Gateway (web, image, TTS, browser)
- Ready to
hermes chat
After setup, config looks like:
model:
provider: nous
default: anthropic/claude-sonnet-4.6
base_url: https://inference-api.nousresearch.com/v1
3. API Key Providers
# OpenRouter — most flexible multi-provider
hermes config set OPENROUTER_API_KEY sk-or-...
# Anthropic — direct Claude access
hermes config set ANTHROPIC_API_KEY sk-ant-...
# OpenAI
hermes config set OPENAI_API_KEY sk-...
# Google Gemini
hermes config set GOOGLE_API_KEY AIza...
# or GEMINI_API_KEY
# DeepSeek
hermes config set DEEPSEEK_API_KEY sk-...
# xAI (Grok)
hermes config set XAI_API_KEY ...
# AWS Bedrock — uses IAM role or aws configure
# No API key needed
# GitHub Copilot
hermes config set COPILOT_GITHUB_TOKEN ghp_...
4. OAuth Providers
# Nous Portal
hermes portal # One-shot OAuth setup
hermes portal info # Check login status
hermes portal tools # Tool Gateway catalog
# Anthropic OAuth
hermes auth add anthropic --type oauth
# OpenAI Codex
hermes model → pick OpenAI Codex (device code flow)
# Google Gemini OAuth
hermes model → pick Google Gemini (OAuth)
# GitHub Copilot OAuth
hermes model → pick GitHub Copilot (OAuth)
# MiniMax OAuth
hermes model → pick MiniMax (OAuth)
# xAI Grok OAuth
hermes model → pick xAI Grok OAuth
# Qwen OAuth
hermes model → pick Qwen OAuth
5. Custom Endpoint Configuration
For local/self-hosted models (Ollama, vLLM, SGLang, LM Studio, llama.cpp):
# Interactive
hermes model
# → Select: Custom endpoint
# → Enter base URL: http://localhost:11434/v1
# → Enter API key: ollama (or empty)
# → Enter model name: qwen3.5:27b
# → Context length: 64000
Or direct config:
model:
default: qwen3.5:27b
provider: custom
base_url: http://localhost:11434/v1
With per-model config:
custom_providers:
- name: "My Ollama"
base_url: "http://localhost:11434/v1"
models:
qwen3.5:27b:
context_length: 64000
llama3.2:3b:
context_length: 8192
Works with: Ollama, vLLM, SGLang, LM Studio, LocalAI, llama.cpp server, any OpenAI-compatible API.
6. Multi-Provider Fallback
model:
provider: openrouter
default: anthropic/claude-sonnet-4.6
fallbacks:
- provider: nous
model: anthropic/claude-sonnet-4.6
- provider: anthropic
model: claude-sonnet-4-6
Fallback chain engages on errors (rate limits, 5xx, connection drops). agent.api_max_retries controls retries before fallback:
agent:
api_max_retries: 3 # Default: 3 retries per provider before fallback
Set to 0 for fast failover (don't retry flaky endpoint, switch immediately).
7. Provider Switching
hermes model # Full setup wizard — add/change providers
Inside session:
/model # Interactive picker
/model anthropic/claude-sonnet-4.6 # Switch model
/model openrouter:anthropic/claude-sonnet-4.6 # Switch provider + model
# Useful patterns:
/model google/gemini-3-pro-preview # Large context for long docs
/model openai/gpt-5.4 # Strong reasoning
/model deepseek/deepseek-v4-pro # Cost-effective coding
/model only shows providers already configured. Add new ones with hermes model from terminal.
8. Credential Pool Strategies
Multiple API keys for same provider:
credential_pool_strategies:
openrouter: round_robin # Cycle through keys evenly
anthropic: least_used # Always pick least-used key
custom: fill_first # Use first key until exhausted
| Strategy | Behavior | | ---------------------- | -------------------------------------- | | fill_first (default) | Use first key, fall to next on failure | | round_robin | Cycle keys evenly | | least_used | Always pick least-used | | random | Random selection |
9. Provider Timeouts
providers:
openrouter:
request_timeout_seconds: 1800 # API call timeout
stale_timeout_seconds: 300 # Stale non-stream detection
models:
claude-sonnet-4:
timeout_seconds: 3600 # Model-specific override
stale_timeout_seconds: 600
| Timeout | Default | Config | | ---------------- | ------- | ---------------------------------------- | | Socket read | 120s | HERMES_STREAM_READ_TIMEOUT | | Stale stream | 180s | HERMES_STREAM_STALE_TIMEOUT | | Stale non-stream | 300s | providers..stale_timeout_seconds | | API call | 1800s | providers..request_timeout_seconds |
Local model detection: Hermes auto-raises read timeout to 1800s, disables stale detection for local endpoints.
10. Model Context Length Requirements
Minimum: 64,000 tokens. Set explicitly if auto-detect wrong:
model:
default: your-model-name
context_length: 131072
For local models, ensure server configured for enough context:
# llama.cpp
--ctx-size 65536
# Ollama
ollama run --num_ctx 64000
11. Compression Model Config
Summarization can use different (cheaper) model:
auxiliary:
compression:
provider: openrouter
model: google/gemini-3-flash-preview
base_url: null
Must have context window ≥ main model's. The compressor sends full middle section to summary model — smaller window causes silent context loss.
12. Provider Config Storage
All secrets → ~/.hermes/.env:
OPENROUTER_API_KEY=sk-or-...
ANTHROPIC_API_KEY=sk-ant-...
GOOGLE_API_KEY=AIza...
All non-secrets → ~/.hermes/config.yaml:
model:
provider: openrouter
default: anthropic/claude-sonnet-4.6
hermes config set KEY VAL auto-routes to correct file.
13. Provider Verification
# Test basic chat
hermes chat -q "hello"
# Test with specific provider
hermes chat --provider openrouter -q "hello"
# Test with specific model
hermes chat --model anthropic/claude-sonnet-4.6 -q "hello"
# Check provider config
hermes config show | head -20
# Full diagnostics
hermes doctor
Pitfalls
- Model 400 = name mismatch or plan issue: Verify exact slug with provider docs
/modelonly shows configured providers: Add new providers viahermes modelfrom terminal- Minimum 64K context: Models with smaller windows rejected at startup
- Local model timeouts: Auto-detected and relaxed. Set
HERMES_STREAM_READ_TIMEOUT=1800if needed - OAuth on headless needs workaround: Use paste-back or SSH port forward
- API key conflicts: OpenAI key won't work with OpenRouter. Check
.envfor conflicting entries - Context length auto-detected may be wrong: Set
model.context_lengthexplicitly for custom endpoints - Compression model needs ≥ main model context: Smaller window causes silent context loss
- Fallback too aggressive can cause problems: Keep routing off until base provider is stable
agent.api_max_retries 0skips retries: Fast failover means no retries on transient errorsprovider: customis first-class provider: Not an alias. Setprovider: customin config.yaml- Portal refresh token quarantine: On invalidation, next call shows "re-authentication required". Run
hermes auth add nous
Verification
hermes doctor # Check provider connectivity
hermes chat -q "hello" # Basic chat works
hermes config show | grep -A 10 model # Provider/model config
hermes portal info # (if using Portal)
# In session:
/model # Should show configured providers
Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: timfewi
- Source: timfewi/tentaflake
- License: MIT
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet, be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.