A homelab LLM stack for $0*: LiteLLM, Hermes Agent and Open WebUI
How a LiteLLM proxy with per-key budgets, a locked-down Hermes Agent in Telegram and an Open WebUI chat came to run on free tiers, and where the free ends.
The asterisk in the title is doing a lot of work. Over three evenings in August a LiteLLM proxy, an agent that reads my Proxmox node and complains about it in Telegram, and a chat UI went from nothing to running, on API tiers that cost nothing, with a spend meter that has since ticked up to fifty-five cents nobody has billed me for. This is how it is wired, what the restrictions on the agent actually protect (less than they look), and where free stops being free.
Tested on: LiteLLM
main-v1.83.14-stablewith Postgres 17, Hermes Agent v0.20.1 (2026.8.13), Open WebUI v0.11.0, Traefik v3.7.5, all in Docker 29.6 on one EL9 host; models on the Gemini API free tier (3.5 Flash, 3.1 Flash-Lite) and Groq’s free tier (Llama 3.3 70B); the agent reads a standalone Proxmox VE 9 node. Written August 2026. The complete compose files, configs and.envtemplates are in the appendix.
LiteLLM first: aliases, keys with budgets, a fallback chain
It started with scripts, not agents. proxmox_review.py pulls the node’s state from the Proxmox API, computes findings deterministically (a backup older than a week is a finding whether or not a model agrees) and asks a model to explain and prioritise them; guest_audit.py does the same for VMs and containers. Each wanted an API key, each wanted to be pointed at “the model”, and each would happily burn a free tier’s requests-per-minute if I ran two of them at once. So before the second script existed, a proxy did.
LiteLLM speaks the OpenAI API on one side and about a hundred providers on the other; the parts I cared about are smaller than the feature list. An alias per role instead of a provider name in the code. Virtual keys with budgets and rate limits. A fallback chain for when a free tier says 429. One place that remembers what everything spent. The rule I apply to package managers applies here too: if three things need the same secret, none of them should hold it. They hold a LiteLLM key that can spend ten dollars a month; LiteLLM holds the real ones.
One container plus Postgres (keys, spend logs and the admin UI live in the database), an internal network, no published ports, Traefik in front. Trimmed to what matters:
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
services:
litellm:
image: ghcr.io/berriai/litellm:main-v1.83.14-stable # pin it; main-latest wobbles
volumes: ["./data/config.yaml:/app/config.yaml:ro"]
command: ["--config", "/app/config.yaml"]
env_file: [.env] # provider keys, master key, DB pw
environment:
DATABASE_URL: postgresql://litellm:${POSTGRES_PASSWORD}@db:5432/litellm
LITELLM_TELEMETRY: "False"
networks: [internal, traefik] # no ports:; Traefik is the door
db:
image: postgres:17-alpine
networks: [internal]
networks:
internal: { internal: true } # no route out; providers are reached via traefik
traefik: { external: true }
The config is where the design lives: four aliases over three free-tier models, each deployment with its own rpm so the router stops before the provider does, and fallback chains that go Gemini → Groq → Flash-Lite. Keys are inlined in mine (os.environ/… is the tidier form) and two retry settings are trimmed here; the file as it runs is in the appendix. Yes, the two Gemini lanes add up to 14 against a 10 RPM ceiling; the bill section says what that costs.
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
model_list:
- model_name: analyst # what the scripts ask for
litellm_params:
model: gemini/gemini-3.5-flash
api_key: os.environ/GEMINI_API_KEY
rpm: 8 # free tier shows 10 RPM; stay under it
- model_name: hermes # same model, own lane for the agent (Gotchas)
litellm_params:
model: gemini/gemini-3.5-flash
api_key: os.environ/GEMINI_API_KEY
rpm: 6
- model_name: analyst-groq
litellm_params:
model: groq/llama-3.3-70b-versatile
api_key: os.environ/GROQ_API_KEY
rpm: 25
- model_name: analyst-lite
litellm_params:
model: gemini/gemini-3.1-flash-lite
api_key: os.environ/GEMINI_API_KEY
rpm: 12
litellm_settings:
fallbacks:
- analyst: ["analyst-groq", "analyst-lite"]
- analyst-groq: ["analyst-lite"]
- hermes: ["analyst-lite"] # deliberately not through Groq
cache: true # a repeated prompt costs no quota
cache_params: { type: local, ttl: 3600 }
redact_user_api_key_info: true
general_settings:
master_key: os.environ/LITELLM_MASTER_KEY
store_model_in_db: true
store_prompts_in_spend_logs: true # else the UI logs show no request/response
router_settings:
retry_after: 5
allowed_fails: 3
cooldown_time: 60
Consumers never see a provider key. Each gets a virtual key minted with the master key, restricted to its aliases, rate-limited, and given a budget that resets monthly:
1
2
3
4
5
6
7
curl -sS -X POST "$PROXY_URL/key/generate" \
-H "Authorization: Bearer $LITELLM_MASTER_KEY" -H "Content-Type: application/json" \
-d '{"key_alias": "hermes",
"models": ["hermes"],
"rpm_limit": 20, "tpm_limit": 300000,
"max_budget": 10, "budget_duration": "30d",
"metadata": {"owner": "agents", "purpose": "Hermes Agent pilot"}}'
The budget is a fuse, not a thermostat. When the ten dollars are gone the proxy answers 400 with ExceededTokenBudget and the agent shows that error in Telegram, which is exactly the failure I want from a hobby project at three in the morning. A small script issues the keys idempotently and lists them; this listing is from the day the post went out:
1
2
3
4
5
6
$ ./create-agent-keys.sh --list
ALIAS SPEND$ BUDGET$ RPM TPM MODELS CREATED
open-webui 0.014604 10.0 20 300000 analyst,analyst-groq,analyst-lite 2026-08-16T19:53
hermes 0.15691241 10.0 20 300000 hermes,analyst,analyst-groq,analyst-lite 2026-08-15T20:13
agents 0.37382563 10.0 20 300000 analyst,analyst-groq,analyst-lite 2026-08-15T17:11
test-agent 0.00147964 - 4 60000 analyst,analyst-groq,analyst-lite 2026-08-14T22:48
Corrected after publication. The
hermeskey in that listing allowed theanalyst*aliases as well, which made “the agent does not go through Groq” a routing convention rather than a boundary. LiteLLM enforces themodelslist on the key, so the key now allowshermesonly, as the example above shows. Thanks to the reader who noticed.
From a hand-rolled Telegram bot to Hermes Agent
The first interface was a bot I wrote myself: long polling, a dozen slash commands (/pve, /audit 110 deep, /triage, /last), an allow list of user ids, every command a subprocess around one of the scripts. It worked on the first evening. It also meant every new capability was a new handler, and any conversation beyond a command was me pasting a report back into the model with /discuss. It was a remote control. I wanted a colleague.
Hermes Agent, from Nous Research, is an MIT-licensed agent runtime: an agentic loop over tools, memory across sessions, skills as Markdown procedures, a cron scheduler, gateways to Telegram and twenty-odd other messengers. Two things made it a fit rather than a science project. model.provider: custom with a base_url means it talks to LiteLLM like everything else, through its own key with its own budget. And its terminal tool runs my scripts, so the division of labour stays: scripts collect facts and count findings, the model reads, explains and prioritises. Nobody asked a language model to add up gigabytes, and nobody should.
flowchart LR
TG[Telegram] --> H[Hermes Agent<br/>gateway · skills · memory]
H --> S[proxmox_review.py<br/>guest_audit.py]
S --> PVE[(Proxmox API<br/>audit-only token)]
H --> L[LiteLLM<br/>aliases · keys · budgets]
S --> L
OW[Open WebUI] --> L
L --> G[Gemini 3.5 Flash<br/>free tier]
L --> Q[Groq · Llama 3.3 70B<br/>free tier]
L --> FL[Gemini 3.1 Flash-Lite<br/>free tier]
The bot was decommissioned on the third evening. Its scripts moved into the Hermes project as read-only tools behind four skills (proxmox-status, proxmox-review, proxmox-guest-audit, proxmox-diff) plus one that reads an Alertmanager archive on Monday mornings. The persona is called Proxmox Sentinel, which sounds a great deal more serious than “reads the API and complains once a week”.
Locking Hermes down, and what the locks actually hold
The agent’s config reads like a lot of security. Read it as three layers with different jobs.
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
model:
provider: custom
model: hermes # the LiteLLM alias
base_url: http://litellm:4000/v1
api_key: ${OPENAI_API_KEY} # the virtual key `hermes`, from .env
max_tokens: 16000 # a thinking model: this includes the reasoning
reasoning_effort: low
agent:
max_turns: 40
disabled_toolsets: [web, search, browser, vision, file, cronjob, code_execution, …]
platform_toolsets:
telegram: [terminal, skills, memory, todo, clarify, session_search]
terminal:
backend: local # commands run inside this container
cwd: /opt/data/work
env_passthrough: [PVE_URL, PVE_TOKEN_ID, PVE_TOKEN_SECRET, LITELLM_KEY, …]
approvals:
mode: manual # "dangerous" waits for /approve in the chat
cron_mode: deny
deny: ["*pvesh set*", "*qm stop*", "*pct destroy*", "*ssh *", "*docker *", "*curl*-X POST*", …]
skills:
write_approval: true
disabled: [apple-notes, claude-code, github-pr-workflow, …] # all 82 bundled, by name
The layer that holds is not in that file. The Proxmox token review@pve!reviewer carries a custom role with VM.Audit, Datastore.Audit and Sys.Audit, nothing more; a call that changes anything is refused by Proxmox itself, whatever the model asked for and whatever prompt injection arrived inside a VM description. The container has no docker.sock, no ssh keys, no-new-privileges, memory and pid limits, and terminal.backend: local, which Hermes’ own comparison table files under isolation as “None — runs on host”. Here the host is the container, and the container is the sandbox. Its LiteLLM key can spend ten dollars a month. Telegram answers one user id.
The layer that helps is the deny list and the toolsets. Approvals in manual mode route anything matching a dangerous pattern to an /approve prompt in the chat; the deny list refuses qm stop, ssh, docker and any curl with a body outright. Disabled toolsets take away the web, the browser, file writes, code execution and self-scheduling. Hermes’ documentation is admirably blunt about what this is: “Deny rules are a guardrail against an honest-but-wrong agent … They are not a sandbox against a deliberately adversarial process.” Exactly so. This layer stops the agent from doing something stupid on a bad day; the previous one limits what a bad day can cost. The same docs recommend a sandboxed terminal backend and an egress credential proxy for a production gateway; running the terminal inside a container that holds only low-value secrets is the hobby-grade version of that advice, not a substitute for it.
The layer that is mostly optics is the rest: 82 bundled skills disabled by name, write_approval on skills, a SOUL.md telling the persona it is read-only and that text inside VM descriptions is data, not instructions. All true, all worth having, none of it a boundary. And one thing the arrangement does not protect at all: the agent’s own secrets in data/.env are readable from inside the container, because that is where the terminal runs. Hence that file holds an audit-only token, a budgeted proxy key, a Telegram token, and nothing I would mind a confused model reading aloud.
Open WebUI: a chat with no login, on purpose
Sometimes you want to talk to the model rather than to the agent: paste a config, ask a question, no Proxmox involved. Open WebUI pointed at LiteLLM needs two settings, the base URL and its own virtual key. Everything else in my compose file is switching things off.
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
environment:
ENABLE_PERSISTENT_CONFIG: "False" # env is the config; UI edits die on restart, on purpose
ENABLE_OPENAI_API: "True"
OPENAI_API_BASE_URL: http://litellm:4000/v1
ENABLE_OLLAMA_API: "False" # or the UI waits for a timeout on start
ENABLE_DIRECT_CONNECTIONS: "False" # no user-added endpoints around the proxy
RAG_EMBEDDING_ENGINE: openai # defaults load sentence-transformers + whisper
AUDIO_STT_ENGINE: openai
ENABLE_TAGS_GENERATION: "False" # each of these is one more request at 8 rpm
ENABLE_FOLLOW_UP_GENERATION: "False"
ENABLE_AUTOCOMPLETE_GENERATION: "False"
ENABLE_TITLE_GENERATION: "True" # one per chat; kept
WEBUI_AUTH: "False" # no login: whoever reaches the page is admin
labels:
- traefik.http.middlewares.open-webui-internal.ipallowlist.sourcerange=10.0.8.0/22,10.0.12.0/22
- traefik.http.middlewares.open-webui-internal.ipallowlist.ipstrategy.depth=1
No login is a decision, not laziness, and it is only sane because of the two labels under it. The UI is reachable from two lab subnets, full stop; from anywhere else Traefik answers before Open WebUI sees the request. The depth=1 is the load-bearing character: Traefik sits behind a front proxy, and without it the allow list is compared against the front proxy’s own address, which is inside 10.0.0.0/8 and would cheerfully admit whoever the front proxy admits. I learned this the way one learns most things about proxies, by putting one in front of another and being surprised; the next section is about why the trust chain holds once it is set.
First load: Open WebUI logs you into admin@localhost, because there is nobody else to be. The model list is whatever the LiteLLM key may see.
The generation switches are about quota, not taste. With the defaults, one message costs a chat completion plus a title, tags, follow-up suggestions and, if you type slowly enough, autocomplete: on an alias limited to 8 requests a minute, a reliable way to rate-limit yourself out of your own chat.
Traefik at the edge, and the proxy in front of Traefik
Both LiteLLM and Open WebUI publish no ports. They join a Docker network called traefik and opt in with labels; one Traefik container owns 80 and 443 on the host and does TLS, routing and the allow lists for everything on the box, which is a lot of responsibility for 256 MB of RAM. The static configuration is a list of flags in its compose file; the parts that matter here:
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
command:
- "--providers.docker=true"
- "--providers.docker.exposedbydefault=false" # containers opt in with traefik.enable=true
- "--providers.docker.network=traefik"
- "--providers.file.directory=/etc/traefik/dynamic" # shared middlewares, hot-reloaded
- "--providers.file.watch=true"
- "--entrypoints.web.address=:80"
- "--entrypoints.websecure.address=:443"
# trust X-Forwarded-* only from the front proxy; from anyone else they are dropped
- "--entryPoints.websecure.forwardedHeaders.trustedIPs=10.0.12.30/32"
# Let's Encrypt via Cloudflare DNS-01: no port 80 exposure, wildcard-capable
- "--certificatesresolvers.letsencrypt.acme.dnschallenge.provider=cloudflare"
- "--certificatesresolvers.letsencrypt.acme.storage=/etc/traefik/acme/acme-dnschallenge.json"
ports:
- { target: 443, published: 443, mode: host, host_ip: 10.0.12.50 } # host mode: real source IPs
Three habits from that file carry the whole post. Services opt in (exposedbydefault=false), so a container that forgets its labels is invisible rather than accidentally public. Shared middlewares live in the file provider: robots-tag-header@file on everything, redirect-to-https, forward-auth for the services with single sign-on (not these two), an LDAP plugin on Traefik’s own dashboard. And certificates come from ACME DNS-01 against Cloudflare, so nothing needs to be reachable on port 80 for a certificate to exist.
The line that earns its keep is forwardedHeaders.trustedIPs. There is a proxy in front of Traefik: the router forwards 443 to an nginx at 10.0.12.30, only from Cloudflare’s address ranges, and nginx appends the client to X-Forwarded-For. Traefik honours forwarded headers from that one address and strips them from everyone else, so a LAN client cannot forge its way past an allow list; a forged header sent through nginx just gets the real address appended after it, which is why depth counts from the right. Internal DNS points the lab subnets at the same nginx, so my laptop arrives with a one-hop header and depth=1 resolves to my laptop; the internet arrives via Cloudflare with the edge as the last hop, which no allow list here contains. Two proxies, one rule about whom to believe, and the “illusion of security” from the Hermes section becomes an actual one, at least on the question of who can open the chat.
The bill: what free costs
LiteLLM keeps a spend column priced from its own cost table, so I can say precisely what this stack has cost: $0.55 across all keys since the 14th, of which Google and Groq have invoiced $0.00. It is the most exact accounting of nothing I have seen, and it makes the budgets meaningful: a key with max_budget: 10 blows at ten dollars of list price, charged or not.
The asterisks:
- Requests per minute are the currency, not dollars. The Gemini free tier is per project, and Google no longer prints the numbers in the docs; AI Studio shows them, and mine says 10 RPM and 1,500 requests a day for 3.5 Flash. Groq allows 30 RPM and 1,000 a day for Llama 3.3 70B. My two Gemini aliases add up to 14 RPM against a 10 RPM ceiling; when they collide, LiteLLM cools the deployment down and falls through to Flash-Lite. A pilot survives that. A second user would not.
- Free means reviewable. Google’s terms for the unpaid tier say submitted content is used to “provide, improve, and develop Google products and services” and that “human reviewers may read, annotate, and process your API input and output”. A Proxmox review contains hostnames, VM names and disk sizes: nothing I would put a password next to, and one more reason the agent’s environment holds no secret that matters.
- Latency is what you pay with. A thinking model with
reasoning_effort: lowanswers a review in about 30 seconds; free-tier retries and cooldowns add tens of seconds when they hit. Fine for a Monday digest, poor for a chat. - Free tiers move. Limits, model names and terms have all shifted since the first evening; hence the versions and the date at the top.
How much work that buys, the spend logs can answer, because LiteLLM keeps every request with its token counts. Three days, all consumers, 100 requests: 540,000 prompt tokens and 56,000 completion tokens, priced at $0.55; 87 of them on the one evening when everything was built and re-run. An agent turn (a question in Telegram, or one step of a review) is the expensive unit and the typical request: the median request carries about 7,000 prompt tokens, because the context brings the system prompt, the skills, the memory and the tool output along, and gets a couple of hundred back. One review conversation, by my notes, was 11 calls. A chat message in Open WebUI is about 1,250 tokens in and 330 out. Against the allowances as of August 2026 (Gemini: 10 RPM and 1,500 requests a day per project; Groq: 30 RPM, 12,000 tokens a minute and 100,000 a day per organisation):
| Workload | Requests / tokens | Gemini | Groq |
|---|---|---|---|
| Monday digest | 1–2 / ~10K | trivial | fine |
| A question to the agent | 1–3 / ~7K each | fine | 1 turn a minute, ~14 a day |
| Full review via the agent | ~10 / 40–80K | ~150 a day on paper | 1 a day, then TPD is spent |
| 60 chat messages | ~65 / ~95K | ~4% of a day | roughly one day’s tokens |
Stated plainly: on Groq the walls are tokens per minute and per day, not requests; a single review would exhaust its daily allowance, which is why Groq is a fallback for chat and not for the agent, quite apart from the bug in the Gotchas. On Gemini it is the other way round: the daily 1,500 is comfortable for one engineer, and what a second person breaks is the 10 a minute, because two agent conversations at once already exceed it. Flash-Lite is the overflow lane; on the build evening it absorbed 43 agent calls and 418,000 tokens while Flash was cooling down, which is what fallbacks are for.
Acceptable, then, for one engineer, one node, one weekly report and the occasional question. Priced, not paid for.
What else can hang off the proxy
Once every consumer talks to litellm:4000 with its own key, adding a capability is a change on the proxy rather than in three places. What I would actually reach for, in order:
- Local models under the same aliases:
ollama/…deployments inmodel_list, soanalystbecomes a local model for sensitive prompts and stays Gemini for the rest, without touching the agent. - A webhook to Telegram when a provider is cooling down or a key crosses its budget (alerting); today I find out from a 429 in the report.
- Guardrails: PII masking and prompt-injection detection as pre-call hooks on the
hermeskey, the one that reads text out of VM descriptions. - Teams and model access groups, the day someone else in the flat wants a key with a smaller budget.
- An MCP gateway, so tools share the same keys and logs, and logging to Langfuse or OpenTelemetry once spend logs stop being enough. Both untried here.
Appendix: the complete files
Everything above is excerpted. The complete files as they run on the host are published alongside this post: secrets are placeholders, hostnames come from .env, comments are in English, and the fixed container addresses are my 10.0.4.0/24 habit rather than a requirement. Not included: the front nginx (a stock reverse proxy that adds X-Forwarded-For), the LDAP filter and internal CA for the dashboard login, and the 150-line key-issuing script, whose two interesting calls (/key/generate, /key/list) are in the text above.
| Stack | Files |
|---|---|
| Traefik | docker-compose.yaml · dynamic/http.yaml · dynamic/global.yaml · dynamic/tcp.yaml · env.example · network.sh |
| LiteLLM | docker-compose.yaml · config.yaml · env.example |
| Open WebUI | docker-compose.yaml · env.example |
| Hermes Agent | docker-compose.yaml · config.yaml · env.example |
Or take the lot:
1
2
3
4
5
6
7
8
base=https://blog.srepowered.com/assets/files/litellm-homelab-stack
for f in traefik/docker-compose.yaml traefik/env.example traefik/network.sh \
traefik/dynamic/http.yaml traefik/dynamic/global.yaml traefik/dynamic/tcp.yaml \
litellm/docker-compose.yaml litellm/config.yaml litellm/env.example \
open-webui/docker-compose.yaml open-webui/env.example \
hermes/docker-compose.yaml hermes/config.yaml hermes/env.example; do
curl -fsS --create-dirs -o "$f" "$base/$f"
done
Bring-up order: the traefik network and Traefik; LiteLLM and its Postgres; mint the virtual keys; then Open WebUI and Hermes, each with its own key. Hermes also needs the scripts and skills from its own project, a Telegram bot token from BotFather, and a Proxmox API token bound to an audit-only role.
Gotchas
Groq as a fallback broke the agent loop. Hermes streams with tool calls; on those, Llama via Groq answers
finish_reason=tool_use_failed, and LiteLLM 1.83 turned that into a 500 (int("tool_use_failed")in the stream builder). Hermes retried, 8 of 11 calls went round again, a reply took 25–110 s. The fix is a second alias for the agent whose fallback skips Groq (hermes: [analyst-lite]): no errors since, 1–4 s a call. The same class of bug, a non-standardfinish_reasonbreaking the stream builder, is on record as litellm#22671.
max_tokenson a thinking model includes the thinking. With 4,000 the model spent most of it reasoning and left about 160 tokens for the report, which arrived as a fragment from its middle. 16,000 withreasoning_effort: lowgives 30 s instead of 130 s and the same findings; the scripts now append a warning whenfinish_reasonislengthinstead of letting it pass.
disabled_toolsetssubtracts tools, not names.debuggingis a composite that containsterminal; disable it and the terminal is gone from every platform. The model said, politely, that it had no terminal tool. Keep composites (debugging,coding,safe) out of the list.
WEBUI_AUTH=Falseonly works on an empty database. Set it with an existing user and every request gets HTTP 400 (“You can’t turn off authentication because there are existing users”, open-webui#9896). Flipping the flag back is not enough; wipedata/and start clean, after a backup.
An IP allow list behind a front proxy allows everyone the proxy allows until
ipstrategy.depth=1is set. Test the list from outside the range before trusting it.
Virtual keys are stored hashed. The secret is visible once, in the
/key/generateresponse; save it then or plan to rotate.
Takeaways
Proxy first, agents second. Everything that went well here came from having one door with a meter on it before anything walked through, and everything that went wrong was found by that meter. The restrictions on the agent are worth having and worth being honest about: they make an honest model safer, and the read-only token, the container and the budget cut what a dishonest one could do down to what the container holds, which is why what it holds is chosen so carefully. Next time the ten-dollar budget goes on the very first key, before the first script, and the alert webhook goes in before it is needed. It is an excellent deal for one node and one engineer, asterisks and all.
References
- LiteLLM: virtual keys, budgets and rate limits, fallbacks and cooldowns — the three pages this setup is built from
- LiteLLM: caching, alerting, guardrails, MCP, releases
- Hermes Agent: configuration, Docker, security — provider
custom, toolsets, approvals, the “honest-but-wrong” sentence - Open WebUI: environment variables; issue #9896 on
WEBUI_AUTH - Gemini API rate limits and terms; Groq rate limits
- Traefik IPAllowList —
ipStrategy.depth; entrypoints —forwardedHeaders.trustedIPs; Docker provider, file provider, ACME resolvers - Proxmox VE: user and permission management — building an audit-only role
