Skip to content
All projects

Hermes and the self-hosted AI stack

Private repo

A personal assistant running on my own hardware, with a hard boundary around what it can touch.

A model gateway with per-function keys, fail-closed tool access, and MCP servers scoped so no caller sees more than it needs. Seventeen containers on one mini PC, behind a dual-ISP firewall.

LiteLLMMCPProxmoxpfSenseHome AssistantCloudflare Tunnel
gateway keys, scoped by function
5

gateway keys, scoped by function

LXC containers
17

LXC containers

MCP servers, one self-built
3

MCP servers, one self-built

ISPs, load balanced
2

ISPs, load balanced

A gateway with five keys, not one

Every model call goes through a LiteLLM gateway holding five virtual keys, split by what the caller is for rather than who the caller is. Calendar, contacts, home automation, the assistant core, and general tooling each get their own. They share a user id so spend rolls up into one number, and each can be rate limited independently.

That split is the thing I would defend hardest. The moment an assistant can read your mail and drive your house on the same credential, you have no way to reason about blast radius, and no way to answer why the bill moved.

Access fails closed, and the restriction lives in the right place

The gateway refuses any key that has not been explicitly granted a tool surface. Nothing works by default, which is the only posture where forgetting something is safe.

Restriction is enforced where tools are registered rather than on the key itself. That distinction matters: a client cannot discover or call anything outside its envelope regardless of what it sends, because the tools were never in its surface to begin with. The MCP servers additionally accept connections only from the gateway address, so that firewall rule is a load-bearing control rather than a comfort.

Tool surface compounds model weakness

This was the most useful thing the project taught me. One MCP server exposes 87 tools, another 104. Handing all of them to a weaker model does not make it more capable, it makes it fail harder, because the size of the choice grows faster than the model’s ability to make it.

The fix is not a better model, it is a smaller surface: several narrow registrations against the same upstream server, each exposing the handful of tools one use case needs. Same server, different doors.

Routing by sensitivity, and prompts shaped for cache

Tool-execution turns go to a cheaper hosted model that happens to be good at them. Conversational turns go to a stronger one. Anything touching mail, calendar, or the house goes to a direct provider rather than a reseller, with a fallback on an entirely different provider so one outage is not a blackout.

Persona and profile are always loaded and capped so they form a stable prefix that caches; everything else loads on demand. The boundaries between those files are drawn by how often the content changes rather than by what it is about, and that alone cut input spend substantially.

The machine underneath

One ASUS mini PC running Proxmox: a Ryzen 5 with 12 threads and 16GB of memory, which is less than people expect for seventeen containers. Each service runs in an LXC cut from one Debian template that already carries Docker and the toolchains, so a new service is a clone away from running.

A Cloudflare tunnel runs as a container inside the lab, so publishing a service is a hostname in a dashboard: no port forward, no inbound hole, nothing listening at the edge. A separate pfSense box with four 2.5 gigabit ports load balances two ISPs and fails over between them, and alerts through three independent paths, the first of which talks to Telegram directly so it still works when my own automation is the thing that broke.

Interested in the rest?

The full history, the architecture decisions, and the code behind this are available on request.

{
  "created_at": "2026-09-14T20:42:21+05:30",
  "updated_at": "2026-09-21T18:10:53+05:30"
}