Hydrogen
LLM Proxy
Self-hosted LLM proxy

HydrogenLLM Proxy

One endpoint for every model — retry, fall back and orchestrate, without touching a line of client code.

Hydrogen holds your provider API keys and decides, per request, which model at which provider actually serves it. Clients name a Model Service, never a real model. It speaks both the OpenAI and Anthropic wire formats and translates between them.

release v1.8.0-rc.2 MIT OpenAI + Anthropic Node ≥ 20 SQLite TypeScript
2API formats, translatedOpenAI and Anthropic, in either direction
8model types supportedchat, image, video, speech, embedding, rerank
3model servicing powersoverrides · micro agents · automatic failover

How it fits together

Clients never name a real model. They name a Model Service, and Hydrogen decides what that means — per request, over your own catalogue.

01 / 05

What a request resolves through

  1. Client requestmodel = "sonnet-any"only Model Services are exposed to clients
  2. Model Serviceordered steps: try this, else that
  3. Modelyour internal name, e.g. sonnet4.6
  4. Providerbase URL + encrypted API key
  5. Upstream model idwhat the provider calls it, e.g. claude-sonnet-4-6

Four concepts, each doing one job

ConceptWhat it isExample
ProviderAn upstream endpoint and its API key.openai-official, anthropic-official
ModelYour internal name for a model.sonnet4.6
MappingWhich provider serves that model, under what id.sonnet4.6 → anthropic-official as claude-sonnet-4-6
Model ServiceThe name clients request, and the rules behind it.sonnet-any

Each step of a Model Service pins one explicit (model, provider) pair, with its own retry count, retry interval, and the failure classes that retry or advance. Provider fallback and model fallback are both just “add another step”. If every step is exhausted, the real upstream error goes back to the client, translated into the client's own wire format.

Three services, three behaviours

Model ServiceBehaviour
sonnet-anytry sonnet4.6 @ anthropic → on failure fall back to gpt5.4 @ openai
sonnet-persisttry sonnet4.6 @ anthropic, retrying 5× at 1 s intervals
essay Micro Agentdraft → critique → revise, returning the draft when the critique approves

The payoff: swapping a provider, adding a fallback, capping a thinking budget or inserting a whole agent pipeline is a dashboard edit. Client code keeps asking for sonnet-any.

What you get

Everything is one container with SQLite inside. Provider keys are encrypted at rest, client keys are hashed, and the whole instance exports to a single passphrase-sealed file.

02 / 05

Two wire formats, translated both ways

OpenAI Chat Completions, OpenAI Responses and Anthropic Messages — streaming and non-streaming, tool calls, images, and thinking blocks round-tripped rather than dropped.

Eight kinds of service

chat, ocr, image, video, tts, stt, embedding, rerank — the non-chat ones are OpenAI-style passthroughs that still run your step chain.

Micro Agents

Forward-only stage pipelines with conditions and routers. Each stage runs a saved Model Service, so every stage inherits that service's resilience. Nesting allowed, cycles rejected.

Per-step overrides

Temperature, top-p/top-k, max tokens, stop sequences, thinking level, a system override, plus arbitrary extra body params — pinned per step, not per client.

Reliable streaming

Buffer the upstream stream, treat a truncation as a retryable failure, and replay a complete response — or a clean 502, never half of one.

Image OCR with a cache

A vision pre-pass transcribes images to text for downstream stages; descriptions are keyed by image hash and reused under an LRU byte budget.

API keys with teeth

Scope a key to specific services, cap its requests and tokens, set an expiry, and let its holder check its own status on the public Key Check page.

Observability

Every request logged with each attempt, payloads, latency and token usage — agent stages nested under the client request — plus live Active Requests and dashboard stats.

Backup & restore

The whole instance in one passphrase-sealed file that restores onto any other Hydrogen. Restore is one transaction: it fully succeeds or changes nothing.

Safe by default

AES-256-GCM for provider keys, argon2id for passwords, SHA-256 for client keys, an SSRF guard on provider base URLs, and a boot-time master-key sentinel.

Two roles, two languages

admin and manager — managers run everything except issuing API keys and Settings. Dashboard in English or 中文.

One container, one port

Dashboard and API on the same port, SQLite inside the image. The database and the master key live in /data — persist that and the rest is disposable.

Micro Agents

Composable pipelines — routing, evaluation, image OCR, nested agents — built on your Model Services.

03 / 05

A Model Service routes one request to one upstream call. A Micro Agent runs several model calls — stages — and still presents itself to clients as one model name. Because a Micro Agent is a Model Service as far as the outside world is concerned, it needs no client support at all: point an existing app at essay instead of sonnet-any, and it transparently gets a draft → critique → revise loop.

Stages run top to bottom. After each one, its transitions are checked in order and the first match wins. No match means fall through to the next stage; running off the end finishes the agent, returning the output of the stage where it stopped.

  • Transitions are forward-only. Skip ahead, never back. No loops, so a fixed-length pipeline always terminates; validation rejects a backwards goto.
  • Each stage runs a saved Model Service — or another Micro Agent — so every stage inherits that service's retry and fallback rules for free.
  • The input list is the prompt engineering. Leave it empty and the stage sees the original conversation; add blocks and you compose the messages yourself, in order.
Name (exposed model name)essay
Per-attempt timeout (ms)120000
agent: draft → critique → revise (branching)
Translate images to text (OCR pre-pass)

Before the chain runs, every image in the request is sent to a multimodal/OCR model and replaced with its transcription — so text-only stage models can process the request. Descriptions are cached by image hash.

Stages (run from the top; transitions branch to later stages)
Stage 1draft
Runssonnet-any
Inputempty — the original messages pass straight through, exactly as the client sent them
Transitionsnone — fall through to stage 2
Stage 2critique
Runsfast-cheap
Input
Last user text only Output from stage draft as assistant user “Critique the answer above against the request. If it is already correct and complete, reply with exactly APPROVED and nothing else. Otherwise list the specific flaws, one per line.”
Advanced
system “You are a strict technical editor. Be concise and specific.” tools listed but not callable temperature 0
Transitions
when output contains APPROVED go to end, returning draft
Stage 3revise
Runssonnet-any
Input
Last user text only Output from stage draft as assistant user “A reviewer raised these points:” Output from stage critique as user user “Rewrite your answer to address every point. Reply with the rewritten answer only.”
Transitionsnone — last stage, so its output is what the client gets
Raw JSON Workflow is valid

If the critique says APPROVED, the agent stops at stage 2 and returns the draft — stage 1's text, not the word “APPROVED”. Otherwise it falls through to revise, the last stage, whose output is the response. The client sees one ordinary answer either way, and Logs shows every stage as its own call, nested under the one client request.

Arrangement 01

Re-evaluate before response

Several models think it through together.

One model drafts, a second argues with it, and the draft comes back sharper. More than one mind on every answer — the client only ever sees the final one.

draft · strong model critique · another strong model critique says “APPROVED” → return draft → end else fall through ↓ revise · strong model → end

Costs one extra call. The second model can be shown your tools without being able to call them.

Arrangement 02

Pay by difficulty

A cheap model picks the expensive one.

Stage 1 runs your cheapest model and does nothing but grade the request. Its verdict picks the path — hard work goes to the strong model, the rest falls through to the cheap one.

triage · cheap model output contains “HARD” → strong model → end else fall through ↓ answer · cheap model → end

Grading costs one cheap call. Branch on a regex instead and a router stage spends nothing at all.

Call it

Point any OpenAI SDK at /v1, any Anthropic SDK at the root, and set model to a Model Service name. The client's format and the provider's format are independent.

04 / 05

Endpoints

OpenAI base URL
http://localhost:8080/v1
Chat Completions and Responses hang off this base. Use a Model Service name as the model.
Anthropic base URL
http://localhost:8080
Messages is appended by the SDK. Use a Model Service name as the model.
OpenAI wire format
curl http://localhost:8080/v1/chat/completions \
  -H "Authorization: Bearer sk-hproxy-..." \
  -H "content-type: application/json" \
  -d '{"model":"sonnet-any","messages":[{"role":"user","content":"hello"}]}'
Anthropic wire format — same service, same upstreams
curl http://localhost:8080/v1/messages \
  -H "x-api-key: sk-hproxy-..." \
  -H "anthropic-version: 2023-06-01" \
  -H "content-type: application/json" \
  -d '{"model":"sonnet-any","max_tokens":256,"messages":[{"role":"user","content":"hello"}]}'

Client-facing routes

MethodPathCategoryNotes
POST/v1/chat/completionschatOpenAI Chat Completions, streaming + non-streaming
POST/v1/responseschatOpenAI Responses API
POST/v1/messageschatAnthropic Messages
GET/v1/models—Your Model Services (Anthropic shape if anthropic-version is sent)
POST/v1/embeddingsembeddingOpenAI-compatible providers
POST/v1/rerankrerank
POST/v1/images/generationsimage
POST/v1/videosvideoReturns a job id carrying its own routing suffix
GET/v1/videos/:id · /v1/videos/:id/contentvideoPoll and download, statelessly routed
POST/v1/audio/speechttsBinary audio streamed through
POST/v1/audio/transcriptionssttmultipart, forwarded with model rewritten

Because /v1/models returns Model Services, any tool with a model picker shows your service names — which is the intent: sonnet-any is the model, as far as a client is concerned. Authenticate with Authorization: Bearer … or x-api-key: ….

Deploy

Three paths, in order of effort. Whatever the path, one rule outranks the rest.

05 / 05

/data must be persistent. It holds the SQLite database and hydrogen-secrets.json, which carries the master key that decrypts your provider API keys. Lose it and Hydrogen refuses to boot rather than run with keys it can no longer read.

1

Rainyun app store

One click, no server to run. Managed, with a persistent volume and HTTPS.

Hydrogen is published as a Rainyun Cloud Application. Search for Hydrogen in the store, keep the /data volume the template ships with, and leave the master key and session secret empty so Hydrogen generates and persists them itself.

2

Container image

Any VPS or home server. Pull from GHCR; pin a tag for anything real.
docker run -d --name hydrogen \
  -p 8080:8080 \
  -v hydrogen-data:/data \
  ghcr.io/arrosam/hydrogen-llm-proxy:v1.7.3

Then docker logs hydrogen — the initial admin credentials are printed in a banner — and open port 8080. A ready compose stack with a Caddy TLS overlay lives in deploy/vps/.

3

Build from source

Local development, or a patched build of your own. Node 20+.
git clone https://github.com/Arrosam/Hydrogen-LLM-proxy.git
cd Hydrogen-LLM-proxy
cp .env.example .env
docker compose up -d --build

The repo-root compose file builds from the working tree instead of pulling. Without Docker: npm install, npm run build, then start the server with DATA_DIR set.

Safe by default AES-256-GCM provider keys argon2id passwords SHA-256 client keys, shown once SSRF guard on base URLs Master-key sentinel at boot
Copied to clipboard