One endpoint for every model — retry, fall back and orchestrate, without touching a line of client code.
Hydrogen holds your provider API keys and decides, per request, which model at which provider actually serves it. Clients name a Model Service, never a real model. It speaks both the OpenAI and Anthropic wire formats and translates between them.
Clients never name a real model. They name a Model Service, and Hydrogen decides what that means — per request, over your own catalogue.
01 / 05
What a request resolves through
Client requestmodel = "sonnet-any"only Model Services are exposed to clients
Model Serviceordered steps: try this, else that
Modelyour internal name, e.g. sonnet4.6
Providerbase URL + encrypted API key
Upstream model idwhat the provider calls it, e.g. claude-sonnet-4-6
Four concepts, each doing one job
Concept
What it is
Example
Provider
An upstream endpoint and its API key.
openai-official, anthropic-official
Model
Your internal name for a model.
sonnet4.6
Mapping
Which provider serves that model, under what id.
sonnet4.6 → anthropic-official as claude-sonnet-4-6
Model Service
The name clients request, and the rules behind it.
sonnet-any
Each step of a Model Service pins one explicit (model, provider) pair, with its own retry count, retry interval, and the failure classes that retry or advance. Provider fallback and model fallback are both just “add another step”. If every step is exhausted, the real upstream error goes back to the client, translated into the client's own wire format.
Three services, three behaviours
Model Service
Behaviour
sonnet-any
try sonnet4.6 @ anthropic → on failure fall back to gpt5.4 @ openai
sonnet-persist
try sonnet4.6 @ anthropic, retrying 5× at 1 s intervals
essayMicro Agent
draft → critique → revise, returning the draft when the critique approves
The payoff: swapping a provider, adding a fallback, capping a thinking budget or inserting a whole agent pipeline is a dashboard edit. Client code keeps asking for sonnet-any.
What you get
Everything is one container with SQLite inside. Provider keys are encrypted at rest, client keys are hashed, and the whole instance exports to a single passphrase-sealed file.
02 / 05
Two wire formats, translated both ways
OpenAI Chat Completions, OpenAI Responses and Anthropic Messages — streaming and non-streaming, tool calls, images, and thinking blocks round-tripped rather than dropped.
Eight kinds of service
chat, ocr, image, video, tts, stt, embedding, rerank — the non-chat ones are OpenAI-style passthroughs that still run your step chain.
Micro Agents
Forward-only stage pipelines with conditions and routers. Each stage runs a saved Model Service, so every stage inherits that service's resilience. Nesting allowed, cycles rejected.
Per-step overrides
Temperature, top-p/top-k, max tokens, stop sequences, thinking level, a system override, plus arbitrary extra body params — pinned per step, not per client.
Reliable streaming
Buffer the upstream stream, treat a truncation as a retryable failure, and replay a complete response — or a clean 502, never half of one.
Image OCR with a cache
A vision pre-pass transcribes images to text for downstream stages; descriptions are keyed by image hash and reused under an LRU byte budget.
API keys with teeth
Scope a key to specific services, cap its requests and tokens, set an expiry, and let its holder check its own status on the public Key Check page.
Observability
Every request logged with each attempt, payloads, latency and token usage — agent stages nested under the client request — plus live Active Requests and dashboard stats.
Backup & restore
The whole instance in one passphrase-sealed file that restores onto any other Hydrogen. Restore is one transaction: it fully succeeds or changes nothing.
Safe by default
AES-256-GCM for provider keys, argon2id for passwords, SHA-256 for client keys, an SSRF guard on provider base URLs, and a boot-time master-key sentinel.
Two roles, two languages
admin and manager — managers run everything except issuing API keys and Settings. Dashboard in English or 中文.
One container, one port
Dashboard and API on the same port, SQLite inside the image. The database and the master key live in /data — persist that and the rest is disposable.
Micro Agents
Composable pipelines — routing, evaluation, image OCR, nested agents — built on your Model Services.
03 / 05
A Model Service routes one request to one upstream call. A Micro Agent runs several model calls — stages — and still presents itself to clients as one model name. Because a Micro Agent is a Model Service as far as the outside world is concerned, it needs no client support at all: point an existing app at essay instead of sonnet-any, and it transparently gets a draft → critique → revise loop.
Stages run top to bottom. After each one, its transitions are checked in order and the first match wins. No match means fall through to the next stage; running off the end finishes the agent, returning the output of the stage where it stopped.
Transitions are forward-only. Skip ahead, never back. No loops, so a fixed-length pipeline always terminates; validation rejects a backwards goto.
Each stage runs a saved Model Service — or another Micro Agent — so every stage inherits that service's retry and fallback rules for free.
The input list is the prompt engineering. Leave it empty and the stage sees the original conversation; add blocks and you compose the messages yourself, in order.
Name (exposed model name)essay
Per-attempt timeout (ms)120000
agent: draft → critique → revise (branching)
Translate images to text (OCR pre-pass)
Before the chain runs, every image in the request is sent to a multimodal/OCR model and replaced with its transcription — so text-only stage models can process the request. Descriptions are cached by image hash.
Stages (run from the top; transitions branch to later stages)
Stage 1draft
Runssonnet-any
Inputempty — the original messages pass straight through, exactly as the client sent them
Transitionsnone — fall through to stage 2
Stage 2critique
Runsfast-cheap
Input
Last user text onlyOutput from stage draftas assistantuser“Critique the answer above against the request. If it is already correct and complete, reply with exactly APPROVED and nothing else. Otherwise list the specific flaws, one per line.”
Advanced
system“You are a strict technical editor. Be concise and specific.”tools listed but not callabletemperature 0
Transitions
when output contains APPROVED go to end, returning draft
Stage 3revise
Runssonnet-any
Input
Last user text onlyOutput from stage draftas assistantuser“A reviewer raised these points:”Output from stage critiqueas useruser“Rewrite your answer to address every point. Reply with the rewritten answer only.”
Transitionsnone — last stage, so its output is what the client gets
If the critique says APPROVED, the agent stops at stage 2 and returns the draft — stage 1's text, not the word “APPROVED”. Otherwise it falls through to revise, the last stage, whose output is the response. The client sees one ordinary answer either way, and Logs shows every stage as its own call, nested under the one client request.
Arrangement 01
Re-evaluate before response
Several models think it through together.
One model drafts, a second argues with it, and the draft comes back sharper. More than one mind on every answer — the client only ever sees the final one.
draft · strong modelcritique · another strong modelcritique says “APPROVED” →return draft → endelse fall through ↓revise · strong model → end
Costs one extra call. The second model can be shown your tools without being able to call them.
Arrangement 02
Pay by difficulty
A cheap model picks the expensive one.
Stage 1 runs your cheapest model and does nothing but grade the request. Its verdict picks the path — hard work goes to the strong model, the rest falls through to the cheap one.
triage · cheap modeloutput contains “HARD” →strong model → endelse fall through ↓answer · cheap model → end
Grading costs one cheap call. Branch on a regex instead and a router stage spends nothing at all.
Call it
Point any OpenAI SDK at /v1, any Anthropic SDK at the root, and set model to a Model Service name. The client's format and the provider's format are independent.
04 / 05
Endpoints
OpenAI base URL
http://localhost:8080/v1
Chat Completions and Responses hang off this base. Use a Model Service name as the model.
Anthropic base URL
http://localhost:8080
Messages is appended by the SDK. Use a Model Service name as the model.
Your Model Services (Anthropic shape if anthropic-version is sent)
POST
/v1/embeddings
embedding
OpenAI-compatible providers
POST
/v1/rerank
rerank
POST
/v1/images/generations
image
POST
/v1/videos
video
Returns a job id carrying its own routing suffix
GET
/v1/videos/:id · /v1/videos/:id/content
video
Poll and download, statelessly routed
POST
/v1/audio/speech
tts
Binary audio streamed through
POST
/v1/audio/transcriptions
stt
multipart, forwarded with model rewritten
Because /v1/models returns Model Services, any tool with a model picker shows your service names — which is the intent: sonnet-anyis the model, as far as a client is concerned. Authenticate with Authorization: Bearer … or x-api-key: ….
Deploy
Three paths, in order of effort. Whatever the path, one rule outranks the rest.
05 / 05
/data must be persistent. It holds the SQLite database andhydrogen-secrets.json, which carries the master key that decrypts your provider API keys. Lose it and Hydrogen refuses to boot rather than run with keys it can no longer read.
1
Rainyun app store
One click, no server to run. Managed, with a persistent volume and HTTPS.
Hydrogen is published as a Rainyun Cloud Application. Search for Hydrogen in the store, keep the /data volume the template ships with, and leave the master key and session secret empty so Hydrogen generates and persists them itself.
Then docker logs hydrogen — the initial admin credentials are printed in a banner — and open port 8080. A ready compose stack with a Caddy TLS overlay lives in deploy/vps/.
Local development, or a patched build of your own. Node 20+.
git clone https://github.com/Arrosam/Hydrogen-LLM-proxy.git
cd Hydrogen-LLM-proxy
cp .env.example .env
docker compose up -d --build
The repo-root compose file builds from the working tree instead of pulling. Without Docker: npm install, npm run build, then start the server with DATA_DIR set.