llmleaf
0.1.6indexedUnifies diverse LLM providers behind a single OpenAI-style API: streaming-first chat/embeddings/voice/realtime/batch modalities, per-model health-aware fallback chains, outbound control-plane auth and extensibility.
Unifies diverse LLM providers behind a single OpenAI-style API: streaming-first chat/embeddings/voice/realtime/batch modalities, per-model health-aware fallback chains, outbound control-plane auth and extensibility.

llmleaf is a llm proxy. It proxies different llm providers and their slighty different apis and converts it to a single api surface (enhanced openai-compatible or anthropic).
The origin of this project is that other AI gateways focus on being all encompassing and fast. I wanted a project that is slim, focused and near native performance instead.

Please use web sockets, it fixes latency and prompt caching issues.
# Run with the embedded dev config (echo provider, key `local-dev:s3cret`)
cargo run -p llmleaf
# …or point at your own config
cargo run -p llmleaf -- llmleaf.toml
Copy llmleaf.example.toml, fill in provider credentials (use env:VAR indirection — secrets
never live in the file), and pass it as the argument. Container image: docker buildx bake image
(listens on :8080). Send a request:
curl localhost:8080/v1/chat/completions \
-H "Authorization: Bearer $(printf 'local-dev:s3cret' | base64)" \
-d '{"model":"demo","messages":[{"role":"user","content":"hi"}]}'
Base64 the
id:passwordcredential with no trailing newline — useprintf(orbase64 -w0), notecho. A stray newline is encoded into the value, so the decoded password becomespw\nand fails the hash check → , even when the configured is correct.
See llmleaf.example.toml for the full configuration surface (providers, routes, keys, control plane).
Consumer endpoints (OpenAI-compatible unless noted):
Read-only admin (optional token): GET /admin/routes, /admin/health, /admin/keys.
Official client SDKs for 6 languages live in clients/.
Native API compaction is opt-in. GET /v1/models includes supports_compaction for every entry.
The flag uses the primary provider's model capabilities and configured API. Without an override,
it is false when support is unconfirmed, including when the provider catalog cannot be fetched.
For an OpenAI-compatible provider whose API supports native Responses compaction, declare it with
settings = { chat_api = "responses", supports_compaction = true }. This overrides the catalog flag
for that provider's language models, including configured routes whose upstream catalog is unavailable.
Set supports_compaction = false to report no support. Omit the setting to use detected support.
Compaction still requires an explicit request setting and an upstream that implements it.
For Claude, send compaction: {"type":"summarize"} to /v1/messages for on-demand compaction,
or context_management: {"edits":[{"type":"compact_20260112"}]} for threshold compaction.
llmleaf sends the required beta header and preserves returned compaction blocks and signatures.
For on-demand compaction, replace the summarized transcript with the returned block, keep the
same system prompt and tools, then append the next turn. See the
Claude compaction docs.
For OpenAI Responses, send context_management: [{"type":"compaction","compact_threshold":200000}]
to /v1/responses. Replay the returned encrypted compaction items with their IDs, or continue
with previous_response_id. llmleaf preserves those items in collected and streaming responses.
This supports inline compaction. The standalone /v1/responses/compact endpoint is not exposed.
See the OpenAI compaction docs.
The chat-completions surface accepts these native request fields for routes that support them.
It returns compaction blocks in the message.compaction array or streamed delta.compaction
array. Replay those blocks as message.compaction on the next request.
Send model, state, and named questions to /v1/decisions. The OpenRouter path
/api/alpha/decisions and JEV path accept the same body. Questions can
use , , or , including structured instructions and criteria.
Responses preserve answers, probabilities, confidence, and provider-reported usage.
The jev route in llmleaf.example.toml uses OpenRouter. To call TypeSafe directly,
add a provider with kind = "typesafe", credential = "env:TYPESAFE_API_KEY", and
route to it with upstream model . The provider kind is an alias
for . Decisions use the same key permissions and fallback rules as chat.
Set to a consumer bearer token allowed to use the route.
curl localhost:8080/v1/decisions \
-H "Authorization: Bearer $LLMLEAF_API_KEY" \
-H 'Content-Type: application/json' \
-d '{"model":"jev","state":"The export button does nothing.","questions":{"bug":{"type":"noul","instructions":"Does this report broken software?"}}}'
Use GET /v1/models?type=decisions to list models identified as producing decisions.
Their architecture.output_modalities contains "decisions", and their
architecture.modality is "text->decisions". Configured aliases use metadata from
their primary upstream target. Models without known metadata are excluded from
this filter, including pinned TypeSafe versions absent from its catalog.
API references: OpenRouter decisions and TypeSafe JEV.
For an example with all three question types, run
cargo run -p llmleaf --example decisions -- --state "The export button does nothing."
against a server with the jev route configured.
Two strictly separated planes. The core (data plane) is a Compio/Cyper server: ingress,
control-plane background I/O, and provider HTTP run on Compio. The separate llmleaf-web
control-plane app intentionally remains Tokio/Axum; its narrow compatibility bridges never put a
Tokio reactor or I/O loop in the data plane. The control plane is reached only outbound — the core
pulls identity/verdicts/topology and pushes usage, never the reverse. A pulled topology
([control.topology]) lets the controller also serve provider and route configuration, diffed
against the previous pull on every refresh so resources are added, updated, and removed
incrementally on top of the immutable config file. See SOUL.md for the full design
constitution. To build a compatible controller, see the
.
flowchart LR
Cons["Consumers<br/>OpenAI · OpenRouter · Anthropic"] --> Surf["Compat surfaces"]
subgraph Core["llmleaf core — data plane"]
direction LR
Surf --> Auth["authenticate"] --> In["map in"] --> Route["route + fallback"] --> Stream["stream"] --> Out["map out"] --> Ev["emit events"]
end
Route --> Prov["Providers<br/>compiled-in traits · WASM plugins"]
Prov --> Up["LLM providers"]
Ctrl[["Control plane (outbound)"]]
Auth -. "pull identity / verdicts" .-> Ctrl
Route -. "pull topology (providers + routes)" .-> Ctrl
Ev -. "push usage" .-> Ctrl
This project is being developed with AI assistance.
Copyright (C) 2026 Fionn Langhans fionnlanghans@codefionn.eu.
llmleaf and its clients are dual-licensed under either the MIT License or the Apache License 2.0, at your option.
Surfaced from shared tags and platforms — no rankings paid for.
typesafe, alias jev) and OpenRouter.zai-coding (GLM Coding Plan, /api/coding/paas/v4) and kimi-coding (Kimi for Coding,
api.kimi.com/coding/v1); MiniMax's Token Plan shares the standard endpoint, so
minimax-token-plan is an alias of minimax (only the key differs).echo for local testing.401 unknown api keypw_hash| Endpoint | Purpose |
|---|
POST /v1/chat/completions | Chat (SSE streaming) |
POST /v1/messages | Anthropic Messages dialect |
POST /v1/responses | OpenAI Responses dialect (encrypted stateless replay and proxied store/previous_response_id; GET remains a 404-by-design stub) |
POST /v1/embeddings | Embeddings |
POST /v1/rerank | Rerank (Cohere/Jina/OpenRouter dialect) |
POST /v1/decisions, POST /api/alpha/decisions, POST /v1/systemone | Decisions (OpenRouter and TypeSafe JEV dialect) |
POST /v1/audio/speech, GET /v1/audio/voices | Text-to-speech |
POST /v1/audio/transcriptions | Speech-to-text |
GET /v1/realtime | OpenAI Realtime (WebSocket) |
POST /v1/batches, GET /v1/batches/{id}[/results] | Batch jobs (ids HMAC-signed + owner-bound with [server].batch_id_secret) |
GET /v1/models, GET /v1/openapi.json, GET /healthz | Discovery & health |
/v1/systemonenoulchoicescorejev-latestjevtypesafeLLMLEAF_API_KEYjev