Route to your own hardware — or approved cloud when you choose. Keep sensitive routes off unapproved providers; redact known credentials before dispatch. Refuse anything over budget. Kill loops in seconds. Connect native Codex Responses with gateway-owned auth and keep Ollama on its separate Chat Completions path, without hiding the provider boundary. One ~5.5 MB Rust binary that decides before the money is spent, with an embedded control room at /ui.
Works with any OpenAI-compatible agent — point base_url at localhost:8787/v1.
Claude Code connects via ANTHROPIC_BASE_URL. Latest stable · MIT · dogfooded daily.
Local nodes can be stranded, cold, or overloaded while a runaway agent turns retries and fan-out into spend. Stoke puts the circuit breaker before the provider call: it rejects loops, clamps fan-out, holds concurrent stream reservations, and only escalates to paid inference when your policy allows it.
Loop detection and rate limits refuse repeated traffic before another provider call is made.
Pre-dispatch ceilings limit how many billed provider calls one request may create.
Concurrent streams reserve their possible cost before dispatch, so a burst cannot overshoot the key's cap.
Use model support, warm state, health, load, and policy to choose a route across your configured capacity.
An embedded panel for your running Stoke gateway: decisions, budget state, routing, and node status in one view.
Illustrative preview. This button simulates panel events in your browser. It does not send traffic to a gateway.
The repository checks exercise the pricing gate, PII redaction, loop refusal, streamed spend, in-flight holds, fan-out clamp, cache isolation, native Responses passthrough, gateway auth, and fail-closed boot validation against local mock providers. These are wire-behavior checks, not savings benchmarks.
Native Codex can use the existing ~/.codex/auth.json through Stoke's Responses provider. STOKE_API_KEYS still authenticates the gateway hop; separate client OAuth remains supported.
Ollama model IDs are discovered from your running Ollama and served on Chat Completions. Native Codex Responses stays native — Stoke does not translate the two client paths.
Named routes can opt into exact caching and bounded coalescing for eligible non-streaming requests. Only validated, unambiguous Retry-After and x-should-retry values may reach clients; this adds no internal retry policy.
Default-off Headroom accepts only a smaller, validated, lossless JSON-whitespace transformation before native Responses dispatch. Unsupported, unavailable, or rejected work keeps the original outputs; the local worker has documented size and platform limits.
Point OpenAI-compatible clients, native Codex, Hermes, or Claude Code at Stoke. Native Codex can use the gateway-owned login from ~/.codex/auth.json, while STOKE_API_KEYS authenticates the client hop. Responses stays Responses — including reasoning and SSE — while Ollama stays on its separate Chat Completions path. Connect Codex and Ollama →
Stoke discovers exact Ollama model IDs, inventory, and health through the running provider and protected /v1/nodes, then routes across configured Ollama nodes and gateways without manual topology selection. Stoke ships no model catalogue.
Poll each configured Ollama or federated Stoke node for discovered models, warm state, health, and load.
Stoke filters by route policy, then prefers a healthy warm node with the lowest live load—without asking the client to map topology.
The response and logs identify the selected node and model; GET /v1/nodes remains the operator's live inventory view.
Keep Ollama private on each machine. Add the second Stoke gateway as an authenticated remote provider; both machines then sit behind the same client URL.
OPENAI_BASE_URL=http://localhost:8787/v1 stays unchanged for the client; Stoke discovers the remote inventory and chooses the eligible node.
See the two-machine Ollama walkthrough.
Use native Codex auth and discovered Ollama IDs without copying OAuth into client config.
Task-shaped walkthroughs for spend, loops, routing, and security.
Configuration, endpoints, and the operational contract.
Open source, MIT licensed, and local by default. No hosted control plane or account required.
Read the docs →Optional, default-off Headroom accepts only a smaller, validated, lossless JSON-whitespace transformation before native Responses dispatch; it does not summarize, change the model, or rewrite reasoning. If the worker is unavailable or rejects a transformation, Stoke retains the original outputs. Buffered and streaming requests are supported before dispatch, but upstream SSE is not compressed. The current worker lock targets macOS arm64 with CPython 3.11, with bounded 1 MiB requests and 128 outputs per batch. Setup and limits.
Only eligible non-streaming named routes can use the exact cache policy. The identity includes caller scope, resolved model, and the complete effective request; exact coalescing is opt-in and bounded, and followers wait at most five seconds. Route TTL can shorten the global retention but not extend it. This is not a Responses caching claim, and reuse is not secure erasure. Cache policy details.
No. Stoke propagates only valid, unambiguous upstream Retry-After and x-should-retry hints when eligible. That is client-facing metadata, not a new internal retry or failover policy.
No. Stoke is a self-hosted Rust gateway that runs on infrastructure you control.
Yes. Stoke accepts OpenAI-compatible Chat Completions, native Codex Responses, and the Anthropic Messages API. Point the client at Stoke, then configure the upstream provider or discovered local node. Follow the Codex + Ollama guide.
Yes. Connect and configure the nodes or gateways, then Stoke can route across their discovered model and health state.
No. It is a browser-only preview of the panel's event shape. A running gateway exposes the real authenticated panel at GET /ui and the repeat-request demo at POST /ui/demo.
Start with the guides, then read the README and architecture notes.
A route can be pinned to the tiers you approve. When no approved provider is available, Stoke fails closed with 403 rather than silently escalating to unapproved cloud capacity.
Set allowed_tiers = ["local", "remote"] on a route and cloud fallback is denied before any provider is contacted.
Unknown tier names are rejected at boot; a route with no eligible approved provider returns 403, never an open gateway.
Known credentials are redacted before a request is dispatched to a provider.
Align configured machines behind one policy gateway, then decide what reaches a paid provider.