Gateway¶
The gateway is the single HTTP front door. Clients talk OpenAI-style (and a few extra endpoints); the gateway looks up replicas in the registry and forwards with load balancing + retries.
Start it¶
Multi-worker (production):
Or via uvicorn directly:
REGISTRY_PATH=redis://login-node:6379 \
uvicorn literegistry.gateway:app --host 0.0.0.0 --port 8080 --workers 4
CLI arguments¶
| Argument | Default | Meaning |
|---|---|---|
registry |
cluster Redis URL | redis://… or filesystem path |
port |
8080 |
Listen port |
workers |
1 |
Uvicorn workers (>1 switches to multi-process mode) |
timeout |
61 |
Seconds for model completion / classify requests |
python_timeout |
20 |
Timeout for /python proxied calls |
python_max_retries |
3 |
Max replica attempts for /python |
python_retry_budget_seconds |
20 |
Wall-clock budget for /python retries |
terminal_timeout |
20 |
Timeout for /terminal |
terminal_max_retries |
2 |
Max replica attempts for /terminal |
terminal_retry_budget_seconds |
20 |
Wall-clock budget for /terminal retries |
When workers > 1, the same values are exported as env vars for child
processes: REGISTRY_PATH, TIMEOUT, PYTHON_TIMEOUT, etc.
Endpoints¶
| Method | Path | Routes to |
|---|---|---|
GET |
/health |
Registry force-refresh; returns model count |
GET |
/session-stats |
Shared aiohttp session / connector stats |
GET |
/v1/models |
Distinct model_path values (+ metadata) |
POST |
/v1/completions |
Replica with matching model |
POST |
/v1/chat/completions |
Replica with matching model |
POST |
/classify |
Replica with matching model |
POST |
/python |
Workers registered as model_path=python |
POST |
/terminal |
Workers registered as model_path=terminal |
Completions¶
curl -X POST http://localhost:8080/v1/completions \
-H "Content-Type: application/json" \
-d '{
"model": "meta-llama/Llama-3.1-8B-Instruct",
"prompt": "Hello",
"max_tokens": 64
}'
- Required body field:
model— must match a registeredmodel_path. - All other fields are forwarded to the backend as-is.
Chat completions¶
curl -X POST http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "meta-llama/Llama-3.1-8B-Instruct",
"messages": [{"role": "user", "content": "Hello"}],
"max_tokens": 64
}'
The gateway requires model, forwards messages and all other fields unchanged, and sends the request to POST /v1/chat/completions on the selected replica.
Classify¶
Same routing: body must include model. Forwarded to POST /classify on the
chosen replica.
Python¶
curl -X POST http://localhost:8080/python \
-H "Content-Type: application/json" \
-d '{"code": "print(2 + 2)", "max_runtime": 1.0}'
- Required:
code - Gateway always looks up
model_path="python"(nomodelfield needed). - Uses the shorter python retry/timeout knobs above.
Terminal¶
curl -X POST http://localhost:8080/terminal \
-H "Content-Type: application/json" \
-d '{
"contents": "INFO ok\nERROR disk full\n",
"command": "rg ERROR | head -n 1",
"max_runtime": 5
}'
Routes to model_path="terminal". See Code & Terminal.
Health / session stats¶
Healthy response includes models_count. Session stats should show
shared_session_initialized: true (LiteLLM-style single shared aiohttp session).
How routing works (short)¶
- Parse JSON body; read
model(or hardcodepython/terminal). - Build
RegistryHTTPClient(registry, model, …). - Call
request_with_rotation(endpoint, payload). - Client samples replicas via the Exp3 bandit, tries until success / retries / budget exhausted, and reports latency back for the next request.
Details: Load balancing.
Ops tips¶
- Raise
ulimit -n(e.g.65536) before busy gateways. - Prefer one shared gateway process family with
--workersrather than many independent gateways fighting for the same FDs. - Watch logs for
Request counts (last 5.0s): …andProbs: …— those are the console’s main signal sources. - Failures return HTTP 500 with
{"error": "...", "status": "failed"}; missingmodel/codereturns 400.
Next: vLLM & SGLang · Load balancing