Admin API Reference¶
SMG provides administrative endpoints for managing tokenizers, workers, WASM modules, and engine caches, plus read-only model and server information.
Tokenizer Management¶
Manage tokenizers for text processing and tokenization.
Add Tokenizer¶
POST /v1/tokenizersAdds a new tokenizer from a local path or HuggingFace model ID.
Request Body:
{
"name": "llama3-tokenizer",
"source": "meta-llama/Meta-Llama-3-8B",
"chat_template_path": "/path/to/template.jinja"
}| Field | Type | Required | Description |
|---|---|---|---|
name |
string | Yes | Unique tokenizer identifier |
source |
string | Yes | HuggingFace model ID or local path |
chat_template_path |
string | No | Path to custom Jinja2 chat template |
Response: 202 Accepted
{
"id": "550e8400-e29b-41d4-a716-446655440000",
"status": "pending",
"message": "Tokenizer 'llama3-tokenizer' registration job submitted. Loading from: meta-llama/Meta-Llama-3-8B"
}Loading runs in the background. Poll Get Tokenizer Status with the returned id.
Response: 409 Conflict when a tokenizer with this name is already registered. The id is the existing tokenizer's:
{
"id": "550e8400-e29b-41d4-a716-446655440000",
"status": "failed",
"message": "Tokenizer 'llama3-tokenizer' already exists"
}List Tokenizers¶
GET /v1/tokenizersReturns all registered tokenizers.
Response: 200 OK
{
"tokenizers": [
{
"id": "550e8400-e29b-41d4-a716-446655440000",
"name": "llama3-tokenizer",
"source": "meta-llama/Meta-Llama-3-8B",
"vocab_size": 128256
}
]
}Get Tokenizer¶
GET /v1/tokenizers/{tokenizer_id}Returns details for a specific tokenizer. {tokenizer_id} is the tokenizer's ID or its name.
Response: 200 OK
{
"id": "550e8400-e29b-41d4-a716-446655440000",
"name": "llama3-tokenizer",
"source": "meta-llama/Meta-Llama-3-8B",
"vocab_size": 128256
}Response: 404 Not Found
{
"error": {
"message": "Tokenizer 'llama3-tokenizer' not found",
"type": "tokenizer_not_found"
}
}Get Tokenizer Status¶
GET /v1/tokenizers/{tokenizer_id}/statusReturns the loading status of a tokenizer. A loaded tokenizer can be looked up by ID or name; a job that is still queued, loading, or failed only by the ID that Add Tokenizer returned.
Response: 200 OK
{
"id": "550e8400-e29b-41d4-a716-446655440000",
"status": "completed",
"message": "Tokenizer 'llama3-tokenizer' is loaded and ready",
"vocab_size": 128256
}| Status | Description |
|---|---|
pending |
Tokenizer loading queued |
processing |
Tokenizer currently loading |
completed |
Tokenizer ready for use |
failed |
Loading failed (see message) |
vocab_size is present only for completed. A failed status is kept for about five minutes. After that, and for an ID that SMG does not know:
Response: 404 Not Found
{
"error": {
"message": "Tokenizer '550e8400-e29b-41d4-a716-446655440000' not found and no pending job",
"type": "not_found"
}
}Remove Tokenizer¶
DELETE /v1/tokenizers/{tokenizer_id}Removes a tokenizer. {tokenizer_id} is the tokenizer's ID or its name.
Response: 200 OK
{
"success": true,
"message": "Tokenizer 'llama3-tokenizer' removed successfully"
}Response: 404 Not Found
{
"success": false,
"message": "Tokenizer 'llama3-tokenizer' not found"
}Worker Management¶
Register, inspect, update, and remove backend workers at runtime.
| Method | Path | Purpose |
|---|---|---|
POST |
/workers |
Register a worker |
GET |
/workers |
List workers |
GET |
/workers/{worker_id} |
Get one worker and its job status |
PATCH |
/workers/{worker_id} |
Change priority, cost, labels, API key, or health settings |
PUT |
/workers/{worker_id} |
Replace the spec and re-run registration |
DELETE |
/workers/{worker_id} |
Drain and remove a worker |
worker_id is the UUID that POST /workers returns. Changes are asynchronous: the gateway checks the request, queues a job, and answers 202 Accepted. Follow the job with GET /workers/{worker_id}.
Create Worker¶
POST /workersRegisters a worker. The body is a worker spec; only url is required. The URL scheme selects the transport:
| Scheme | Transport | Backends |
|---|---|---|
http://, https:// |
HTTP | SGLang, vLLM, or other OpenAI-compatible HTTP servers, and external providers. Hosts ending in openai.com, anthropic.com, x.ai, or googleapis.com register as external automatically. |
grpc://, grpcs:// |
gRPC | SGLang, vLLM, TensorRT-LLM, TokenSpeed, or MLX gRPC servers. See gRPC Workers. |
ipc://<path> |
ZMQ | A vLLM or TokenSpeed engine core on the same host. See ZMQ Workers. |
The scheme must be lowercase. Send the spec as JSON:
curl -X POST http://localhost:30000/workers \
-H "Content-Type: application/json" \
-H "Authorization: Bearer ${ADMIN_TOKEN}" \
-d @worker.json{
"url": "http://gpu1:8000",
"api_key": "worker-secret-key",
"labels": {"region": "us-east"}
}The gateway probes the worker, detects the engine, and reads the served model from it.
{
"url": "grpc://gpu2:50051",
"runtime_type": "sglang",
"models": [
{
"id": "meta-llama/Llama-3.1-8B-Instruct",
"aliases": ["llama-3.1-8b"],
"tool_parser": "llama"
}
]
}An explicit runtime_type skips engine detection. The model card adds an alias and pins the tool-call parser for this model.
{
"url": "ipc:///tmp/smg-zmq/engine-31000",
"runtime_type": "vllm",
"labels": {"model_path": "Qwen/Qwen3-8B"}
}The gateway binds the sockets and the engine dials in. For this path the handshake listens on tcp://127.0.0.1:22714, a port derived from the ipc:// path. An engine core does not report a model name, so a ZMQ worker needs models, a model label such as model_path (which also locates the tokenizer), or a gateway started with --model-path. For a group of data-parallel engines or a fixed handshake address, see the ZMQ fields under Worker Spec.
{
"url": "http://gpu3:8000",
"runtime_type": "vllm",
"overload": {
"waiting_requests": 32,
"token_usage": 0.85
},
"http_pool": {
"http2": true,
"connect_timeout_secs": 5
},
"health": {
"drain_settle_secs": 30
}
}This worker leaves routing while 32 or more requests wait in its queue or its KV cache is at least 85% used, even if gateway-wide overload protection is off. The gateway always speaks HTTP/2 (h2c) to it instead of negotiating the version, so registration fails if the worker only speaks HTTP/1.1. When deleted while ready, it drains for 30 seconds.
Response: 202 Accepted, with a Location header equal to location.
{
"status": "accepted",
"worker_id": "01997a4e-3c2b-7f1d-9a8e-2b6c4d5e7f80",
"url": "http://gpu1:8000",
"location": "/workers/01997a4e-3c2b-7f1d-9a8e-2b6c4d5e7f80",
"message": "Worker addition queued for background processing"
}202 means the request passed the checks listed under Worker Errors. Probing, engine detection, and the remaining validation run in the background job; Get Worker shows how to follow it.
Worker Spec¶
The body of POST /workers and PUT /workers/{worker_id}. Only url is required. The nested blocks (health, overload, http_pool, resilience) are partial: a field you leave out falls back to the gateway default. Unknown fields are ignored rather than rejected, so check the spelling. A known field with the wrong type, or an unknown enum value, is rejected with 422.
Identity and transport
| Field | Type | Default | Description |
|---|---|---|---|
url |
string | required | Worker address. The scheme selects the transport (see Create Worker). Each URL can be registered once. |
runtime_type |
string | unspecified |
sglang, vllm, trtllm, mlx, tokenspeed, generic (an OpenAI-compatible HTTP server whose engine is unknown), or external (a third-party API, required for a provider behind a host the gateway does not recognize). unspecified detects the engine: over HTTP from /v1/models, /version, and /server_info, registering an unidentified OpenAI-compatible server as generic; over gRPC by trying sglang, vllm, trtllm, tokenspeed, then mlx. ZMQ workers default to vllm. Also accepted as runtime. |
worker_type |
string | regular |
regular, prefill, decode, or encode (EPD encode worker). See PD Disaggregation. |
api_key |
string | none | Key the gateway presents to the worker. Write-only: never returned. Workers added through the API do not inherit the gateway --api-key. For provider keys, see External Providers. |
Models and labels
| Field | Type | Default | Description |
|---|---|---|---|
models |
array | [] |
Model cards the worker serves (fields below). A self-hosted worker uses the first card. With none, it takes the model the engine reports: the served_model_name, model_id, or model_path label, in that order. External workers ignore this field: the gateway lists the provider's models when it has a key (api_key, or the provider's *_ADMIN_KEY environment variable), and without one the worker accepts any model. |
labels |
object | {} |
String-to-string metadata, merged over the labels the gateway discovers from the engine (your values win). Some keys have built-in meaning (see below). |
Model card fields (models[]):
| Field | Description |
|---|---|
id |
Model ID (required). |
aliases |
Other model names that route to this model. |
model_type |
Capabilities, any of chat, completions, responses, embeddings, rerank, generate, vision, tools, reasoning, image_gen, audio, moderation. Default: ["chat", "completions", "responses", "tools"]. |
context_length |
Context window in tokens. |
chat_template |
Chat template path used when the gateway loads the tokenizer for a gRPC or ZMQ worker. |
tool_parser, reasoning_parser |
Parser names for this model on the gRPC path, overriding --tool-call-parser and --reasoning-parser. An unknown name fails registration. |
Cards also accept display_name, hf_model_type, architectures, provider, tokenizer_path, metadata, id2label, and num_labels.
Labels with built-in meaning:
| Label | Effect |
|---|---|
realtime |
"true" (exactly) lets the worker serve the Realtime API. |
tool_parser, reasoning_parser |
Same as the model card fields, when the card does not set them. |
served_model_name, model_id, model_path |
Model ID when models is empty, checked in this order. |
tokenizer_path, model_path |
Where the gateway loads the tokenizer for a gRPC or ZMQ worker, ahead of --tokenizer-path and --model-path. |
pairing_protocol |
Explicit PD pairing key: a prefill and a decode worker that both set it pair only when the values match. |
kv_connector, kv_role, kv_engine_id |
Moved into the spec fields of the same name at registration (a non-empty kv_connector or kv_engine_id field wins over the label). |
Routing metadata
| Field | Type | Default | Description |
|---|---|---|---|
priority |
integer | 50 |
Stored and returned in responses. No built-in routing policy reads it in v1.11.0. |
cost |
number | 1.0 |
Stored and returned in responses. No built-in routing policy reads it in v1.11.0. |
Health checks (health): per-worker overrides of the gateway health-check flags. See Health Checks.
| Field | Gateway default | Description |
|---|---|---|
timeout_secs |
--health-check-timeout-secs (5) |
Probe timeout, also used for the registration probes. |
check_interval_secs |
--health-check-interval-secs (60) |
Seconds between probes. |
success_threshold |
--health-success-threshold (2) |
Consecutive successful probes before the worker becomes ready. |
failure_threshold |
--health-failure-threshold (3) |
Consecutive failed probes before the worker leaves rotation. |
disable_health_check |
--disable-health-check (off); true for external workers |
Skip probing: the worker is routable as soon as it registers. Ignored for ZMQ workers. |
drain_settle_secs |
--drain-settle-secs (5) |
Seconds a ready worker stays draining when it is removed. 0 removes it immediately. |
Overload thresholds (overload): self-hosted workers only. See Overload Protection.
| Field | Type | Gateway default | Description |
|---|---|---|---|
waiting_requests |
integer, at least 1 | --worker-overload-waiting-requests |
Waiting requests, summed across DP ranks, at or above which the worker leaves routing. |
token_usage |
number in (0.0, 1.0] | --worker-overload-token-usage (0.9 under --worker-overload-protection) |
Mean KV-cache usage across DP ranks at or above which the worker leaves routing. |
Either field turns overload protection on for this worker, even when the gateway flags leave it off; a field you leave out uses the gateway value. The worker returns to routing once a load report is below both thresholds. An out-of-range value fails registration. PATCH cannot change this block; use PUT.
HTTP client (http_pool): applies to HTTP workers, including external providers. See Request Streaming.
| Field | Type | Default | Description |
|---|---|---|---|
http2 |
boolean | --upstream-http2 (off) |
true: always speak HTTP/2 with prior knowledge (h2c on http://); registration fails if the worker only speaks HTTP/1.1. false: no HTTP/2 prior knowledge, so an http:// worker stays on HTTP/1.1. Unset: under --upstream-http2, the gateway probes an http:// worker with both versions and prefers HTTP/2; without the flag, an http:// worker uses HTTP/1.1. Responses report the result as http2. |
pool_max_idle_per_host |
integer | 500 |
Idle connections kept per host. |
pool_idle_timeout_secs |
integer | --upstream-pool-idle-timeout-secs (3) |
How long an idle connection is kept. Keep it below the engine's keep-alive timeout; 0 keeps idle connections forever. |
timeout_secs |
integer | --request-timeout-secs (1800) |
Default request timeout. |
connect_timeout_secs |
integer | 10 |
Connection timeout. |
Resilience (resilience): per-worker circuit-breaker settings. See Circuit Breakers.
| Field | Type | Gateway default | Description |
|---|---|---|---|
cb_failure_threshold |
integer | --cb-failure-threshold (10) |
Failures that open this worker's circuit. |
cb_success_threshold |
integer | --cb-success-threshold (3) |
Successes in the half-open state that close it. |
cb_timeout_secs |
integer | --cb-timeout-duration-secs (60) |
Seconds before an open circuit tries half-open. |
cb_window_secs |
integer | --cb-window-duration-secs (120) |
Accepted but unused: the circuit breaker counts consecutive failures, not failures in a window. |
retryable_status_codes |
integer array | [408, 429, 500, 502, 503, 504] |
Statuses counted as circuit-breaker failures. Replaces the default set; it does not change which responses are retried. |
capacity_status_codes |
integer array | [429] |
Capacity-pushback statuses, which the circuit breaker never counts, as either failure or success. Replaces the default set. |
ZMQ workers: see ZMQ Workers.
| Field | Type | Default | Description |
|---|---|---|---|
dp_size |
integer | none | On an ipc:// worker, the number of data-parallel engines that dial into its one socket set. A value above 1 makes a grouped worker, and the gateway spreads requests across the group's engines. On other workers the gateway sets this field itself during data-parallel discovery. |
zmq_handshake_address |
string | derived | tcp:// address the gateway binds for the engine handshake. The default is tcp://127.0.0.1:<port>, with the port (20000 to 29999) derived from the ipc:// path. Set it for an engine that dials a fixed address, such as tcp://127.0.0.1:30500, TokenSpeed's default. |
Registration of a ZMQ worker fails when the runtime is not vllm or tokenspeed, when worker_type is not regular (disaggregated workers need gRPC), when no model ID is available, when dp_size is above 1 on a gateway running with --dp-aware, or when the handshake address is not tcp:// or is already bound by another ZMQ worker. Setting zmq_handshake_address on a non-ZMQ worker also fails registration. Health checks stay on for ZMQ workers, because the probe is what reconnects a restarted engine.
PD disaggregation and KV transfer: see PD Disaggregation.
| Field | Type | Default | Description |
|---|---|---|---|
bootstrap_port |
integer | none | KV bootstrap port of a prefill or encode worker. The gRPC and vLLM Mooncake paths use 8998 when it is unset. |
kv_connector |
string | engine-reported | KV connector, such as MooncakeConnector or NixlConnector. Overrides the value a vLLM gRPC engine reports. |
kv_engine_id |
string | engine-reported | vLLM kv_transfer_config.engine_id, used for Mooncake PD. For gRPC prefill and decode workers, the gateway reads it from the engine again when the worker recovers from failed or not_ready (a restarted engine gets a new ID), and the engine's value wins. |
health, http_pool, and resilience take effect but are not echoed back by GET /workers; overload is.
Worker Errors¶
Worker endpoints report errors as {"error": "<message>", "code": "<CODE>"}, not in the shape described under Error Responses:
{
"error": "Invalid value for field 'worker_url': 10.0.0.5:8000 - URL must start with a lowercase http://, https://, grpc://, grpcs://, or ipc:// scheme",
"code": "BAD_REQUEST"
}| Status | code |
When |
|---|---|---|
400 |
BAD_REQUEST |
worker_id is not a UUID. The url is empty, does not start with a lowercase http://, https://, grpc://, grpcs://, or ipc:// scheme, is not a valid URL or has no host, or is an ipc:// URL without a path. A PUT changes the URL or is sent to a gateway running with --dp-aware. |
400 |
PROVIDER_NOT_COMPILED |
The spec targets a provider whose router is not compiled into this build. |
404 |
WORKER_NOT_FOUND |
No worker has this ID. |
409 |
WORKER_ALREADY_EXISTS |
POST for a URL that is already registered. The message names the existing ID. |
409 |
WORKER_CREATE_IN_PROGRESS |
POST for a URL whose registration is still running. The message names the ID to poll. |
500 |
INTERNAL_SERVER_ERROR |
The job queue is unavailable or full. |
A body that does not parse never reaches these checks. The gateway answers with a plain-text body: 415 when Content-Type: application/json is missing, 400 for malformed JSON, and 422 (Failed to deserialize the JSON body into the target type: ...) for a missing url, a wrong type, or an unknown value such as "worker_type": "prefil".
Other problems surface in the background job, after the 202:
overload.waiting_requestsof0, oroverload.token_usageoutside (0.0, 1.0]- the ZMQ constraints listed under Worker Spec
- an unknown
tool_parserorreasoning_parser - a worker that never answers, or
http_pool.http2: truefor a worker that only speaks HTTP/1.1
List Workers¶
GET /workers
GET /workers?model={model_id}Returns every registered worker. model keeps only workers that serve that model ID or alias; workers without a model list match any model. The stats counts cover the returned workers (encode workers are counted in total only). Workers whose registration job is still running are not listed.
Response: 200 OK (abridged: discovered labels and card fields vary by engine)
{
"workers": [
{
"id": "01997a4e-3c2b-7f1d-9a8e-2b6c4d5e7f80",
"model_id": "meta-llama/Llama-3.1-8B-Instruct",
"url": "grpc://gpu2:50051",
"models": [
{
"id": "meta-llama/Llama-3.1-8B-Instruct",
"aliases": ["llama-3.1-8b"],
"model_type": ["chat", "completions", "responses", "tools"],
"context_length": 131072,
"tool_parser": "llama"
}
],
"worker_type": "regular",
"connection_mode": "grpc",
"runtime_type": "sglang",
"labels": {
"model_path": "meta-llama/Llama-3.1-8B-Instruct",
"tp_size": "1"
},
"priority": 50,
"cost": 1.0,
"max_connection_attempts": 20,
"is_healthy": true,
"status": "ready",
"load": 2,
"http2": false
}
],
"total": 1,
"stats": {
"prefill_count": 0,
"decode_count": 0,
"regular_count": 1
}
}Each worker object carries its spec fields (never api_key) plus:
| Field | Description |
|---|---|
id |
Worker UUID. |
model_id |
ID of the first model card. Absent for workers without a model list. |
is_healthy |
true when status is ready. |
status |
pending (registered, not yet proven healthy), ready, not_ready (failing probes, may recover), failed, or draining (being removed). Only ready workers receive traffic. |
load |
Requests the gateway has in flight to this worker. |
http2 |
Whether the gateway speaks HTTP/2 to this worker. |
pd_pairing |
Prefill and decode workers only: the key used to pair them for KV transfer, either the explicit pairing_protocol or runtime/transport/layout. |
engine_load |
The engine's last polled load report (per-DP-rank queue and KV-cache figures). Absent until the load monitor has polled the worker. |
Get Worker¶
GET /workers/{worker_id}Returns one worker in the same shape as a GET /workers entry. While a job for the worker is queued or running, or for about five minutes after one fails, the response also carries job_status. A job that succeeds clears it.
While POST /workers is still registering the worker, the response is a placeholder: only id, url, status, and job_status are meaningful, and the other fields hold defaults.
Response: 200 OK
{
"id": "01997a4e-3c2b-7f1d-9a8e-2b6c4d5e7f80",
"url": "grpc://gpu2:50051",
"worker_type": "regular",
"connection_mode": "http",
"runtime_type": "unspecified",
"priority": 50,
"cost": 1.0,
"max_connection_attempts": 20,
"is_healthy": false,
"status": "pending",
"load": 0,
"http2": false,
"job_status": {
"job_type": "AddWorker",
"worker_url": "grpc://gpu2:50051",
"status": "processing",
"message": null,
"timestamp": 1790250060
}
}job_status field |
Description |
|---|---|
job_type |
AddWorker (create or replace), UpdateWorker, or RemoveWorker. |
worker_url |
URL of the worker the job targets. |
status |
pending (queued), processing, or failed. |
message |
Error text when status is failed; otherwise null. |
timestamp |
Unix time, in seconds, of the last status change. |
Registration keeps probing a worker that does not answer, every --worker-startup-check-interval seconds (default 30), until --worker-startup-timeout-secs (default 1800) runs out. If a POST /workers job fails, the gateway releases the reserved ID: GET /workers/{worker_id} returns 404, and the reason appears only in the gateway log (Failed job: type=AddWorker ...). A failed PATCH, PUT, or DELETE job leaves the worker in place with a failed job_status.
Update Worker (partial)¶
PATCH /workers/{worker_id}Changes a few fields in place. Omitted fields keep their current values. The worker keeps its status and connections, and nothing is probed again.
Request Body:
{
"priority": 75,
"labels": {"tier": "gold"},
"api_key": "new-api-key",
"health": {"check_interval_secs": 15, "drain_settle_secs": 60}
}| Field | Type | Description |
|---|---|---|
priority |
integer | New priority. |
cost |
number | New cost. |
labels |
object | Merged into the current labels: listed keys are added or overwritten, and the rest are kept. PATCH cannot remove a label. |
api_key |
string | New worker API key, for key rotation. |
health |
object | Partial health overrides, merged into the worker's current settings: timeout_secs, check_interval_secs, success_threshold, failure_threshold, disable_health_check, drain_settle_secs. Setting disable_health_check to true makes the worker routable immediately. |
Any other field, such as overload, http_pool, resilience, or models, is ignored. Use PUT to change those.
In v1.11.0, a PATCH also rebuilds the worker's circuit breaker with built-in thresholds (5 failures, 2 successes, 30 seconds), ignoring the gateway's circuit-breaker flags and any resilience overrides until a PUT re-registers the worker.
Response: 202 Accepted
{
"status": "accepted",
"worker_id": "01997a4e-3c2b-7f1d-9a8e-2b6c4d5e7f80",
"message": "Worker update queued for background processing"
}Replace Worker (full)¶
PUT /workers/{worker_id}Replaces the worker's spec and re-runs the full registration workflow (probe, engine detection, metadata discovery) under the same ID. The body is a complete worker spec. Fields you omit take their defaults rather than keeping their current values, so resend any labels, overload, or other settings you want to keep.
The url must equal the worker's current URL; to move a worker, DELETE it and POST the new URL. PUT is refused with 400 when the gateway runs with --dp-aware, because one spec expands to one worker per data-parallel rank; use DELETE and POST there too. If the worker is deleted or replaced again before the job runs, the job fails rather than overwrite the newer state.
Response: 202 Accepted, with the same body as PATCH.
Delete Worker¶
DELETE /workers/{worker_id}Removes a worker. A ready worker first moves to draining: it receives no new requests, and the gateway waits health.drain_settle_secs (default --drain-settle-secs, 5 seconds) so in-flight requests can finish before it removes the worker. A worker in any other status is removed without waiting.
Response: 202 Accepted
{
"status": "accepted",
"worker_id": "01997a4e-3c2b-7f1d-9a8e-2b6c4d5e7f80",
"message": "Worker removal queued for background processing"
}Cache Management¶
Flush the engines' KV caches, and read the engine load the gateway has polled from its workers.
Flush Cache¶
POST /flush_cacheAsks every registered HTTP and gRPC worker to drop its KV prefix cache. The calls go out in parallel, and the response reports the outcome per worker. Requires admin authentication.
| Worker | How it is flushed |
|---|---|
| HTTP | POST {worker_url}/flush_cache, with the worker's API key (if it has one) as a bearer token and a 45-second timeout. Any non-2xx answer counts as a failure, so the engine must serve this route (SGLang's HTTP server does). |
| gRPC SGLang, TokenSpeed | FlushCache RPC. |
| gRPC vLLM | FlushCache RPC, new in v1.11.0. Needs smg-grpc-servicer 0.12.0 or later; older servicers answer UNIMPLEMENTED. Resets vLLM's local prefix cache without preempting running requests or clearing connector caches. The gateway makes a single attempt, so a reset that vLLM refuses while requests are in flight is reported under failed. |
| gRPC TRT-LLM, MLX | Not supported: reported under failed (UNIMPLEMENTED). |
ZMQ (ipc://) |
Skipped, because the transport has no cache-flush RPC. Counted in total_zmq_workers_skipped. |
The fan-out does not filter by runtime: external-provider workers registered over HTTP receive the POST too, and appear under failed if they reject it. The endpoint only calls the workers; it does not reset the gateway's own cache-aware routing state.
Response: 200 OK when every worker that was called succeeds (or there was nothing to call), 206 Partial Content when at least one fails, including when all of them fail.
{
"status": "success",
"message": "Successfully flushed cache on all 3 workers",
"workers_flushed": 3,
"total_http_workers": 2,
"total_grpc_workers": 1,
"total_zmq_workers_skipped": 0,
"total_workers": 3
}On a partial failure, status is "partial_success" and two more fields list the outcome, successful (worker URLs) and failed ({worker, error} entries):
{
"status": "partial_success",
"message": "Cache flush: 2 succeeded, 1 failed (1 ZMQ workers skipped: no cache-flush RPC)",
"workers_flushed": 2,
"total_http_workers": 2,
"total_grpc_workers": 1,
"total_zmq_workers_skipped": 1,
"total_workers": 4,
"successful": ["http://gpu1:8000", "grpc://gpu3:50051"],
"failed": [
{
"worker": "http://gpu2:8000",
"error": "flush_cache failed for worker http://gpu2:8000/flush_cache: HTTP 500 Internal Server Error"
}
]
}| Field | Description |
|---|---|
status |
success or partial_success |
message |
Human-readable summary |
workers_flushed |
Workers that confirmed the flush |
total_http_workers, total_grpc_workers |
Registered workers by transport |
total_zmq_workers_skipped |
ZMQ workers that were not called |
total_workers |
All registered workers (HTTP + gRPC + ZMQ) |
successful, failed |
Present only when at least one worker failed |
Get Loads¶
GET /loads
GET /get_loadsReturns the engine load the gateway's load monitor last collected: one entry per worker per DP rank, plus a fleet-wide aggregate. The body uses the same schema engines report on /v1/loads, with every entry tagged by worker and worker_type.
The response is built from the monitor's cached snapshot, the same numbers the routing policies act on, so no request reaches a worker. Values can be up to one poll interval old (--load-monitor-interval, 10 seconds by default). Serving from the cache is also what makes the route safe to poll from another SMG gateway that registers this one as a worker.
- Auth:
/loadsis a public route, served without auth like/health./get_loadsis a deprecated alias that stays on the admin tier and requires admin authentication. - Filter:
?model=<model_id>limits the response to workers serving that model. - Missing workers: workers without a current report (not
Readyyet, no load source, or a failed last poll) are left out rather than reported as idle. Before the first poll,loadsis empty andaggregateis absent.
Response: 200 OK
{
"timestamp": "2026-09-24T18:20:11.482913+00:00",
"version": "smg-1.11.0",
"dp_rank_count": 2,
"loads": [
{
"worker": "http://gpu1:8000",
"worker_type": "regular",
"dp_rank": 0,
"num_running_reqs": 12,
"num_waiting_reqs": 3,
"num_waiting_uncached_tokens": 5120,
"num_total_reqs": 15,
"num_used_tokens": 48000,
"max_total_num_tokens": 131072,
"token_usage": 0.37,
"gen_throughput": 1850.5,
"cache_hit_rate": 0.62,
"utilization": 0.41,
"max_running_requests": 256
},
{
"worker": "http://gpu2:8000",
"worker_type": "regular",
"dp_rank": 0,
"num_running_reqs": 9,
"num_waiting_reqs": 0,
"num_waiting_uncached_tokens": 0,
"num_total_reqs": 9,
"num_used_tokens": 30000,
"max_total_num_tokens": 131072,
"token_usage": 0.23,
"gen_throughput": 1420.0,
"cache_hit_rate": 0.55,
"utilization": 0.3,
"max_running_requests": 256
}
],
"aggregate": {
"total_running_reqs": 21,
"total_waiting_reqs": 3,
"total_reqs": 24,
"avg_token_usage": 0.3,
"avg_throughput": 1635.25,
"avg_utilization": 0.355
}
}| Field | Description |
|---|---|
timestamp |
When the gateway built this response (RFC 3339) |
version |
smg- followed by the gateway version |
dp_rank_count |
Number of entries in loads, across the whole fleet |
loads[].worker, loads[].worker_type |
Worker URL and type (regular, prefill, decode, or encode) |
loads[].dp_rank |
DP rank within that worker |
loads[].num_running_reqs, num_waiting_reqs, num_total_reqs |
Requests running, queued, and both |
loads[].num_waiting_uncached_tokens |
Queued tokens not served from the prefix cache (0 when the engine does not report it) |
loads[].num_used_tokens, max_total_num_tokens |
KV tokens in use and KV capacity (0 for workers read from Prometheus /metrics) |
loads[].token_usage |
KV token usage ratio, 0.0 to 1.0 |
loads[].gen_throughput, cache_hit_rate, utilization, max_running_requests |
As reported by the engine |
loads[].memory, loads[].queues |
Optional sections, present when the engine reports them: memory (weight_gb, kv_cache_gb, graph_gb, token_capacity) and queues (waiting, grammar, paused, retracted) |
loads[].kv_transfer_latency_ms, kv_transfer_speed_gb_s, prefill_queue_reqs, decode_queue_reqs, disagg_mode |
Optional prefill/decode disaggregation fields, present when the engine reports them |
aggregate |
Fleet roll-up across all entries: total_running_reqs, total_waiting_reqs, and total_reqs are sums; avg_token_usage, avg_throughput, and avg_utilization are means |
The same per-worker report appears as engine_load on GET /workers and GET /workers/{worker_id}. It is unrelated to their load field, which counts requests the gateway has in flight to the worker. Overload Protection explains where each backend's report comes from.
Model Information¶
Query model and server information. These are public routes and need no authentication.
List Models¶
GET /v1/modelsReturns the models that the registered self-hosted workers serve, read from the gateway's worker registry rather than from the workers. A caller that sends its own provider key gets the external providers' model list instead, fetched with that key. See List Models for the details.
Response: 200 OK
{
"object": "list",
"data": [
{
"id": "meta-llama/Llama-3.1-8B-Instruct",
"object": "model",
"created": 0,
"owned_by": "self_hosted"
}
]
}created is always 0. With no self-hosted model to list, the route answers 503 with the plain-text body No models available.
Get Model Info¶
GET /get_model_infoForwards the request to the first healthy regular HTTP worker (the first healthy prefill worker in HTTP PD mode) and returns that worker's answer. SGLang serves this route, and its answer includes fields such as model_path, tokenizer_path, and is_generation. An engine without the route, such as vLLM, answers 404. In regular mode the worker's status and body pass through unchanged; in HTTP PD mode an error status comes back in the gateway's error envelope. With no healthy worker to ask, the route answers 503.
Get Server Info¶
GET /get_server_infoForwards the request the same way as Get Model Info and returns the worker's answer. SGLang returns its server arguments together with scheduler state and its version.
WASM Module Management¶
Manage WebAssembly plugins. Modules are registered from files accessible to the gateway process; the request body contains descriptors with paths, not binary payloads.
These routes need --enable-wasm. Without it, each answers 500 Internal Server Error with an empty body.
Add WASM Module¶
POST /wasmRegisters one or more WASM modules.
Request Body: JSON WasmModuleAddRequest
{
"modules": [
{
"name": "custom-middleware",
"file_path": "/etc/smg/wasm/custom-middleware.wasm",
"module_type": "Middleware",
"attach_points": [
{"Middleware": "OnRequest"},
{"Middleware": "OnResponse"}
]
}
]
}The only supported module_type today is Middleware. Valid Middleware attach points are OnRequest, OnResponse, and OnError, but OnError modules never run in v1.11.0: the middleware calls only OnRequest and OnResponse modules, and skips OnResponse for streaming responses.
Response: 200 OK on full success, 400 Bad Request if any module failed to register. The response body echoes every requested module with an add_result field indicating success (carrying the assigned UUID) or failure (carrying the error message).
{
"modules": [
{
"name": "custom-middleware",
"file_path": "/etc/smg/wasm/custom-middleware.wasm",
"module_type": "Middleware",
"attach_points": [
{"Middleware": "OnRequest"},
{"Middleware": "OnResponse"}
],
"add_result": {
"Success": "550e8400-e29b-41d4-a716-446655440000"
}
}
]
}List WASM Modules¶
GET /wasmReturns all registered WASM modules together with aggregate execution metrics.
Response: 200 OK
{
"modules": [
{
"module_uuid": "550e8400-e29b-41d4-a716-446655440000",
"module_meta": {
"name": "custom-middleware",
"file_path": "/etc/smg/wasm/custom-middleware.wasm",
"sha256_hash": "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855",
"size_bytes": 65536,
"created_at": "2024-01-15T12:00:00.000000000Z",
"last_accessed_at": "2024-01-15T12:05:00.000000000Z",
"access_count": 40,
"attach_points": [
{"Middleware": "OnRequest"}
]
}
}
],
"metrics": {
"total_executions": 40,
"successful_executions": 40,
"failed_executions": 0,
"total_execution_time_ms": 120,
"max_execution_time_ms": 8,
"average_execution_time_ms": 3.0
}
}metrics covers every module. average_execution_time_ms is total_execution_time_ms divided by total_executions, and is left out until a module has run.
Remove WASM Module¶
DELETE /wasm/{module_uuid}Removes a WASM module. The body is a plain text status message, not JSON.
Response: 200 OK
Module removed successfullyOn failure returns 400 Bad Request with the error text as the body. A module_uuid that is not a UUID gets 400 with an empty body.
Error Responses¶
The admin endpoints do not share one error format. The body depends on the part of SMG that rejects the request:
| Source | Status | Body |
|---|---|---|
| Control-plane auth | 401, 403 |
Plain text, such as Missing or invalid Authorization header, Invalid authentication token, or Admin role required for control plane access. A 401 carries WWW-Authenticate: Bearer realm="control-plane" |
--api-key auth, when control-plane auth is not configured |
401 |
Empty |
| A request body that does not parse | 400, 415, 422 |
Plain text from the JSON extractor |
| Worker endpoints | 400, 404, 409, 500 |
{"error": "<message>", "code": "<CODE>"} |
| Get Tokenizer, Get Tokenizer Status | 404 |
{"error": {"message": "<message>", "type": "tokenizer_not_found"}}, with "type": "not_found" from Get Tokenizer Status |
| Add Tokenizer | 409, 503 |
The 202 body shape with "status": "failed" |
| Remove Tokenizer | 404 |
{"success": false, "message": "<message>"} |
/parse/function_call, /parse/reasoning |
400, 503 |
{"error": "<message>", "success": false} |
| Add WASM Module | 400 |
The module list, with "add_result": {"Error": "<message>"} on each module that failed |
Remove WASM Module, and every WASM route without --enable-wasm |
400, 500 |
Plain text or empty |
The gateway, on forwarded routes such as /get_model_info |
503 and others |
The gateway error envelope below |
Errors that the gateway generates itself, for example when no worker is available, use this envelope and repeat code in the X-SMG-Error-Code response header:
{
"error": {
"type": "Service Unavailable",
"code": "no_workers",
"message": "No workers are available",
"param": null
}
}type is the reason phrase of the HTTP status (Not Found, Service Unavailable, ...), and param is always null. The inference endpoints use the same envelope; see Error Responses in the OpenAI-compatible API reference.
Authentication¶
The tokenizer, worker, WASM, cache flush, and deprecated /get_loads endpoints on this page are control-plane routes. Send the credential as Authorization: Bearer <token>:
- With control-plane auth configured (
--control-plane-api-keys, or--jwt-issuerwith--jwt-audience): a control-plane API key or JWT whose role isadmin. A missing or invalid credential gets401, and one with another role gets403. The shared--api-keyis not accepted. - Without it: the shared
--api-key. Per-tenant--tenant-api-keykeys are rejected, so a gateway with tenant keys but no--api-keyanswers401to every call. With no keys configured at all, the routes are open.
Public endpoints need no authentication: the health probes, /loads, /engine_metrics, and the Model Information routes. See Auth Model for every route's tier.