Admin API Reference

SMG provides administrative endpoints for managing tokenizers, workers, WASM modules, and engine caches, plus read-only model and server information.


Tokenizer Management

Manage tokenizers for text processing and tokenization.

Add Tokenizer

POST /v1/tokenizers

Adds a new tokenizer from a local path or HuggingFace model ID.

Request Body:

{
  "name": "llama3-tokenizer",
  "source": "meta-llama/Meta-Llama-3-8B",
  "chat_template_path": "/path/to/template.jinja"
}
Field Type Required Description
name string Yes Unique tokenizer identifier
source string Yes HuggingFace model ID or local path
chat_template_path string No Path to custom Jinja2 chat template

Response: 202 Accepted

{
  "id": "550e8400-e29b-41d4-a716-446655440000",
  "status": "pending",
  "message": "Tokenizer 'llama3-tokenizer' registration job submitted. Loading from: meta-llama/Meta-Llama-3-8B"
}

Loading runs in the background. Poll Get Tokenizer Status with the returned id.

Response: 409 Conflict when a tokenizer with this name is already registered. The id is the existing tokenizer's:

{
  "id": "550e8400-e29b-41d4-a716-446655440000",
  "status": "failed",
  "message": "Tokenizer 'llama3-tokenizer' already exists"
}

List Tokenizers

GET /v1/tokenizers

Returns all registered tokenizers.

Response: 200 OK

{
  "tokenizers": [
    {
      "id": "550e8400-e29b-41d4-a716-446655440000",
      "name": "llama3-tokenizer",
      "source": "meta-llama/Meta-Llama-3-8B",
      "vocab_size": 128256
    }
  ]
}

Get Tokenizer

GET /v1/tokenizers/{tokenizer_id}

Returns details for a specific tokenizer. {tokenizer_id} is the tokenizer's ID or its name.

Response: 200 OK

{
  "id": "550e8400-e29b-41d4-a716-446655440000",
  "name": "llama3-tokenizer",
  "source": "meta-llama/Meta-Llama-3-8B",
  "vocab_size": 128256
}

Response: 404 Not Found

{
  "error": {
    "message": "Tokenizer 'llama3-tokenizer' not found",
    "type": "tokenizer_not_found"
  }
}

Get Tokenizer Status

GET /v1/tokenizers/{tokenizer_id}/status

Returns the loading status of a tokenizer. A loaded tokenizer can be looked up by ID or name; a job that is still queued, loading, or failed only by the ID that Add Tokenizer returned.

Response: 200 OK

{
  "id": "550e8400-e29b-41d4-a716-446655440000",
  "status": "completed",
  "message": "Tokenizer 'llama3-tokenizer' is loaded and ready",
  "vocab_size": 128256
}
Status Description
pending Tokenizer loading queued
processing Tokenizer currently loading
completed Tokenizer ready for use
failed Loading failed (see message)

vocab_size is present only for completed. A failed status is kept for about five minutes. After that, and for an ID that SMG does not know:

Response: 404 Not Found

{
  "error": {
    "message": "Tokenizer '550e8400-e29b-41d4-a716-446655440000' not found and no pending job",
    "type": "not_found"
  }
}

Remove Tokenizer

DELETE /v1/tokenizers/{tokenizer_id}

Removes a tokenizer. {tokenizer_id} is the tokenizer's ID or its name.

Response: 200 OK

{
  "success": true,
  "message": "Tokenizer 'llama3-tokenizer' removed successfully"
}

Response: 404 Not Found

{
  "success": false,
  "message": "Tokenizer 'llama3-tokenizer' not found"
}

Worker Management

Register, inspect, update, and remove backend workers at runtime.

Method Path Purpose
POST /workers Register a worker
GET /workers List workers
GET /workers/{worker_id} Get one worker and its job status
PATCH /workers/{worker_id} Change priority, cost, labels, API key, or health settings
PUT /workers/{worker_id} Replace the spec and re-run registration
DELETE /workers/{worker_id} Drain and remove a worker

worker_id is the UUID that POST /workers returns. Changes are asynchronous: the gateway checks the request, queues a job, and answers 202 Accepted. Follow the job with GET /workers/{worker_id}.

Create Worker

POST /workers

Registers a worker. The body is a worker spec; only url is required. The URL scheme selects the transport:

Scheme Transport Backends
http://, https:// HTTP SGLang, vLLM, or other OpenAI-compatible HTTP servers, and external providers. Hosts ending in openai.com, anthropic.com, x.ai, or googleapis.com register as external automatically.
grpc://, grpcs:// gRPC SGLang, vLLM, TensorRT-LLM, TokenSpeed, or MLX gRPC servers. See gRPC Workers.
ipc://<path> ZMQ A vLLM or TokenSpeed engine core on the same host. See ZMQ Workers.

The scheme must be lowercase. Send the spec as JSON:

curl -X POST http://localhost:30000/workers \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer ${ADMIN_TOKEN}" \
  -d @worker.json
{
  "url": "http://gpu1:8000",
  "api_key": "worker-secret-key",
  "labels": {"region": "us-east"}
}

The gateway probes the worker, detects the engine, and reads the served model from it.

{
  "url": "grpc://gpu2:50051",
  "runtime_type": "sglang",
  "models": [
    {
      "id": "meta-llama/Llama-3.1-8B-Instruct",
      "aliases": ["llama-3.1-8b"],
      "tool_parser": "llama"
    }
  ]
}

An explicit runtime_type skips engine detection. The model card adds an alias and pins the tool-call parser for this model.

{
  "url": "ipc:///tmp/smg-zmq/engine-31000",
  "runtime_type": "vllm",
  "labels": {"model_path": "Qwen/Qwen3-8B"}
}

The gateway binds the sockets and the engine dials in. For this path the handshake listens on tcp://127.0.0.1:22714, a port derived from the ipc:// path. An engine core does not report a model name, so a ZMQ worker needs models, a model label such as model_path (which also locates the tokenizer), or a gateway started with --model-path. For a group of data-parallel engines or a fixed handshake address, see the ZMQ fields under Worker Spec.

{
  "url": "http://gpu3:8000",
  "runtime_type": "vllm",
  "overload": {
    "waiting_requests": 32,
    "token_usage": 0.85
  },
  "http_pool": {
    "http2": true,
    "connect_timeout_secs": 5
  },
  "health": {
    "drain_settle_secs": 30
  }
}

This worker leaves routing while 32 or more requests wait in its queue or its KV cache is at least 85% used, even if gateway-wide overload protection is off. The gateway always speaks HTTP/2 (h2c) to it instead of negotiating the version, so registration fails if the worker only speaks HTTP/1.1. When deleted while ready, it drains for 30 seconds.

Response: 202 Accepted, with a Location header equal to location.

{
  "status": "accepted",
  "worker_id": "01997a4e-3c2b-7f1d-9a8e-2b6c4d5e7f80",
  "url": "http://gpu1:8000",
  "location": "/workers/01997a4e-3c2b-7f1d-9a8e-2b6c4d5e7f80",
  "message": "Worker addition queued for background processing"
}

202 means the request passed the checks listed under Worker Errors. Probing, engine detection, and the remaining validation run in the background job; Get Worker shows how to follow it.


Worker Spec

The body of POST /workers and PUT /workers/{worker_id}. Only url is required. The nested blocks (health, overload, http_pool, resilience) are partial: a field you leave out falls back to the gateway default. Unknown fields are ignored rather than rejected, so check the spelling. A known field with the wrong type, or an unknown enum value, is rejected with 422.

Identity and transport

Field Type Default Description
url string required Worker address. The scheme selects the transport (see Create Worker). Each URL can be registered once.
runtime_type string unspecified sglang, vllm, trtllm, mlx, tokenspeed, generic (an OpenAI-compatible HTTP server whose engine is unknown), or external (a third-party API, required for a provider behind a host the gateway does not recognize). unspecified detects the engine: over HTTP from /v1/models, /version, and /server_info, registering an unidentified OpenAI-compatible server as generic; over gRPC by trying sglang, vllm, trtllm, tokenspeed, then mlx. ZMQ workers default to vllm. Also accepted as runtime.
worker_type string regular regular, prefill, decode, or encode (EPD encode worker). See PD Disaggregation.
api_key string none Key the gateway presents to the worker. Write-only: never returned. Workers added through the API do not inherit the gateway --api-key. For provider keys, see External Providers.

Models and labels

Field Type Default Description
models array [] Model cards the worker serves (fields below). A self-hosted worker uses the first card. With none, it takes the model the engine reports: the served_model_name, model_id, or model_path label, in that order. External workers ignore this field: the gateway lists the provider's models when it has a key (api_key, or the provider's *_ADMIN_KEY environment variable), and without one the worker accepts any model.
labels object {} String-to-string metadata, merged over the labels the gateway discovers from the engine (your values win). Some keys have built-in meaning (see below).

Model card fields (models[]):

Field Description
id Model ID (required).
aliases Other model names that route to this model.
model_type Capabilities, any of chat, completions, responses, embeddings, rerank, generate, vision, tools, reasoning, image_gen, audio, moderation. Default: ["chat", "completions", "responses", "tools"].
context_length Context window in tokens.
chat_template Chat template path used when the gateway loads the tokenizer for a gRPC or ZMQ worker.
tool_parser, reasoning_parser Parser names for this model on the gRPC path, overriding --tool-call-parser and --reasoning-parser. An unknown name fails registration.

Cards also accept display_name, hf_model_type, architectures, provider, tokenizer_path, metadata, id2label, and num_labels.

Labels with built-in meaning:

Label Effect
realtime "true" (exactly) lets the worker serve the Realtime API.
tool_parser, reasoning_parser Same as the model card fields, when the card does not set them.
served_model_name, model_id, model_path Model ID when models is empty, checked in this order.
tokenizer_path, model_path Where the gateway loads the tokenizer for a gRPC or ZMQ worker, ahead of --tokenizer-path and --model-path.
pairing_protocol Explicit PD pairing key: a prefill and a decode worker that both set it pair only when the values match.
kv_connector, kv_role, kv_engine_id Moved into the spec fields of the same name at registration (a non-empty kv_connector or kv_engine_id field wins over the label).

Routing metadata

Field Type Default Description
priority integer 50 Stored and returned in responses. No built-in routing policy reads it in v1.11.0.
cost number 1.0 Stored and returned in responses. No built-in routing policy reads it in v1.11.0.

Health checks (health): per-worker overrides of the gateway health-check flags. See Health Checks.

Field Gateway default Description
timeout_secs --health-check-timeout-secs (5) Probe timeout, also used for the registration probes.
check_interval_secs --health-check-interval-secs (60) Seconds between probes.
success_threshold --health-success-threshold (2) Consecutive successful probes before the worker becomes ready.
failure_threshold --health-failure-threshold (3) Consecutive failed probes before the worker leaves rotation.
disable_health_check --disable-health-check (off); true for external workers Skip probing: the worker is routable as soon as it registers. Ignored for ZMQ workers.
drain_settle_secs --drain-settle-secs (5) Seconds a ready worker stays draining when it is removed. 0 removes it immediately.

Overload thresholds (overload): self-hosted workers only. See Overload Protection.

Field Type Gateway default Description
waiting_requests integer, at least 1 --worker-overload-waiting-requests Waiting requests, summed across DP ranks, at or above which the worker leaves routing.
token_usage number in (0.0, 1.0] --worker-overload-token-usage (0.9 under --worker-overload-protection) Mean KV-cache usage across DP ranks at or above which the worker leaves routing.

Either field turns overload protection on for this worker, even when the gateway flags leave it off; a field you leave out uses the gateway value. The worker returns to routing once a load report is below both thresholds. An out-of-range value fails registration. PATCH cannot change this block; use PUT.

HTTP client (http_pool): applies to HTTP workers, including external providers. See Request Streaming.

Field Type Default Description
http2 boolean --upstream-http2 (off) true: always speak HTTP/2 with prior knowledge (h2c on http://); registration fails if the worker only speaks HTTP/1.1. false: no HTTP/2 prior knowledge, so an http:// worker stays on HTTP/1.1. Unset: under --upstream-http2, the gateway probes an http:// worker with both versions and prefers HTTP/2; without the flag, an http:// worker uses HTTP/1.1. Responses report the result as http2.
pool_max_idle_per_host integer 500 Idle connections kept per host.
pool_idle_timeout_secs integer --upstream-pool-idle-timeout-secs (3) How long an idle connection is kept. Keep it below the engine's keep-alive timeout; 0 keeps idle connections forever.
timeout_secs integer --request-timeout-secs (1800) Default request timeout.
connect_timeout_secs integer 10 Connection timeout.

Resilience (resilience): per-worker circuit-breaker settings. See Circuit Breakers.

Field Type Gateway default Description
cb_failure_threshold integer --cb-failure-threshold (10) Failures that open this worker's circuit.
cb_success_threshold integer --cb-success-threshold (3) Successes in the half-open state that close it.
cb_timeout_secs integer --cb-timeout-duration-secs (60) Seconds before an open circuit tries half-open.
cb_window_secs integer --cb-window-duration-secs (120) Accepted but unused: the circuit breaker counts consecutive failures, not failures in a window.
retryable_status_codes integer array [408, 429, 500, 502, 503, 504] Statuses counted as circuit-breaker failures. Replaces the default set; it does not change which responses are retried.
capacity_status_codes integer array [429] Capacity-pushback statuses, which the circuit breaker never counts, as either failure or success. Replaces the default set.

ZMQ workers: see ZMQ Workers.

Field Type Default Description
dp_size integer none On an ipc:// worker, the number of data-parallel engines that dial into its one socket set. A value above 1 makes a grouped worker, and the gateway spreads requests across the group's engines. On other workers the gateway sets this field itself during data-parallel discovery.
zmq_handshake_address string derived tcp:// address the gateway binds for the engine handshake. The default is tcp://127.0.0.1:<port>, with the port (20000 to 29999) derived from the ipc:// path. Set it for an engine that dials a fixed address, such as tcp://127.0.0.1:30500, TokenSpeed's default.

Registration of a ZMQ worker fails when the runtime is not vllm or tokenspeed, when worker_type is not regular (disaggregated workers need gRPC), when no model ID is available, when dp_size is above 1 on a gateway running with --dp-aware, or when the handshake address is not tcp:// or is already bound by another ZMQ worker. Setting zmq_handshake_address on a non-ZMQ worker also fails registration. Health checks stay on for ZMQ workers, because the probe is what reconnects a restarted engine.

PD disaggregation and KV transfer: see PD Disaggregation.

Field Type Default Description
bootstrap_port integer none KV bootstrap port of a prefill or encode worker. The gRPC and vLLM Mooncake paths use 8998 when it is unset.
kv_connector string engine-reported KV connector, such as MooncakeConnector or NixlConnector. Overrides the value a vLLM gRPC engine reports.
kv_engine_id string engine-reported vLLM kv_transfer_config.engine_id, used for Mooncake PD. For gRPC prefill and decode workers, the gateway reads it from the engine again when the worker recovers from failed or not_ready (a restarted engine gets a new ID), and the engine's value wins.

health, http_pool, and resilience take effect but are not echoed back by GET /workers; overload is.


Worker Errors

Worker endpoints report errors as {"error": "<message>", "code": "<CODE>"}, not in the shape described under Error Responses:

{
  "error": "Invalid value for field 'worker_url': 10.0.0.5:8000 - URL must start with a lowercase http://, https://, grpc://, grpcs://, or ipc:// scheme",
  "code": "BAD_REQUEST"
}
Status code When
400 BAD_REQUEST worker_id is not a UUID. The url is empty, does not start with a lowercase http://, https://, grpc://, grpcs://, or ipc:// scheme, is not a valid URL or has no host, or is an ipc:// URL without a path. A PUT changes the URL or is sent to a gateway running with --dp-aware.
400 PROVIDER_NOT_COMPILED The spec targets a provider whose router is not compiled into this build.
404 WORKER_NOT_FOUND No worker has this ID.
409 WORKER_ALREADY_EXISTS POST for a URL that is already registered. The message names the existing ID.
409 WORKER_CREATE_IN_PROGRESS POST for a URL whose registration is still running. The message names the ID to poll.
500 INTERNAL_SERVER_ERROR The job queue is unavailable or full.

A body that does not parse never reaches these checks. The gateway answers with a plain-text body: 415 when Content-Type: application/json is missing, 400 for malformed JSON, and 422 (Failed to deserialize the JSON body into the target type: ...) for a missing url, a wrong type, or an unknown value such as "worker_type": "prefil".

Other problems surface in the background job, after the 202:

  • overload.waiting_requests of 0, or overload.token_usage outside (0.0, 1.0]
  • the ZMQ constraints listed under Worker Spec
  • an unknown tool_parser or reasoning_parser
  • a worker that never answers, or http_pool.http2: true for a worker that only speaks HTTP/1.1

List Workers

GET /workers
GET /workers?model={model_id}

Returns every registered worker. model keeps only workers that serve that model ID or alias; workers without a model list match any model. The stats counts cover the returned workers (encode workers are counted in total only). Workers whose registration job is still running are not listed.

Response: 200 OK (abridged: discovered labels and card fields vary by engine)

{
  "workers": [
    {
      "id": "01997a4e-3c2b-7f1d-9a8e-2b6c4d5e7f80",
      "model_id": "meta-llama/Llama-3.1-8B-Instruct",
      "url": "grpc://gpu2:50051",
      "models": [
        {
          "id": "meta-llama/Llama-3.1-8B-Instruct",
          "aliases": ["llama-3.1-8b"],
          "model_type": ["chat", "completions", "responses", "tools"],
          "context_length": 131072,
          "tool_parser": "llama"
        }
      ],
      "worker_type": "regular",
      "connection_mode": "grpc",
      "runtime_type": "sglang",
      "labels": {
        "model_path": "meta-llama/Llama-3.1-8B-Instruct",
        "tp_size": "1"
      },
      "priority": 50,
      "cost": 1.0,
      "max_connection_attempts": 20,
      "is_healthy": true,
      "status": "ready",
      "load": 2,
      "http2": false
    }
  ],
  "total": 1,
  "stats": {
    "prefill_count": 0,
    "decode_count": 0,
    "regular_count": 1
  }
}

Each worker object carries its spec fields (never api_key) plus:

Field Description
id Worker UUID.
model_id ID of the first model card. Absent for workers without a model list.
is_healthy true when status is ready.
status pending (registered, not yet proven healthy), ready, not_ready (failing probes, may recover), failed, or draining (being removed). Only ready workers receive traffic.
load Requests the gateway has in flight to this worker.
http2 Whether the gateway speaks HTTP/2 to this worker.
pd_pairing Prefill and decode workers only: the key used to pair them for KV transfer, either the explicit pairing_protocol or runtime/transport/layout.
engine_load The engine's last polled load report (per-DP-rank queue and KV-cache figures). Absent until the load monitor has polled the worker.

Get Worker

GET /workers/{worker_id}

Returns one worker in the same shape as a GET /workers entry. While a job for the worker is queued or running, or for about five minutes after one fails, the response also carries job_status. A job that succeeds clears it.

While POST /workers is still registering the worker, the response is a placeholder: only id, url, status, and job_status are meaningful, and the other fields hold defaults.

Response: 200 OK

{
  "id": "01997a4e-3c2b-7f1d-9a8e-2b6c4d5e7f80",
  "url": "grpc://gpu2:50051",
  "worker_type": "regular",
  "connection_mode": "http",
  "runtime_type": "unspecified",
  "priority": 50,
  "cost": 1.0,
  "max_connection_attempts": 20,
  "is_healthy": false,
  "status": "pending",
  "load": 0,
  "http2": false,
  "job_status": {
    "job_type": "AddWorker",
    "worker_url": "grpc://gpu2:50051",
    "status": "processing",
    "message": null,
    "timestamp": 1790250060
  }
}
job_status field Description
job_type AddWorker (create or replace), UpdateWorker, or RemoveWorker.
worker_url URL of the worker the job targets.
status pending (queued), processing, or failed.
message Error text when status is failed; otherwise null.
timestamp Unix time, in seconds, of the last status change.

Registration keeps probing a worker that does not answer, every --worker-startup-check-interval seconds (default 30), until --worker-startup-timeout-secs (default 1800) runs out. If a POST /workers job fails, the gateway releases the reserved ID: GET /workers/{worker_id} returns 404, and the reason appears only in the gateway log (Failed job: type=AddWorker ...). A failed PATCH, PUT, or DELETE job leaves the worker in place with a failed job_status.


Update Worker (partial)

PATCH /workers/{worker_id}

Changes a few fields in place. Omitted fields keep their current values. The worker keeps its status and connections, and nothing is probed again.

Request Body:

{
  "priority": 75,
  "labels": {"tier": "gold"},
  "api_key": "new-api-key",
  "health": {"check_interval_secs": 15, "drain_settle_secs": 60}
}
Field Type Description
priority integer New priority.
cost number New cost.
labels object Merged into the current labels: listed keys are added or overwritten, and the rest are kept. PATCH cannot remove a label.
api_key string New worker API key, for key rotation.
health object Partial health overrides, merged into the worker's current settings: timeout_secs, check_interval_secs, success_threshold, failure_threshold, disable_health_check, drain_settle_secs. Setting disable_health_check to true makes the worker routable immediately.

Any other field, such as overload, http_pool, resilience, or models, is ignored. Use PUT to change those.

In v1.11.0, a PATCH also rebuilds the worker's circuit breaker with built-in thresholds (5 failures, 2 successes, 30 seconds), ignoring the gateway's circuit-breaker flags and any resilience overrides until a PUT re-registers the worker.

Response: 202 Accepted

{
  "status": "accepted",
  "worker_id": "01997a4e-3c2b-7f1d-9a8e-2b6c4d5e7f80",
  "message": "Worker update queued for background processing"
}

Replace Worker (full)

PUT /workers/{worker_id}

Replaces the worker's spec and re-runs the full registration workflow (probe, engine detection, metadata discovery) under the same ID. The body is a complete worker spec. Fields you omit take their defaults rather than keeping their current values, so resend any labels, overload, or other settings you want to keep.

The url must equal the worker's current URL; to move a worker, DELETE it and POST the new URL. PUT is refused with 400 when the gateway runs with --dp-aware, because one spec expands to one worker per data-parallel rank; use DELETE and POST there too. If the worker is deleted or replaced again before the job runs, the job fails rather than overwrite the newer state.

Response: 202 Accepted, with the same body as PATCH.


Delete Worker

DELETE /workers/{worker_id}

Removes a worker. A ready worker first moves to draining: it receives no new requests, and the gateway waits health.drain_settle_secs (default --drain-settle-secs, 5 seconds) so in-flight requests can finish before it removes the worker. A worker in any other status is removed without waiting.

Response: 202 Accepted

{
  "status": "accepted",
  "worker_id": "01997a4e-3c2b-7f1d-9a8e-2b6c4d5e7f80",
  "message": "Worker removal queued for background processing"
}

Cache Management

Flush the engines' KV caches, and read the engine load the gateway has polled from its workers.

Flush Cache

POST /flush_cache

Asks every registered HTTP and gRPC worker to drop its KV prefix cache. The calls go out in parallel, and the response reports the outcome per worker. Requires admin authentication.

Worker How it is flushed
HTTP POST {worker_url}/flush_cache, with the worker's API key (if it has one) as a bearer token and a 45-second timeout. Any non-2xx answer counts as a failure, so the engine must serve this route (SGLang's HTTP server does).
gRPC SGLang, TokenSpeed FlushCache RPC.
gRPC vLLM FlushCache RPC, new in v1.11.0. Needs smg-grpc-servicer 0.12.0 or later; older servicers answer UNIMPLEMENTED. Resets vLLM's local prefix cache without preempting running requests or clearing connector caches. The gateway makes a single attempt, so a reset that vLLM refuses while requests are in flight is reported under failed.
gRPC TRT-LLM, MLX Not supported: reported under failed (UNIMPLEMENTED).
ZMQ (ipc://) Skipped, because the transport has no cache-flush RPC. Counted in total_zmq_workers_skipped.

The fan-out does not filter by runtime: external-provider workers registered over HTTP receive the POST too, and appear under failed if they reject it. The endpoint only calls the workers; it does not reset the gateway's own cache-aware routing state.

Response: 200 OK when every worker that was called succeeds (or there was nothing to call), 206 Partial Content when at least one fails, including when all of them fail.

{
  "status": "success",
  "message": "Successfully flushed cache on all 3 workers",
  "workers_flushed": 3,
  "total_http_workers": 2,
  "total_grpc_workers": 1,
  "total_zmq_workers_skipped": 0,
  "total_workers": 3
}

On a partial failure, status is "partial_success" and two more fields list the outcome, successful (worker URLs) and failed ({worker, error} entries):

{
  "status": "partial_success",
  "message": "Cache flush: 2 succeeded, 1 failed (1 ZMQ workers skipped: no cache-flush RPC)",
  "workers_flushed": 2,
  "total_http_workers": 2,
  "total_grpc_workers": 1,
  "total_zmq_workers_skipped": 1,
  "total_workers": 4,
  "successful": ["http://gpu1:8000", "grpc://gpu3:50051"],
  "failed": [
    {
      "worker": "http://gpu2:8000",
      "error": "flush_cache failed for worker http://gpu2:8000/flush_cache: HTTP 500 Internal Server Error"
    }
  ]
}
Field Description
status success or partial_success
message Human-readable summary
workers_flushed Workers that confirmed the flush
total_http_workers, total_grpc_workers Registered workers by transport
total_zmq_workers_skipped ZMQ workers that were not called
total_workers All registered workers (HTTP + gRPC + ZMQ)
successful, failed Present only when at least one worker failed

Get Loads

GET /loads
GET /get_loads

Returns the engine load the gateway's load monitor last collected: one entry per worker per DP rank, plus a fleet-wide aggregate. The body uses the same schema engines report on /v1/loads, with every entry tagged by worker and worker_type.

The response is built from the monitor's cached snapshot, the same numbers the routing policies act on, so no request reaches a worker. Values can be up to one poll interval old (--load-monitor-interval, 10 seconds by default). Serving from the cache is also what makes the route safe to poll from another SMG gateway that registers this one as a worker.

  • Auth: /loads is a public route, served without auth like /health. /get_loads is a deprecated alias that stays on the admin tier and requires admin authentication.
  • Filter: ?model=<model_id> limits the response to workers serving that model.
  • Missing workers: workers without a current report (not Ready yet, no load source, or a failed last poll) are left out rather than reported as idle. Before the first poll, loads is empty and aggregate is absent.

Response: 200 OK

{
  "timestamp": "2026-09-24T18:20:11.482913+00:00",
  "version": "smg-1.11.0",
  "dp_rank_count": 2,
  "loads": [
    {
      "worker": "http://gpu1:8000",
      "worker_type": "regular",
      "dp_rank": 0,
      "num_running_reqs": 12,
      "num_waiting_reqs": 3,
      "num_waiting_uncached_tokens": 5120,
      "num_total_reqs": 15,
      "num_used_tokens": 48000,
      "max_total_num_tokens": 131072,
      "token_usage": 0.37,
      "gen_throughput": 1850.5,
      "cache_hit_rate": 0.62,
      "utilization": 0.41,
      "max_running_requests": 256
    },
    {
      "worker": "http://gpu2:8000",
      "worker_type": "regular",
      "dp_rank": 0,
      "num_running_reqs": 9,
      "num_waiting_reqs": 0,
      "num_waiting_uncached_tokens": 0,
      "num_total_reqs": 9,
      "num_used_tokens": 30000,
      "max_total_num_tokens": 131072,
      "token_usage": 0.23,
      "gen_throughput": 1420.0,
      "cache_hit_rate": 0.55,
      "utilization": 0.3,
      "max_running_requests": 256
    }
  ],
  "aggregate": {
    "total_running_reqs": 21,
    "total_waiting_reqs": 3,
    "total_reqs": 24,
    "avg_token_usage": 0.3,
    "avg_throughput": 1635.25,
    "avg_utilization": 0.355
  }
}
Field Description
timestamp When the gateway built this response (RFC 3339)
version smg- followed by the gateway version
dp_rank_count Number of entries in loads, across the whole fleet
loads[].worker, loads[].worker_type Worker URL and type (regular, prefill, decode, or encode)
loads[].dp_rank DP rank within that worker
loads[].num_running_reqs, num_waiting_reqs, num_total_reqs Requests running, queued, and both
loads[].num_waiting_uncached_tokens Queued tokens not served from the prefix cache (0 when the engine does not report it)
loads[].num_used_tokens, max_total_num_tokens KV tokens in use and KV capacity (0 for workers read from Prometheus /metrics)
loads[].token_usage KV token usage ratio, 0.0 to 1.0
loads[].gen_throughput, cache_hit_rate, utilization, max_running_requests As reported by the engine
loads[].memory, loads[].queues Optional sections, present when the engine reports them: memory (weight_gb, kv_cache_gb, graph_gb, token_capacity) and queues (waiting, grammar, paused, retracted)
loads[].kv_transfer_latency_ms, kv_transfer_speed_gb_s, prefill_queue_reqs, decode_queue_reqs, disagg_mode Optional prefill/decode disaggregation fields, present when the engine reports them
aggregate Fleet roll-up across all entries: total_running_reqs, total_waiting_reqs, and total_reqs are sums; avg_token_usage, avg_throughput, and avg_utilization are means

The same per-worker report appears as engine_load on GET /workers and GET /workers/{worker_id}. It is unrelated to their load field, which counts requests the gateway has in flight to the worker. Overload Protection explains where each backend's report comes from.


Model Information

Query model and server information. These are public routes and need no authentication.

List Models

GET /v1/models

Returns the models that the registered self-hosted workers serve, read from the gateway's worker registry rather than from the workers. A caller that sends its own provider key gets the external providers' model list instead, fetched with that key. See List Models for the details.

Response: 200 OK

{
  "object": "list",
  "data": [
    {
      "id": "meta-llama/Llama-3.1-8B-Instruct",
      "object": "model",
      "created": 0,
      "owned_by": "self_hosted"
    }
  ]
}

created is always 0. With no self-hosted model to list, the route answers 503 with the plain-text body No models available.


Get Model Info

GET /get_model_info

Forwards the request to the first healthy regular HTTP worker (the first healthy prefill worker in HTTP PD mode) and returns that worker's answer. SGLang serves this route, and its answer includes fields such as model_path, tokenizer_path, and is_generation. An engine without the route, such as vLLM, answers 404. In regular mode the worker's status and body pass through unchanged; in HTTP PD mode an error status comes back in the gateway's error envelope. With no healthy worker to ask, the route answers 503.


Get Server Info

GET /get_server_info

Forwards the request the same way as Get Model Info and returns the worker's answer. SGLang returns its server arguments together with scheduler state and its version.


WASM Module Management

Manage WebAssembly plugins. Modules are registered from files accessible to the gateway process; the request body contains descriptors with paths, not binary payloads.

These routes need --enable-wasm. Without it, each answers 500 Internal Server Error with an empty body.

Add WASM Module

POST /wasm

Registers one or more WASM modules.

Request Body: JSON WasmModuleAddRequest

{
  "modules": [
    {
      "name": "custom-middleware",
      "file_path": "/etc/smg/wasm/custom-middleware.wasm",
      "module_type": "Middleware",
      "attach_points": [
        {"Middleware": "OnRequest"},
        {"Middleware": "OnResponse"}
      ]
    }
  ]
}

The only supported module_type today is Middleware. Valid Middleware attach points are OnRequest, OnResponse, and OnError, but OnError modules never run in v1.11.0: the middleware calls only OnRequest and OnResponse modules, and skips OnResponse for streaming responses.

Response: 200 OK on full success, 400 Bad Request if any module failed to register. The response body echoes every requested module with an add_result field indicating success (carrying the assigned UUID) or failure (carrying the error message).

{
  "modules": [
    {
      "name": "custom-middleware",
      "file_path": "/etc/smg/wasm/custom-middleware.wasm",
      "module_type": "Middleware",
      "attach_points": [
        {"Middleware": "OnRequest"},
        {"Middleware": "OnResponse"}
      ],
      "add_result": {
        "Success": "550e8400-e29b-41d4-a716-446655440000"
      }
    }
  ]
}

List WASM Modules

GET /wasm

Returns all registered WASM modules together with aggregate execution metrics.

Response: 200 OK

{
  "modules": [
    {
      "module_uuid": "550e8400-e29b-41d4-a716-446655440000",
      "module_meta": {
        "name": "custom-middleware",
        "file_path": "/etc/smg/wasm/custom-middleware.wasm",
        "sha256_hash": "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855",
        "size_bytes": 65536,
        "created_at": "2024-01-15T12:00:00.000000000Z",
        "last_accessed_at": "2024-01-15T12:05:00.000000000Z",
        "access_count": 40,
        "attach_points": [
          {"Middleware": "OnRequest"}
        ]
      }
    }
  ],
  "metrics": {
    "total_executions": 40,
    "successful_executions": 40,
    "failed_executions": 0,
    "total_execution_time_ms": 120,
    "max_execution_time_ms": 8,
    "average_execution_time_ms": 3.0
  }
}

metrics covers every module. average_execution_time_ms is total_execution_time_ms divided by total_executions, and is left out until a module has run.


Remove WASM Module

DELETE /wasm/{module_uuid}

Removes a WASM module. The body is a plain text status message, not JSON.

Response: 200 OK

Module removed successfully

On failure returns 400 Bad Request with the error text as the body. A module_uuid that is not a UUID gets 400 with an empty body.


Error Responses

The admin endpoints do not share one error format. The body depends on the part of SMG that rejects the request:

Source Status Body
Control-plane auth 401, 403 Plain text, such as Missing or invalid Authorization header, Invalid authentication token, or Admin role required for control plane access. A 401 carries WWW-Authenticate: Bearer realm="control-plane"
--api-key auth, when control-plane auth is not configured 401 Empty
A request body that does not parse 400, 415, 422 Plain text from the JSON extractor
Worker endpoints 400, 404, 409, 500 {"error": "<message>", "code": "<CODE>"}
Get Tokenizer, Get Tokenizer Status 404 {"error": {"message": "<message>", "type": "tokenizer_not_found"}}, with "type": "not_found" from Get Tokenizer Status
Add Tokenizer 409, 503 The 202 body shape with "status": "failed"
Remove Tokenizer 404 {"success": false, "message": "<message>"}
/parse/function_call, /parse/reasoning 400, 503 {"error": "<message>", "success": false}
Add WASM Module 400 The module list, with "add_result": {"Error": "<message>"} on each module that failed
Remove WASM Module, and every WASM route without --enable-wasm 400, 500 Plain text or empty
The gateway, on forwarded routes such as /get_model_info 503 and others The gateway error envelope below

Errors that the gateway generates itself, for example when no worker is available, use this envelope and repeat code in the X-SMG-Error-Code response header:

{
  "error": {
    "type": "Service Unavailable",
    "code": "no_workers",
    "message": "No workers are available",
    "param": null
  }
}

type is the reason phrase of the HTTP status (Not Found, Service Unavailable, ...), and param is always null. The inference endpoints use the same envelope; see Error Responses in the OpenAI-compatible API reference.


Authentication

The tokenizer, worker, WASM, cache flush, and deprecated /get_loads endpoints on this page are control-plane routes. Send the credential as Authorization: Bearer <token>:

  1. With control-plane auth configured (--control-plane-api-keys, or --jwt-issuer with --jwt-audience): a control-plane API key or JWT whose role is admin. A missing or invalid credential gets 401, and one with another role gets 403. The shared --api-key is not accepted.
  2. Without it: the shared --api-key. Per-tenant --tenant-api-key keys are rejected, so a gateway with tenant keys but no --api-key answers 401 to every call. With no keys configured at all, the routes are open.

Public endpoints need no authentication: the health probes, /loads, /engine_metrics, and the Model Information routes. See Auth Model for every route's tier.