ZMQ Direct Workers¶
When SMG and the inference engine share a host, SMG can connect straight to the engine core over ZMQ. No engine HTTP server or gRPC servicer sits in the path: the engine runs headless (scheduling and model execution only), SMG does everything else — tokenization, chat templates, parsing, stop handling, and routing — and the two exchange token IDs over local ipc:// sockets. Direct ZMQ connections were added in v1.10.0.
When to Use ZMQ¶
| HTTP | gRPC | ZMQ | |
|---|---|---|---|
| Worker URL | http://host:port |
grpc://host:port |
ipc:///path |
| Engine-side process | Engine's OpenAI-compatible server | Engine with a gRPC servicer | Headless engine core only |
| Tokenization, chat templates, parsers | Engine | Gateway | Gateway |
| Engine location | Any reachable host | Any reachable host | Same host as SMG |
| Prefill-decode disaggregation | Yes | Yes | No |
| KV-event stream for cache-aware routing | No | Yes | No |
Compared with gRPC, the direct path drops the servicer process, a serialization round-trip, and a context switch per message. Use it when the gateway and engine run on one machine. Use gRPC when the engine runs on another host, or when you need disaggregation, KV-event routing, embeddings, or engine admin operations (see Limits).
How It Works¶
For a worker URL ipc://<path>:
- SMG binds two data-plane sockets,
ipc://<path>-in.sock(requests) andipc://<path>-out.sock(outputs), plus a handshake listener attcp://127.0.0.1:<port>, where the port is derived from<path>(see Handshake Address). - The headless engine dials the handshake address. SMG answers with the data-plane addresses (a HELLO, INIT, READY exchange); the engine connects to them and reports its context length and data-parallel size.
- SMG marks the worker ready as soon as the handshake completes.
- Requests go to the engine as token IDs. Outputs come back as token IDs, with the engine's scheduler load piggybacked on the output batches.
The engine never sees text, so SMG also does the work that the engine's API server or gRPC servicer would otherwise do:
| Task | What SMG does |
|---|---|
| Tokenization, chat templates, reasoning and tool-call parsing | Same as in gRPC mode |
| String stops | Sends a stop string that encodes to a single token as a stop token ID, and matches longer stop strings on the decoded text and trims them. For Harmony (gpt-oss) models, stops are matched on the parsed channel text |
| EOS (vLLM) | Attaches the model's EOS token IDs to every request: from config.json and generation_config.json when --model-path is a local directory, otherwise from the tokenizer. TokenSpeed stops at EOS on its own |
Default max_tokens (vLLM) |
Uses the context length the engine reported at handshake minus the prompt length |
n > 1 |
Sends n single-sample requests and merges the results; an explicit seed becomes seed + i for sample i, so samples differ but stay reproducible |
| Data-parallel ranks | Picks the least-loaded engine inside a grouped worker (see Data-Parallel Engines) |
Supported Engines¶
| Engine | Runtime | Headless engine command | Notes |
|---|---|---|---|
| vLLM | vllm |
vllm serve <model> --headless ... |
Assumed when a ZMQ worker declares no runtime |
| TokenSpeed | tokenspeed |
python -m tokenspeed.cli serve --headless ... |
Must be declared. Needs --grammar-backend for structured outputs and --enable-output-logprobs for logprobs |
Both engines use the same handshake, so SMG can't tell which one dialed in. Declare the runtime with --backend (startup workers) or runtime_type (worker API). A ZMQ worker without a runtime is treated as vLLM, and SMG logs runtime_type unspecified for ZMQ worker; defaulting to vLLM EngineCore.
SGLang, TensorRT-LLM, and MLX can't use ZMQ in v1.11.0. smg serve accepts --connection-mode zmq only with --backend vllm or --backend tokenspeed, and the gateway rejects any other runtime with ZMQ worker ... has unsupported runtime ...: only vllm and tokenspeed are supported over the ZMQ direct backend. Connect those engines over gRPC.
SMG doesn't check engine versions at connect time, and its decoders accept fields that newer engine releases append. SMG's CI runs the ZMQ suites against vLLM 0.27.1 and a pinned TokenSpeed revision.
Quick Start with smg serve¶
smg serve --connection-mode zmq starts headless engines and the gateway together:
smg serve \
--backend vllm \
--connection-mode zmq \
--model meta-llama/Llama-3.1-8B-Instruct \
--router-model-path meta-llama/Llama-3.1-8B-Instruct \
--host 0.0.0.0 \
--port 30000smg serve \
--backend tokenspeed \
--connection-mode zmq \
--model meta-llama/Llama-3.1-8B-Instruct \
--router-model-path meta-llama/Llama-3.1-8B-Instruct \
--host 0.0.0.0 \
--port 30000What smg serve sets up in ZMQ mode:
| Item | Behavior |
|---|---|
| Worker URL | ipc://<socket dir>/engine-<worker port> per replica. The socket directory is $SMG_ZMQ_SOCKET_DIR, default /tmp/smg-zmq-<uid> |
| Engine command | A headless engine that dials the handshake port derived from its worker URL (commands below) |
| TokenSpeed defaults | Adds --grammar-backend xgrammar and --enable-output-logprobs unless you pass your own value for either flag |
| Engine health wait | Skipped. The gateway starts at once, and /readiness returns 503 until a worker has completed its handshake and the tokenizer has loaded |
| Gateway runtime | --backend is forwarded so the gateway speaks the engine's wire protocol |
| Replicas | --data-parallel-size N starts N independent single-engine workers, each with its own socket path and GPU slice. The launch stops early if two replicas' paths derive the same handshake port (ZMQ handshake port collision on ...); change --worker-base-port |
The launcher builds these engine commands. If you pass a flag the launcher sets, your copy is dropped, except for the two TokenSpeed defaults above, where your value is used:
# vLLM
python -m vllm.entrypoints.cli.main serve <model> --headless \
--data-parallel-size 1 --data-parallel-size-local 1 \
--data-parallel-address 127.0.0.1 --data-parallel-rpc-port <handshake port>
# TokenSpeed
python -m tokenspeed.cli serve --headless --model <model> \
--port <worker port> --dist-init-addr 127.0.0.1:<store port> \
--data-parallel-address 127.0.0.1 --data-parallel-rpc-port <handshake port> \
--zmq-engine-index 0 --grammar-backend xgrammar --enable-output-logprobsTo put several data-parallel engines behind one worker, start the engines yourself and use smg launch --zmq-engine-count (see Data-Parallel Engines).
Manual Setup with smg launch¶
Socket Path¶
Pick a path for each worker, for example ipc:///tmp/smg-zmq/engine-0. SMG creates the parent directory with mode 0700 if it doesn't exist. An existing directory must be a real directory (not a symlink) owned by the user SMG runs as, or SMG refuses to bind, so don't place sockets directly in a shared directory such as /tmp. The engine must be able to reach the directory: run it as the same user, or create the directory yourself with the access the engine needs.
Before binding, SMG deletes any socket files already at <path>-in.sock and <path>-out.sock, treating them as leftovers from an earlier gateway run. Give each gateway process its own socket paths: a second gateway configured with the same worker URL would delete the first one's live sockets.
Handshake Address¶
SMG binds the handshake on 127.0.0.1 at a port in 20000–29999, computed from the path after ipc:// as 20000 + FNV-1a-64(path) mod 10000. Compute it before starting the engine:
python3 - <<'EOF'
path = "/tmp/smg-zmq/engine-0" # the worker URL without "ipc://"
h = 0xCBF29CE484222325
for b in path.encode():
h = ((h ^ b) * 0x100000001B3) & 0xFFFFFFFFFFFFFFFF
print(20000 + h % 10000)
EOFFor this path it prints 22670. SMG also logs the address when it binds:
Binding ZMQ client for worker ipc:///tmp/smg-zmq/engine-0 (handshake=tcp://127.0.0.1:22670, engines=1)Two paths can hash to the same port; SMG rejects the second worker at registration. Workers added through the worker API can set a fixed zmq_handshake_address instead.
Start the Engine¶
vllm serve meta-llama/Llama-3.1-8B-Instruct \
--headless \
--data-parallel-size 1 \
--data-parallel-size-local 1 \
--data-parallel-address 127.0.0.1 \
--data-parallel-rpc-port 22670python -m tokenspeed.cli serve \
--headless \
--model meta-llama/Llama-3.1-8B-Instruct \
--data-parallel-address 127.0.0.1 \
--data-parallel-rpc-port 22670 \
--zmq-engine-index 0 \
--grammar-backend xgrammar \
--enable-output-logprobsTokenSpeed's grammar backend defaults to none, and then the engine ignores structured-output constraints: a tool_choice: "required" or json_schema request comes back as free text. Output logprobs are off by default. When several TokenSpeed engines share a host, give each its own --port; TokenSpeed derives its internal control ports from it.
Add your usual engine flags, such as --tensor-parallel-size.
Start the Gateway¶
smg launch \
--worker-urls ipc:///tmp/smg-zmq/engine-0 \
--backend vllm \
--model-path meta-llama/Llama-3.1-8B-Instruct \
--host 0.0.0.0 \
--port 30000Use --backend tokenspeed for a TokenSpeed engine. --model-path is required: it names the model and loads the tokenizer. SMG binds the sockets as soon as it registers the worker, then waits up to 600 seconds for each handshake message, which gives the engine time to load the model and profile its KV cache. If the handshake fails, the next health probe starts a new attempt.
To serve several engines, list one ipc:// URL per engine. They are independent workers balanced by --policy.
Register Workers at Runtime¶
Add a ZMQ worker to a running gateway with POST /workers. The ipc:// scheme selects ZMQ, so connection_mode can be left out:
curl -X POST http://localhost:30000/workers \
-H "Content-Type: application/json" \
-d '{
"url": "ipc:///tmp/smg-zmq/engine-1",
"runtime_type": "tokenspeed"
}'ZMQ-related worker fields:
| Field | Description |
|---|---|
url |
ipc://<path> (the path is required) |
runtime_type |
vllm or tokenspeed (runtime is accepted as an alias). Unset means vLLM |
zmq_handshake_address |
A tcp:// address to bind for the handshake instead of the derived one. Rejected on non-ZMQ workers |
dp_size |
Number of engines in a grouped worker; leave dp_rank unset |
labels.model_path |
Model ID and tokenizer source, for gateways started without --model-path |
zmq_handshake_address pairs a worker with an engine that dials a fixed address. A TokenSpeed engine started without --data-parallel-rpc-port dials tcp://127.0.0.1:30500, so this worker pairs with it:
curl -X POST http://localhost:30000/workers \
-H "Content-Type: application/json" \
-d '{
"url": "ipc:///tmp/smg-zmq/tokenspeed-0",
"runtime_type": "tokenspeed",
"zmq_handshake_address": "tcp://127.0.0.1:30500"
}'Registration runs in the background (the API answers 202 Accepted), so a rejected worker shows up only in the gateway log; see Troubleshooting. For the full request schema, see Worker Management.
Data-Parallel Engines¶
There are two ways to run data parallelism over ZMQ:
- Replicas: several single-engine workers, each with its own socket path. The gateway's routing policy spreads requests across them.
smg serve --data-parallel-size Nsets this up. - Grouped worker: one worker URL and one socket set, with
Ndata-parallel engines behind it. The gateway sees one worker and picks the engine (rank) for each request itself.
Configure a Grouped Worker¶
Start the engine group with its data-parallel size, then tell the gateway how many engines to wait for:
vllm serve meta-llama/Llama-3.1-8B-Instruct \
--headless \
--data-parallel-size 2 \
--data-parallel-size-local 2 \
--data-parallel-address 127.0.0.1 \
--data-parallel-rpc-port 22670smg launch
--worker-urls ipc:///tmp/smg-zmq/engine-0
--backend vllm
--model-path meta-llama/Llama-3.1-8B-Instruct
--zmq-engine-count 2
python -m tokenspeed.cli serve
--headless
--model meta-llama/Llama-3.1-8B-Instruct
--data-parallel-size 2
--dist-init-addr 127.0.0.1:31233
--data-parallel-address 127.0.0.1
--data-parallel-rpc-port 22670
--zmq-engine-index 0
--grammar-backend xgrammar
--enable-output-logprobs
smg launch
--worker-urls ipc:///tmp/smg-zmq/engine-0
--backend tokenspeed
--model-path meta-llama/Llama-3.1-8B-Instruct
--zmq-engine-count 2
Each TokenSpeed rank dials in with its own identity (--zmq-engine-index plus its rank). TokenSpeed requires an explicit --dist-init-addr when --data-parallel-size is above 1; pick a free port outside the 20000–29999 handshake range.
--zmq-engine-countapplies to everyipc://URL in--worker-urlsand must be a positive integer. For a worker added through the API, set"dp_size"on the worker instead.- The engine process needs
Ntimes its tensor-parallel size in GPUs. - The gateway must wait for exactly as many engines as dial in. An extra engine can fail the handshake with
duplicate HELLO ... after INIT phase; a missing one leaves the handshake waiting until it times out. - Grouped workers can't be combined with
--dp-aware, which tries to expand the group into one worker per rank and fails withcannot be dp-aware expanded. Single-engine ZMQ workers register normally under--dp-aware.
Rank Selection¶
Output batches carry the producing rank's scheduler load: running requests, waiting requests, and KV-cache usage (vLLM's scheduler stats, or TokenSpeed's load snapshot). For each request, SMG scores every rank in the group and sends the request to the lowest score:
score = max(requests SMG has in flight on the rank, running + waiting)
+ waiting × 6 × max(0, kv_cache_usage − 0.5)- The in-flight floor keeps a stale report from making a busy rank look idle, so a burst of requests between reports spreads across ranks.
- Engines report load only while they produce output. When a rank's last in-flight request finishes, SMG resets that rank's running and waiting counts to zero and keeps its last reported KV usage.
- Ties rotate across ranks. A TokenSpeed build that sends no load snapshot is scored on SMG's in-flight counts alone.
- vLLM MoE models run their data-parallel ranks in lockstep and pause the whole group when idle. SMG sends the wake-up that vLLM's data-parallel coordinator process would otherwise send, so no coordinator runs. Ranks of dense models run independently.
Rank selection happens inside the connection. Routing policies see the group as one worker, and requests are never pinned to a rank.
Worker Lifecycle¶
| Stage | What happens |
|---|---|
| Registered | The worker starts as pending. SMG binds its sockets and starts the handshake right away |
| Connected | SMG promotes the worker to ready the moment the handshake completes, without waiting for the health-check success threshold |
| Serving | Health probes read a local liveness flag; the engine has no health RPC. /readiness stays at 503 with tokenizer not yet registered until the model's tokenizer has loaded |
| Engine lost | SMG marks the connection dead when the engine signals that it died (vLLM does so on a fatal error, TokenSpeed also on shutdown), the socket fails, a request send blocks for 10 seconds, three outputs in a row can't be decoded, or no output arrives for 300 seconds while requests are in flight. In-flight requests fail. The next health probe drops the dead connection, and a later probe binds the sockets again so a restarted engine can reconnect |
| Removed | DELETE /workers/{worker_id} drains the worker, then removes it; a handshake still in progress is cancelled and its sockets are released |
- Health checks stay on for ZMQ workers even with
--disable-health-checkor a per-workerdisable_health_check, because the probe is what reconnects a restarted engine. SMG logsIgnoring disabled health checks for ZMQ worker .... - With
--remove-unhealthy-workers, a ZMQ worker whose engine stays down is removed and nothing re-adds it. The flag is off by default unless service discovery is enabled; leave it off for ZMQ workers so they rejoin when their engine restarts. - The 300-second silence limit is fixed. A single request whose prefill takes longer than that on an otherwise idle engine is failed as if the engine had died.
- Updating a ZMQ worker's properties, such as labels or priority, keeps its live connection.
Features and Limits¶
Supported over ZMQ¶
| Feature | vLLM | TokenSpeed |
|---|---|---|
| Chat Completions, Completions, Responses, and Messages, including streaming | Yes | Yes |
| Tool-call and reasoning parsing | Yes | Yes |
Structured outputs (response_format, constrained tool_choice, regex, grammar) |
Yes | Yes, with --grammar-backend set on the engine |
| String stops and EOS | Yes | Yes |
| Harmony (gpt-oss) stop strings | Yes | Yes |
| Multimodal inputs | Yes, one modality per request | Yes, except models whose processor emits image_grid_thw or video_grid_thw (MRoPE) |
| Output logprobs | Yes, including top_logprobs |
Sampled-token logprobs, with --enable-output-logprobs; top_logprobs above 1 is rejected |
| Prompt logprobs | Forwarded to the engine, but no API field requests them in v1.11.0 | Not supported; token_ids_logprob is rejected |
n > 1 and sampling seed |
Yes | Yes |
Limits¶
| Limit | What you see |
|---|---|
| Same host only | Sockets are local ipc:// paths with a loopback handshake. ZMQ workers are never synced to high-availability mesh peers; each belongs to the gateway that registered it |
| No prefill, decode, or encode workers | Registration fails with ZMQ worker ... cannot serve worker type .... PD and EPD routing select only gRPC workers |
| No KV-event stream | cache_aware tracks ZMQ workers with its approximate prefix tree, and SMG logs ... the ZMQ transport has no KV-event stream ... |
| No embeddings | 501 with code unsupported_backend: ZMQ backend does not support embeddings yet |
| No engine admin operations | POST /flush_cache skips ZMQ workers, and profiling fails with start_profile is not supported over ZMQ. LoRA loading, weight updates, and sleep/wake aren't available over ZMQ |
| No tokenizer from the engine | The gateway loads it from --model-path or --tokenizer-path (or the worker's model_path/tokenizer_path labels) |
| No rank pinning | See Rank Selection |
| vLLM multimodal | Mixed image and video in one request: 400, the vLLM ZMQ backend takes one modality per request .... Worker-side media processing (media references) needs a gRPC vLLM worker: 400, multimodal_not_supported |
| TokenSpeed logprobs | top_logprobs above 1, or token_ids_logprob: 400, ... not supported over the TokenSpeed ZMQ backend |
| TokenSpeed MRoPE models | 400: MRoPE position tensors are not derivable over the TokenSpeed ZMQ wire yet; use the gRPC transport for this model |
Verify¶
# Ready once the handshake has completed and the tokenizer has loaded
curl http://localhost:30000/readiness
# Transport, runtime, and status of each worker
curl -s http://localhost:30000/workers | jq '.workers[] | {url, connection_mode, runtime_type, status}'Expected worker entry:
{
"url": "ipc:///tmp/smg-zmq/engine-0",
"connection_mode": "zmq",
"runtime_type": "vllm",
"status": "ready"
}Send a request:
curl http://localhost:30000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "meta-llama/Llama-3.1-8B-Instruct",
"messages": [{"role": "user", "content": "Hello!"}]
}'Gateway metrics label ZMQ workers with connection_mode="zmq":
curl -s http://localhost:29000/metrics | grep 'connection_mode="zmq"'Troubleshooting¶
Worker stays pending
Symptoms: /readiness returns 503 with insufficient healthy workers, and /workers shows "status": "pending".
Solutions:
- Check that the engine dials the address SMG bound: compare the engine's
--data-parallel-rpc-portand--data-parallel-addresswith thehandshake=value in SMG'sBinding ZMQ client for worker ...log line. With no engine, SMG logsZMQ backend handshake failed for ...: ... startup handshake timed out while waiting for HELLO after 600s; will retry on the next health probe. - For a grouped worker, check that
--zmq-engine-count(ordp_size) matches the engine's--data-parallel-size. - A worker added through the API to a gateway started with
--disable-health-checkand no ZMQ workers at startup is never promoted. Restart the gateway without that flag.
Socket bind errors or stale socket files
Symptoms: The worker stays pending, and the gateway log shows ZMQ backend handshake failed for ... followed by a socket or directory error.
Solutions:
SMG deletes leftover socket files at <path>-in.sock and <path>-out.sock before binding, so a crashed gateway doesn't block a restart. Other causes:
ipc socket path ... exists but is not a socket; refusing to unlink: a regular file sits at the socket path. Move it away.ipc socket dir ... is owned by uid ...oripc socket dir ... exists but is not a directory: use a directory owned by SMG's user, not/tmpitself or a symlink.- Another process holds the handshake port. Choose a different socket path, or register the worker with
zmq_handshake_address. - The socket path is too long. Unix socket paths have a small length limit (108 bytes on Linux), so keep
<path>-out.sockshort.
Registration rejected
Symptoms: The worker never appears in /workers, and the gateway log shows Failed job: type=AddWorker, worker=ipc://... with the reason.
Solutions:
has no model identity: start SMG with--model-path(withsmg serve,--router-model-path), or set amodel_pathlabel on the worker.has unsupported runtime: usevllmortokenspeed.would bind handshake address ..., already claimed by worker ...: two socket paths hash to the same port. Rename one, or setzmq_handshake_address.zmq_handshake_address must be a tcp:// addressorZMQ worker URL must be ipc://<path>: fix the address or URL.cannot serve worker type: ZMQ workers must beregular.cannot be dp-aware expanded: drop--dp-aware, or use single-engine workers.
TokenSpeed requests fail or hang
Symptoms: The worker connects, but requests error or time out.
Solutions:
- Look for
runtime_type unspecified for ZMQ worker; defaulting to vLLM EngineCorein the gateway log. A TokenSpeed worker must be declared with--backend tokenspeedor"runtime_type": "tokenspeed", or SMG speaks the vLLM wire protocol to it. - If structured outputs come back unconstrained or logprobs are missing, start the engine with
--grammar-backend xgrammarand--enable-output-logprobs.
Readiness reports tokenizer not yet registered
Symptoms: /readiness returns 503 with tokenizer not yet registered after the worker is ready.
Solutions:
- Wait for the tokenizer download from
--model-pathto finish, and check the gateway log for load errors. - Set
--tokenizer-pathif the tokenizer lives somewhere other than the model path.
Next Steps¶
- gRPC Workers — Engines on other hosts, disaggregation, and KV-event routing
- Multiple Workers — Mix worker types and register workers at runtime
- Load Balancing — Policies for spreading traffic across ZMQ replicas
- Architecture Overview — Where the ZMQ path sits in the gateway
- Configuration Reference —
--worker-urlsand related worker flags - Monitoring — Gateway and engine-load metrics