Service Discovery¶
SMG discovers workers from Kubernetes pods and keeps its worker registry in step with the cluster. Pods that match a label selector and are Ready become workers; workers whose pods are deleted, stop matching, or turn unready are drained and removed.
Overview¶
Level-Triggered Reconcile¶
An informer keeps a local cache of pods. Every pass compares the whole cache with the registry, so a missed watch event is corrected on the next pass.
Readiness-Aware Scaling¶
Workers are added when their pods become Ready and drained when pods scale down, terminate, or turn unready.
Multi-Port Pods¶
A pod that runs several engine servers lists their ports in one annotation, and each port becomes its own worker.
PD Support¶
Separate discovery for prefill and decode workers in disaggregated deployments.
How It Works¶
Informer and Reconcile Loop¶
SMG runs a Kubernetes informer for pods. An initial LIST fills a local cache and a WATCH keeps it current. If the watch drops, the informer reconnects (with backoff after errors) and resumes from the last version it saw; it lists again only when the API server no longer has that version (HTTP 410 Gone). SMG passes the label selector to the API server, so the cache holds only candidate pods.
A reconcile pass turns the cache into the set of desired workers, compares it with the workers that discovery registered earlier, and submits AddWorker and RemoveWorker jobs to the control-plane job queue for the difference.
| Trigger | Behavior |
|---|---|
| Watch event | Any change to a cached pod starts a pass. Bursts, such as a rollout or a re-list, are coalesced: SMG waits 1 second, then runs one pass. |
| Periodic resync | A pass runs every 60 seconds. It reads the local cache only and makes no API calls. |
Because every pass compares complete state instead of replaying individual events, discovery converges even when events are missed:
- A pod deleted while the watch was disconnected drops out of the cache once the watch recovers, from the replayed delete event or a fresh list, and its workers are removed on the next pass.
- A registration that failed is submitted again on a later pass, because the worker is still missing from the registry.
- A worker registered for a pod that no longer exists, for example a registration that finished after its pod was deleted, is removed on the next pass.
While an add or remove job for an address is pending or running, later passes skip that address; completed and failed jobs do not block, so failures are retried.
Which Pods Become Workers¶
A pod contributes workers when all of the following hold:
- Its labels match
--selector(or a role selector in PD mode). - It is not terminating (no
deletionTimestamp). - It has a pod IP.
- Its phase is
Runningand itsReadycondition isTrue.
SMG routes to pod IPs directly, so it honors the pod's aggregate Ready condition, including readiness gates, the same way a Service endpoint does. Each matching pod contributes one worker per port (see Multi-Port Pods). A worker's address is <pod-ip>:<port>, with IPv6 addresses in brackets. Registration probes HTTP and gRPC on that address and keeps the protocol that answers, preferring HTTP when both do.
Ownership¶
Every worker that discovery registers carries two labels, smg.ai/pod-name and smg.ai/pod-uid. The reconciler only adds and removes workers that carry smg.ai/pod-uid and were registered by this gateway, so:
- Workers added with
--worker-urlsor through the worker API are never touched by discovery. - Workers synchronized from mesh peers are left to the gateway that owns them.
- When a pod is replaced by one with a new UID at the same address (for example, a restart that keeps its IP), the old worker is removed and the new one registered.
Configuration¶
Basic Setup¶
smg \
--service-discovery \
--selector app=sglang-worker \
--service-discovery-namespace inference \
--service-discovery-port 8000Parameters¶
| Parameter | Default | Description |
|---|---|---|
--service-discovery |
false |
Enable Kubernetes service discovery (also turns on IGW mode, see below) |
--selector |
- | Label selector for worker pods, as key=value pairs (required unless PD or EPD mode is on) |
--service-discovery-namespace |
(all namespaces) | Kubernetes namespace to watch |
--service-discovery-port |
80 |
Worker port for pods without a smg.ai/worker-ports annotation |
--model-id-from |
- | Override each worker's model ID from pod metadata: namespace, label:<key>, or annotation:<key> |
--model-alias |
- | Extra client-facing model name, alias=canonical, repeatable; applies to discovered workers too (details) |
Fleet Scale¶
Discovery submits its jobs to the gateway's control-plane job queue, which it shares with tokenizer, MCP, and WASM jobs.
| Parameter | Default | Description |
|---|---|---|
--job-queue-capacity |
1000 |
Maximum pending control-plane jobs. A reconcile pass waits while the queue is full, so set this at least as high as the number of discovered workers. |
--job-queue-concurrency |
200 |
Maximum control-plane jobs dispatched concurrently |
The cost of registering one worker does not grow with the fleet, and a job's in-flight status is kept until the job finishes, so a long registration wave on a new gateway replica is not submitted twice.
Worker Startup¶
| Parameter | Default | Description |
|---|---|---|
--worker-startup-delay |
0 |
Seconds to wait after a worker is submitted before its first startup probe |
--worker-startup-check-interval |
30 |
Seconds between startup probes while registration waits for the engine to answer |
--worker-startup-timeout-secs |
1800 |
How long registration waits for the engine before the job fails |
Discovery only submits workers for Ready pods, so give engine pods a readiness probe that passes once the model is loaded. If a pod reports Ready earlier, registration keeps probing at --worker-startup-check-interval until the engine answers or --worker-startup-timeout-secs elapses. A failed registration is submitted again on a later reconcile pass.
Multi-Port Pods¶
A pod that runs several engine servers, for example one per GPU, lists their ports in the smg.ai/worker-ports annotation as a comma-separated list. Each port becomes an independent worker with its own health checks, circuit breaker, and load tracking.
apiVersion: v1
kind: Pod
metadata:
name: sglang-multi-0
labels:
app: sglang-worker
annotations:
smg.ai/worker-ports: "8000,8001,8002,8003"smg.ai/worker-ports |
Workers registered for the pod |
|---|---|
| Absent | One, at --service-discovery-port |
"8000,8001,8002,8003" |
One per listed port, in order; duplicate ports are ignored |
| Invalid (an entry that is not a port from 1 to 65535) | One, at --service-discovery-port, with a warning that names the pod |
All workers of a pod follow the pod: they are registered when it becomes Ready and removed when it is deleted or turns unready. For per-port bootstrap ports and KV metadata in PD deployments, see PD Disaggregation Discovery.
Label Selectors¶
SMG uses Kubernetes label selectors to identify worker pods.
Simple Selector¶
Match pods with a single label:
smg --service-discovery --selector app=vllmMatches pods with label app=vllm.
Multiple Labels¶
Match pods that carry several labels by passing multiple key=value pairs:
smg --service-discovery --selector app=sglang environment=productionMatches pods with both app=sglang AND environment=production.
PD Disaggregation Discovery¶
For prefill-decode disaggregated deployments, use a separate selector for each worker role. SMG gives each pod the role of the first selector it matches (encode, then prefill, then decode) and ignores pods that match none.
Configuration¶
smg launch \
--service-discovery \
--pd-disaggregation \
--prefill-selector app=vllm role=prefill \
--decode-selector app=vllm role=decode \
--service-discovery-namespace inference \
--service-discovery-port 8000Service discovery turns on IGW mode automatically; the PD (or EPD) routing mode and the per-role policies stay in effect. /readiness reports ready only when at least one prefill worker and one decode worker are healthy (and an encode worker in EPD mode). See PD Disaggregation for how the legs are paired and dispatched.
Parameters¶
| Parameter | Default | Description |
|---|---|---|
--prefill-selector |
— | Label selector for prefill pods |
--decode-selector |
— | Label selector for decode pods |
--encode-selector |
— | Label selector for encode pods (EPD mode) |
--kv-connector-annotation |
smg.ai/kv-connector |
Pod annotation that names the vLLM KV connector |
--kv-engine-id-annotation |
smg.ai/kv-engine-id |
Pod annotation that lists the vLLM KV engine ids |
PD mode needs at least one of --prefill-selector and --decode-selector. EPD mode (--epd-disaggregation) needs all three role selectors. Annotation names must not be empty or padded with whitespace.
Pod Annotations¶
| Annotation | Pods | Value |
|---|---|---|
sglang.ai/bootstrap-port |
Prefill, encode | The worker's bootstrap port: SGLang and TokenSpeed --disaggregation-bootstrap-port, or vLLM Mooncake VLLM_MOONCAKE_BOOTSTRAP_PORT. A single value applies to every worker port of the pod; a comma-separated list needs one port per worker port, or it is ignored |
smg.ai/kv-connector |
vLLM prefill and decode | NixlConnector or MooncakeConnector, shared by every worker in the pod. Any other value is kept but handled as passthrough, with a warning |
smg.ai/kv-engine-id |
vLLM Mooncake prefill | The engine's kv_transfer_config.engine_id. One id for a single-port pod, or a comma-separated list of distinct ids in the order of smg.ai/worker-ports. A list of the wrong length, or with an empty or duplicate id, is ignored with a warning |
SGLang and TokenSpeed pods need only their role labels and, on prefill (and encode) pods, the bootstrap-port annotation. vLLM workers served over HTTP need smg.ai/kv-connector for a KV handoff, because the vLLM HTTP server does not report its connector. vLLM gRPC workers report their connector and engine id themselves, and the annotations override what they report. To pin a pairing protocol, set SMG_PAIRING_PROTOCOL in the engine container of gRPC workers; there is no annotation for it.
SMG reads the annotations when it registers a pod. After changing them, replace the pod so that SMG registers it again under its new UID.
Worker Labels¶
Label and annotate your pods. This example runs vLLM over HTTP with Mooncake:
# Prefill worker
apiVersion: v1
kind: Pod
metadata:
name: vllm-prefill-0
namespace: inference
labels:
app: vllm
role: prefill
annotations:
sglang.ai/bootstrap-port: "8998"
smg.ai/kv-connector: MooncakeConnector
smg.ai/kv-engine-id: prefill-0
spec:
containers:
- name: vllm
image: vllm/vllm-openai:latest
command: ["vllm", "serve", "meta-llama/Llama-3.1-8B-Instruct"]
args:
- --port=8000
- '--kv-transfer-config={"kv_connector":"MooncakeConnector","kv_role":"kv_producer","engine_id":"prefill-0"}'
env:
- name: VLLM_MOONCAKE_BOOTSTRAP_PORT
value: "8998"
---
# Decode worker
apiVersion: v1
kind: Pod
metadata:
name: vllm-decode-0
namespace: inference
labels:
app: vllm
role: decode
annotations:
smg.ai/kv-connector: MooncakeConnector
spec:
containers:
- name: vllm
image: vllm/vllm-openai:latest
command: ["vllm", "serve", "meta-llama/Llama-3.1-8B-Instruct"]
args:
- --port=8000
- '--kv-transfer-config={"kv_connector":"MooncakeConnector","kv_role":"kv_consumer"}'Required RBAC¶
SMG's informer lists and watches pods and reads no other Kubernetes resources. Grant get, list, and watch on pods in the watched namespace, the same rule the Helm chart creates. Mesh router discovery (--router-selector) watches pods too and needs no extra permissions.
Role¶
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
name: smg-discovery
namespace: inference
rules:
- apiGroups: [""]
resources: ["pods"]
verbs: ["get", "list", "watch"]RoleBinding¶
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
name: smg-discovery
namespace: inference
subjects:
- kind: ServiceAccount
name: smg
namespace: inference
roleRef:
kind: Role
name: smg-discovery
apiGroup: rbac.authorization.k8s.ioServiceAccount¶
apiVersion: v1
kind: ServiceAccount
metadata:
name: smg
namespace: inferenceCross-Namespace Discovery¶
Without --service-discovery-namespace, SMG watches pods in all namespaces. Grant the same rule cluster-wide with a ClusterRole:
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
name: smg-discovery
rules:
- apiGroups: [""]
resources: ["pods"]
verbs: ["get", "list", "watch"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRoleBinding
metadata:
name: smg-discovery
subjects:
- kind: ServiceAccount
name: smg
namespace: inference
roleRef:
kind: ClusterRole
name: smg-discovery
apiGroup: rbac.authorization.k8s.ioComplete Deployment Example¶
SMG Deployment¶
apiVersion: apps/v1
kind: Deployment
metadata:
name: smg
namespace: inference
spec:
replicas: 1
selector:
matchLabels:
app: smg
template:
metadata:
labels:
app: smg
spec:
serviceAccountName: smg
containers:
- name: smg
image: ghcr.io/smg-project/smg:latest
args:
- --service-discovery
- --selector=app=sglang-worker
- --service-discovery-namespace=inference
- --service-discovery-port=8000
- --policy=cache_aware
ports:
- containerPort: 30000
name: httpWorker StatefulSet¶
apiVersion: apps/v1
kind: StatefulSet
metadata:
name: sglang-worker
namespace: inference
spec:
serviceName: sglang-worker
replicas: 3
selector:
matchLabels:
app: sglang-worker
template:
metadata:
labels:
app: sglang-worker
spec:
containers:
- name: sglang
image: lmsysorg/sglang:latest
args:
- --model-path=meta-llama/Llama-3.1-8B-Instruct
- --port=8000
ports:
- containerPort: 8000
# SMG registers the pod only once it is Ready
readinessProbe:
httpGet:
path: /health
port: 8000
periodSeconds: 10Worker Lifecycle¶
Registration Flow¶
- Pod becomes Ready: The pod matches the selector, is
Running, and itsReadycondition isTrue. - Reconcile pass: SMG finds no registered worker for the pod's address and submits an
AddWorkerjob. - Startup probing: After
--worker-startup-delay, registration detects whether the worker speaks HTTP or gRPC, retrying every--worker-startup-check-intervalseconds until--worker-startup-timeout-secs. - Capability query: SMG reads the engine's metadata, for example SGLang's
/model_infoendpoint (falling back to the deprecated/get_model_infoif the new path returns 404). - Registration: The worker joins the registry with its
smg.ai/pod-nameandsmg.ai/pod-uidlabels, becomesReady, and background health checks start immediately.
Removal Flow¶
SMG removes a pod's workers when the pod:
- Starts terminating (its
deletionTimestampis set), so draining begins at the start of the pod's termination grace period rather than the end - Turns unready (its
Readycondition becomesFalseorUnknown) - No longer matches the selector
- Is replaced by a pod with a different UID at the same address
Each removal runs as a workflow:
- RemoveWorker job: The reconcile pass submits one job per address, pinned to each registration's revision. A worker that was replaced in the meantime is skipped and re-evaluated on the next pass.
- Drain:
Readyworkers move toDraining. They receive no new requests; requests already in flight continue. - Settle: SMG waits
--drain-settle-secs(default5). A worker'shealth.drain_settle_secsoverrides it, and a job that drains several workers waits for the longest window. Workers that were notReadyare not drained, and a job with noReadyworker skips the wait. - Remove: The workers leave the worker registry and the routing policies.
A removal job matches every registration at the address: the http:// or grpc:// worker and each DP rank registered there (<address>@<rank>).
When the pod becomes Ready again, the next pass registers it again. If readiness returns while the removal is still inside its settle window, re-registration can wait for the next periodic pass.
Worker States¶
| State | Description | Receives Traffic |
|---|---|---|
| Pending | Just registered, not yet verified | No |
| Ready | Verified and passing health checks | Yes |
| NotReady | Previously Ready, now failing health checks; still probed |
No |
| Failed | Sustained probe failure (about 12 minutes at the defaults); removed when auto-recovery is on, otherwise still probed | No |
| Draining | Being removed; kept for the settle window so in-flight requests can finish | No new requests |
See Health Checks for the thresholds behind each transition.
Recovery¶
With --service-discovery, worker auto-recovery (--remove-unhealthy-workers, alias --worker-auto-recovery) is on by default. A worker that reaches Failed is removed, and while its pod is still Ready the next reconcile pass registers it again; registration then waits for the engine to answer. To keep failed workers registered and let them rejoin in place instead, pass --remove-unhealthy-workers=false (Python launcher: --no-remove-unhealthy-workers). See Worker Auto-Recovery.
Mesh Router Discovery¶
--router-selector finds peer SMG gateway pods for an HA mesh. It runs as its own task, independent of worker discovery, but it is configured through the service discovery flags: it starts only when --service-discovery is set and the mesh is enabled (--enable-mesh). Router pods can publish their mesh port in the sglang.ai/mesh-port annotation; pods without it are assumed to use this gateway's --mesh-port. See High Availability.
Monitoring¶
Metrics¶
| Metric | Type | Labels | Description |
|---|---|---|---|
smg_discovery_workers_discovered |
Gauge | source |
Desired workers (Ready pods × ports) seen by the latest reconcile pass |
smg_discovery_registrations_total |
Counter | source, result |
AddWorker jobs submitted by discovery; result is success when the job was queued and failed when the queue refused it |
smg_discovery_deregistrations_total |
Counter | source, reason |
RemoveWorker jobs submitted by discovery; reason is reconciled |
smg_discovery_sync_duration_seconds |
Histogram | source |
Duration of reconcile passes that submitted jobs |
source is kubernetes. The counters record job submissions; whether a registration then succeeded shows up in GET /workers and in the worker health metrics.
Logs¶
# Enable discovery debug logging
RUST_LOG=info,smg::service_discovery=debug smg launch --service-discovery ...Example log output:
INFO Starting K8s service discovery | selector: 'app=sglang-worker'
INFO Starting K8s worker watcher | selector: 'app=sglang-worker'
INFO K8s worker store synced, reconciling on change and every 60s
INFO Reconciling workers: 2 to add, 0 to remove (2 desired)
INFO Registering worker 10.0.0.5:8000 (Regular) for pod sglang-worker-0
INFO Registering worker 10.0.0.6:8000 (Regular) for pod sglang-worker-1When a pod is deleted or turns unready:
INFO Reconciling workers: 0 to add, 1 to remove (1 desired)
INFO Removing worker 10.0.0.6:8000 (1 registration(s), pod <pod-uid>): pod unready, gone, terminating, or replaced
INFO Draining 1 worker(s) for 5s before removalTroubleshooting¶
| Symptom | Cause | Solution |
|---|---|---|
| No workers discovered | Selector does not match, or pods are not Ready | Check kubectl get pods -l <selector> and the pods' READY column |
Failed to start service discovery, then Continuing without service discovery |
No in-cluster or kubeconfig credentials | Run SMG with a ServiceAccount or a valid kubeconfig |
K8s worker watcher error (auto-retrying with backoff) |
RBAC denies the request, or the API server is unreachable | Apply the Role and RoleBinding and check API connectivity; the informer retries with backoff, and discovery converges once the watch recovers |
| One worker per pod instead of several | smg.ai/worker-ports is missing or invalid |
Check the pod's annotations and the gateway log for invalid smg.ai/worker-ports annotation |
| Workers drop out during a rollout before their pods are gone | Expected: terminating and unready pods are drained | Tune the pods' readiness probes and --drain-settle-secs |
| Workers registered but not receiving traffic | Health checks failing | Check the worker health endpoint and Health Checks |
| Reconcile passes stall on a large fleet | Control-plane job queue is full | Raise --job-queue-capacity |
Verify Discovery¶
# Reach the gateway from your machine
kubectl -n inference port-forward deployment/smg 30000:30000 &
# List discovered workers with their state and owning pod
curl -s http://localhost:30000/workers | jq '.workers[] | {url, status, pod: .labels["smg.ai/pod-name"]}'
# Check pod labels match selector
kubectl get pods -n inference -l app=sglang-worker
# Verify RBAC
kubectl auth can-i list pods -n inference --as=system:serviceaccount:inference:smg
kubectl auth can-i watch pods -n inference --as=system:serviceaccount:inference:smg