Service Discovery¶
SMG can automatically discover workers in Kubernetes by watching pods with label selectors. Workers are registered when their pods become Ready and are drained and removed when pods scale down, terminate, or turn unready — no manual URL management needed.
Basic Setup¶
Enable service discovery with a label selector that matches your worker pods:
smg \
--service-discovery \
--selector app=sglang-worker \
--service-discovery-namespace inference \
--service-discovery-port 8000SMG keeps a local cache of the matching pods, registers a worker for each Ready pod, and removes workers whose pods go away. Every pod change triggers a reconcile pass, and a full pass also runs every 60 seconds, so changes missed while the watch was disconnected still converge.
Parameters¶
| Parameter | Default | Description |
|---|---|---|
--service-discovery |
false |
Enable Kubernetes service discovery |
--selector |
— | Label selector for worker pods (required unless PD or EPD mode is on) |
--service-discovery-namespace |
(all namespaces) | Kubernetes namespace to watch |
--service-discovery-port |
80 |
Worker port for pods without a smg.ai/worker-ports annotation |
Connection mode (HTTP vs gRPC) is probed automatically during worker registration, so no protocol flag is required — the first protocol that responds successfully is used, with HTTP taking priority when both succeed.
Service discovery also turns on IGW mode (--enable-igw) automatically.
Multiple Engines per Pod¶
If a pod runs several engine servers, list their ports in the smg.ai/worker-ports annotation. SMG registers one worker per port:
metadata:
labels:
app: sglang-worker
annotations:
smg.ai/worker-ports: "8000,8001,8002,8003"Pods without the annotation get a single worker at --service-discovery-port. If the annotation is invalid, SMG logs a warning and falls back to that port.
Label Selectors¶
Single Label¶
smg --service-discovery --selector app=vllmMultiple Labels¶
Pass multiple key=value pairs separated by spaces:
smg --service-discovery --selector app=sglang environment=productionMatches pods that carry every listed label.
PD Disaggregation Discovery¶
For prefill-decode deployments, use separate selectors:
smg launch \
--service-discovery \
--pd-disaggregation \
--prefill-selector app=vllm role=prefill \
--decode-selector app=vllm role=decode \
--service-discovery-namespace inference \
--service-discovery-port 8000Label your pods accordingly, and annotate prefill pods with their bootstrap port. vLLM pods also declare their KV connector, and Mooncake prefill pods their KV engine id:
# Prefill worker pod
metadata:
labels:
app: vllm
role: prefill
annotations:
sglang.ai/bootstrap-port: "8998" # SGLang, TokenSpeed, vLLM Mooncake
smg.ai/kv-connector: MooncakeConnector # vLLM: NixlConnector or MooncakeConnector
smg.ai/kv-engine-id: prefill-0 # vLLM Mooncake: engine_id from --kv-transfer-config
# Decode worker pod
metadata:
labels:
app: vllm
role: decode
annotations:
smg.ai/kv-connector: MooncakeConnector # vLLM only- vLLM's HTTP server does not report its KV connector, so HTTP vLLM workers need
smg.ai/kv-connectorfor a KV handoff. gRPC workers report it themselves. - On a pod that runs several workers (
smg.ai/worker-ports),smg.ai/kv-engine-idneeds a comma-separated list with one distinct id per port, in order.sglang.ai/bootstrap-porttakes either one port for all of them or one port per worker port. --kv-connector-annotationand--kv-engine-id-annotationrename the twosmg.ai/kv-*annotations.- SMG reads annotations when it registers a pod; replace the pod after changing them.
- Service discovery turns on IGW mode automatically; PD mode and the per-role policies are kept.
- For EPD, use
--epd-disaggregationand add an--encode-selector; encode pods take the bootstrap-port annotation too.
See PD Disaggregation Discovery for the annotation formats and a full pod example.
RBAC¶
SMG's informer lists and watches pods and reads no other resources. Apply these resources to your cluster:
apiVersion: v1
kind: ServiceAccount
metadata:
name: smg
namespace: inference
---
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
name: smg-discovery
namespace: inference
rules:
- apiGroups: [""]
resources: ["pods"]
verbs: ["get", "list", "watch"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
name: smg-discovery
namespace: inference
subjects:
- kind: ServiceAccount
name: smg
namespace: inference
roleRef:
kind: Role
name: smg-discovery
apiGroup: rbac.authorization.k8s.ioWithout --service-discovery-namespace, SMG watches all namespaces: use a ClusterRole and ClusterRoleBinding with the same rule instead.
Deployment Example¶
SMG Deployment¶
apiVersion: apps/v1
kind: Deployment
metadata:
name: smg
namespace: inference
spec:
replicas: 1
selector:
matchLabels:
app: smg
template:
metadata:
labels:
app: smg
spec:
serviceAccountName: smg
containers:
- name: smg
image: ghcr.io/smg-project/smg:latest
args:
- --service-discovery
- --selector=app=sglang-worker
- --service-discovery-namespace=inference
- --service-discovery-port=8000
- --policy=cache_aware
ports:
- containerPort: 30000
name: httpWorker StatefulSet¶
apiVersion: apps/v1
kind: StatefulSet
metadata:
name: sglang-worker
namespace: inference
spec:
serviceName: sglang-worker
replicas: 3
selector:
matchLabels:
app: sglang-worker
template:
metadata:
labels:
app: sglang-worker
spec:
containers:
- name: sglang
image: lmsysorg/sglang:latest
args:
- --model-path=meta-llama/Llama-3.1-8B-Instruct
- --port=8000
ports:
- containerPort: 8000
readinessProbe:
httpGet:
path: /health
port: 8000
periodSeconds: 10The readiness probe matters: SMG registers a pod only once its Ready condition is True, and drains its workers as soon as it turns unready.
Rollouts and Draining¶
When a pod is deleted, stops matching the selector, or fails its readiness probe, SMG stops routing new requests to its workers (they enter Draining; in-flight requests continue), then removes them after --drain-settle-secs (default 5). The pod is registered again when it becomes Ready.
With service discovery on, worker auto-recovery (--remove-unhealthy-workers) is also on by default: a worker that keeps failing health checks is removed, and discovery registers it again while its pod is still Ready. See Health Checks.
Verify¶
# If SMG runs in the cluster, forward its port first
kubectl -n inference port-forward deployment/smg 30000:30000 &
# Check discovered workers, their state, and the pod behind each one
curl -s http://localhost:30000/workers | jq '.workers[] | {url, status, pod: .labels["smg.ai/pod-name"]}'
# Check pod labels match selector
kubectl get pods -n inference -l app=sglang-worker
# Verify RBAC permissions
kubectl auth can-i list pods -n inference --as=system:serviceaccount:inference:smg
kubectl auth can-i watch pods -n inference --as=system:serviceaccount:inference:smgTroubleshooting¶
| Symptom | Cause | Solution |
|---|---|---|
| No workers discovered | Wrong selector, or pods not Ready | Verify labels match: kubectl get pods -l <selector>, and check the READY column |
| RBAC error | Missing permissions | Apply Role and RoleBinding above |
| Only one worker per multi-engine pod | smg.ai/worker-ports missing or invalid |
Check the pod annotation; SMG logs a warning for invalid values |
| Workers not ready | Health check failing | Check worker health endpoint |
| Workers not added or removed | Watch cannot reach the API server | Check Kubernetes API connectivity; discovery converges once the watch reconnects |
Next Steps¶
- Service Discovery Concepts — Reconcile loop, worker lifecycle, fleet-scale tuning, monitoring metrics
- Health Checks — Worker states and auto-recovery
- Load Balancing — Choose a routing policy for discovered workers
- PD Disaggregation — Full PD setup with SGLang and vLLM