Reference

Reference documentation provides detailed specifications for SMG's APIs, CLI, configuration options, and metrics. Use these pages when you need precise information about specific features.


API Reference

OpenAI-Compatible API

Complete reference for the OpenAI-compatible endpoints including chat completions, completions, and models.

API Reference

Responses API

Agentic /v1/responses requests with tool calling and MCP, plus stored responses and conversations.

Responses

Anthropic Messages API

The /v1/messages and /v1/messages/count_tokens endpoints, with backend support per worker type.

Messages

Admin API

Request and response details for tokenizer, worker, cache, WASM, and model information endpoints.

Admin

Extension API

SMG-specific endpoints, from health probes and tokenization to worker management and other control-plane routes, with each route's auth tier.

Extensions


Configuration Reference

Configuration Reference

Complete CLI options, environment variables, and configuration for tuning SMG behavior.

Configuration

Tenant Rate Limiting

Per-tenant token and request budgets: flags, the YAML policy schema, and the 429 response.

Tenant Rate Limiting

Priority Scheduler

The priority-aware admission scheduler: the request header, response codes, and configuration flags.

Priority Scheduler

Internal MCP Servers

Static MCP servers whose tools the model uses but clients do not see in responses.

Internal MCP Servers


Observability Reference

Metrics Reference

Complete list of Prometheus metrics exposed by SMG for monitoring and alerting.

Metrics


Reference Description
CLI Options All command-line flags
Environment Variables Configurable environment variables
Chat Completions API /v1/chat/completions endpoint
HTTP Metrics HTTP request metrics
Worker Metrics Worker health and performance metrics