Reference¶
Reference documentation provides detailed specifications for SMG's APIs, CLI, configuration options, and metrics. Use these pages when you need precise information about specific features.
API Reference¶
OpenAI-Compatible API
Complete reference for the OpenAI-compatible endpoints including chat completions, completions, and models.
Responses API
Agentic /v1/responses requests with tool calling and MCP, plus stored responses and conversations.
Anthropic Messages API
The /v1/messages and /v1/messages/count_tokens endpoints, with backend support per worker type.
Admin API
Request and response details for tokenizer, worker, cache, WASM, and model information endpoints.
Extension API
SMG-specific endpoints, from health probes and tokenization to worker management and other control-plane routes, with each route's auth tier.
Configuration Reference¶
Configuration Reference
Complete CLI options, environment variables, and configuration for tuning SMG behavior.
Tenant Rate Limiting
Per-tenant token and request budgets: flags, the YAML policy schema, and the 429 response.
Priority Scheduler
The priority-aware admission scheduler: the request header, response codes, and configuration flags.
Internal MCP Servers
Static MCP servers whose tools the model uses but clients do not see in responses.
Observability Reference¶
Metrics Reference
Complete list of Prometheus metrics exposed by SMG for monitoring and alerting.
Quick Links¶
| Reference | Description |
|---|---|
| CLI Options | All command-line flags |
| Environment Variables | Configurable environment variables |
| Chat Completions API | /v1/chat/completions endpoint |
| HTTP Metrics | HTTP request metrics |
| Worker Metrics | Worker health and performance metrics |