llmuxv0.1.0

Configuration

The sections of llmux.yaml, with the values from config/llmux.example.yaml.

The template is config/llmux.example.yaml; the demo config is config/llmux.demo.yaml. Both are embedded in the binary and written by llmux init and llmux init --demo. The config is validated at startup: a model on an unknown provider, an unknown capability name or an alias pointing to an unknown model stops the server before it binds. The snippets below are taken from the example config, with its comments left out.

server and auth

server:
  host: "0.0.0.0"
  port: 3456

auth:
  llmux_key: "mp_dev_changeme"

Clients send Authorization: Bearer <llmux_key> to /v1/.... An empty key disables auth; only do that on a machine nobody else can reach.

providers

providers:
  openrouter:
    enabled: true
    base_url: "https://openrouter.ai/api/v1"
    api_key_env: "OPENROUTER_API_KEY"
  openai:
    enabled: true
    base_url: "https://api.openai.com/v1"
    api_key_env: "OPENAI_API_KEY"
  ollama:
    enabled: true
    base_url: "http://localhost:11434/v1"
    api_key_env: null
    local: true
    strip_params: ["frequency_penalty", "presence_penalty", "logit_bias"]
  anthropic:
    enabled: false
    kind: anthropic
    base_url: "https://api.anthropic.com/v1"
    api_key_env: "ANTHROPIC_API_KEY"
    prompt_caching: true
    cache_billed_fraction: 0.1
Key Meaning
kind default OpenAI-compatible; anthropic translates to /v1/messages (non-streaming); echo is the demo provider
local the provider stays eligible for local_only requests
api_key_env environment variable holding the key; a provider whose key is missing is skipped
keys instead of api_key_env: several keys with env, weight and optional allow list; weighted-random choice, rotation on 401, 402, 403 and 429
strip_params request fields removed before sending (also per model)
prompt_caching discount the repeated prompt prefix in the routing estimate by cache_billed_fraction; billing is unchanged

models

The catalog the selector chooses from. Prices are USD per million tokens.

models:
  - { provider: ollama, model: "qwen2.5", tier: 1, context: 32000, supports_tools: false, input_per_mtok: 0.0, output_per_mtok: 0.0 }
  - { provider: openrouter, model: "google/gemini-flash-1.5", tier: 2, context: 1000000, supports_tools: true, capabilities: ["json_schema", "vision", "large_context"], input_per_mtok: 0.075, output_per_mtok: 0.30 }
  - { provider: openai, model: "gpt-4.1-mini", tier: 3, context: 128000, supports_tools: true, capabilities: ["json_schema", "vision"], input_per_mtok: 0.40, output_per_mtok: 1.60 }
  - { provider: openrouter, model: "anthropic/claude-3.5-sonnet", tier: 4, context: 200000, supports_tools: true, capabilities: ["json_schema", "vision"], input_per_mtok: 3.00, output_per_mtok: 15.00 }
  - { provider: openrouter, model: "anthropic/claude-3-opus", tier: 5, context: 200000, supports_tools: true, capabilities: ["vision"], input_per_mtok: 15.00, output_per_mtok: 75.00 }

Tier 1 means cheap or local, tier 5 top reasoning. Known capabilities: tools, json_schema, vision, streaming, streaming_usage, large_context, reasoning, strict_tool_schema. supports_tools: true is shorthand for tools. The example config lists eight models; five are shown here.

classification

classification:
  simple_text:       { min_tier: 1, expected_output_ratio: 0.5 }
  summarize:         { min_tier: 1, expected_output_ratio: 0.3 }
  code_review:       { min_tier: 3, expected_output_ratio: 1.0 }
  architecture:      { min_tier: 4, expected_output_ratio: 1.5 }
  private_sensitive: { min_tier: 1, local_only: true, expected_output_ratio: 1.0 }

Optional per task: require_tools, require_capabilities, local_only.

aliases and projects

aliases:
  fast: "google/gemini-flash-1.5"
  best: "anthropic/claude-3-opus"
  cheap: "meta-llama/llama-3.1-8b-instruct:free"

projects:
  client-api:
    local_only: true
  checkout:
    min_tier: 3
  docs:
    forbid_providers: ["openai"]

An alias in x-llmux-model or in the request’s model field forces its target. A project scope applies when the request sends x-llmux-project: <name>; it can set local_only, raise min_tier, and restrict providers with require_providers or forbid_providers.

budgets and routing

budgets:
  daily_max_usd: 2.00
  monthly_max_usd: 30.00
  pressure_downgrade:
    - { at: 0.50, max_tier: 4 }
    - { at: 0.80, max_tier: 2 }
    - { at: 0.95, max_tier: 1 }

routing:
  default_profile: balanced

Profiles: balanced (cheapest viable), interactive (lowest expected latency first), batch (currently the same as balanced). See Routing.

retry and cache

retry:
  max_retries: 2
  backoff_initial_ms: 500
  backoff_max_ms: 8000

cache:
  enabled: true
  ttl_seconds: 1800
  max_conversation_messages: 3
  eviction_interval_seconds: 300
  max_entries: 10000

The optional cache.semantic block (enabled, base_url, model, api_key_env, timeout_ms, threshold, default 0.85) adds an embedding-based second stage.

privacy and classifier

privacy:
  local_first: true
  block_cloud_patterns: [".env", "PRIVATE_KEY", "API_KEY", "BEGIN RSA", "password", "secret"]
  scan_system: false

classifier:
  user_messages: 1

With local_first: true, local models win ties on estimated cost. The privacy scan covers user and tool message content and tool or function schemas. scan_system: true also scans system and assistant content. The optional classifier.llm block (enabled, base_url, model, api_key_env, timeout_ms, default 1500) lets a small local model choose the task type.

Edit this page on GitHub · Docs for v0.1.0