Routing
How llmux classifies a request, filters the model catalog and picks one model, and what it logs about the decision.
Classification
The rule-based classifier looks at the latest user message only
(classifier.user_messages, default 1), so the large static prefix of agent clients
(system prompt, tool schemas, history) does not skew it. It checks keyword groups in this
order and stops at the first match:
task_type |
Examples of keywords |
|---|---|
architecture |
architecture, design pattern, database schema, security, scalab, trade-off |
code_review |
refactor, bug, fix, test, rust, typescript, review this, function, implement |
summarize |
summarize, summary, tl;dr, explain, translate |
simple_text |
everything else |
The keyword lists also contain German terms. private_sensitive is never set by the
classifier; only the privacy scan sets it. With classifier.llm enabled, a small local
model returns the label instead, and any error or timeout falls back to the rules.
Selection
The selector applies its steps in this order:
- Quality floor. The task’s
min_tier, raised by a project’smin_tierif set. - Hard filters. Required capabilities (
toolswhen the request uses tools, plusjson_schemaorvisionwhen the request needs them, plus the task’srequire_capabilities),local_onlyfrom privacy or project, context window (estimated input plus expected output), provider enabled and key present, project allow and deny lists. - Budget pressure. Utilisation is today’s or this month’s real cost divided by
daily_max_usdormonthly_max_usd, whichever is higher. Everypressure_downgraderule withatat or below that value caps the tier; the lowest cap wins. If the cap falls below the quality floor, the request is markeddegraded. - Cheapest viable. Among the remaining models, the lowest estimated cost wins:
input tokens plus
expected_output_ratiotimes input tokens, at catalog prices. Under theinteractiveprofile, lowest expected p50 latency from the log comes first and cost breaks ties. On equal rank, local models come first whenprivacy.local_firstis set, then the higher tier. Withx-llmux-session, the model chosen earlier in the session is kept while it stays valid. - Budget gate. Only if even the cheapest candidate exceeds the remaining budget
does llmux answer
402.
With the example config’s budget rules, the allowed ceiling drops to tier 4 at 50 % utilisation, tier 2 at 80 % and tier 1 at 95 %.
Retry and fallback
| Provider answer | Reaction |
|---|---|
| 5xx, 429, network error | retry the same model with jittered exponential backoff |
| 401, 402, 403 | rotate to the next key if configured, then the next model |
| other 4xx | abort, no fallback |
retry.max_retries, backoff_initial_ms and backoff_max_ms tune the backoff.
x-llmux-no-fallback: true limits a request to its primary model.
Cache
The exact-match cache keys on the selected model and the normalised request. A hit
costs nothing and carries x-llmux-cache: hit. Conversations longer than
cache.max_conversation_messages are not cached. Expired rows are removed every
eviction_interval_seconds, and max_entries caps the table. The optional semantic
stage (cache.semantic) compares embeddings of the user text against earlier requests
to the same model after an exact miss (non-streaming requests only).
Policy result
Every request is logged with one policy label:
| Result | Meaning |
|---|---|
allowed |
selected dynamically and answered |
forced |
pinned by x-llmux-model or an alias |
cached |
answered from the cache |
fallback |
answered by a later candidate after the first one failed |
degraded |
answered below the task’s quality floor because of budget |
rejected |
refused, with the reason in error |
GET /api/stats/policy aggregates these labels; see the Stats API.