llmuxv0.1.0

Showcase

Each example shows real output produced by llmux itself, from files in the project repository. Input on the left, output on the right.

Simple question → tier 1

examples/simple-question.jsonDetail: Simple question → tier 1 →
POST /v1/chat/completions
content-type: application/json

{
  "model": "gpt-4o",
  "messages": [
    { "role": "user", "content": "What is the capital of France?" }
  ]
}
HTTP/1.1 200 OK
x-llmux-echo: 1

{
  "choices": [
    {
      "finish_reason": "stop",
      "index": 0,
      "message": {
        "content": "llmux demo · echo provider.\nRouted to echo/echo-nano (tier 1, task: simple_text). No external call was made — this is a synthetic response so you can see routing and the dashboard end-to-end.\n\nYour message: \"What is the capital of France?\"",
        "role": "assistant"
      }
    }
  ],
  "id": "chatcmpl-echo",
  "model": "echo-nano",
  "object": "chat.completion",
  "usage": {
    "completion_tokens": 61,
    "prompt_tokens": 8,
    "total_tokens": 69
  }
}
  • simple_text
  • tier 1

Code review → tier 3

examples/code-review.jsonDetail: Code review → tier 3 →
POST /v1/chat/completions
content-type: application/json

{
  "model": "gpt-4o",
  "messages": [
    { "role": "user", "content": "Review this function for bugs: fn add(a: i32, b: i32) -> i32 { a - b }" }
  ]
}
HTTP/1.1 200 OK
x-llmux-echo: 1

{
  "choices": [
    {
      "finish_reason": "stop",
      "index": 0,
      "message": {
        "content": "llmux demo · echo provider.\nRouted to echo/echo-base (tier 3, task: code_review). No external call was made — this is a synthetic response so you can see routing and the dashboard end-to-end.\n\nYour message: \"Review this function for bugs: fn add(a: i32, b: i32) -> i32 { a - b }\"",
        "role": "assistant"
      }
    }
  ],
  "id": "chatcmpl-echo",
  "model": "echo-base",
  "object": "chat.completion",
  "usage": {
    "completion_tokens": 71,
    "prompt_tokens": 18,
    "total_tokens": 89
  }
}
  • code_review
  • tier 3

Architecture question → tier 4

examples/architecture.jsonDetail: Architecture question → tier 4 →
POST /v1/chat/completions
content-type: application/json

{
  "model": "gpt-4o",
  "messages": [
    { "role": "user", "content": "Explain the architecture trade-offs between a monolith and microservices." }
  ]
}
HTTP/1.1 200 OK
x-llmux-echo: 1

{
  "choices": [
    {
      "finish_reason": "stop",
      "index": 0,
      "message": {
        "content": "llmux demo · echo provider.\nRouted to echo/echo-pro (tier 4, task: architecture). No external call was made — this is a synthetic response so you can see routing and the dashboard end-to-end.\n\nYour message: \"Explain the architecture trade-offs between a monolith and microservices.\"",
        "role": "assistant"
      }
    }
  ],
  "id": "chatcmpl-echo",
  "model": "echo-pro",
  "object": "chat.completion",
  "usage": {
    "completion_tokens": 72,
    "prompt_tokens": 19,
    "total_tokens": 91
  }
}
  • architecture
  • tier 4

Secret in the prompt → local only

examples/private-key.jsonDetail: Secret in the prompt → local only →
POST /v1/chat/completions
content-type: application/json

{
  "model": "gpt-4o",
  "messages": [
    { "role": "user", "content": "Does this look valid? -----BEGIN RSA PRIVATE KEY----- MIIEow..." }
  ]
}
HTTP/1.1 200 OK
x-llmux-echo: 1

{
  "choices": [
    {
      "finish_reason": "stop",
      "index": 0,
      "message": {
        "content": "llmux demo · echo provider.\nRouted to echo/echo-nano (tier 1, task: private_sensitive). No external call was made — this is a synthetic response so you can see routing and the dashboard end-to-end.\n\nYour message: \"Does this look valid? -----BEGIN RSA PRIVATE KEY----- MIIEow...\"",
        "role": "assistant"
      }
    }
  ],
  "id": "chatcmpl-echo",
  "model": "echo-nano",
  "object": "chat.completion",
  "usage": {
    "completion_tokens": 71,
    "prompt_tokens": 16,
    "total_tokens": 87
  }
}
  • privacy
  • private_sensitive
  • local_only

Repeated request → cache hit

examples/cache-hit.jsonDetail: Repeated request → cache hit →
POST /v1/chat/completions
content-type: application/json

{
  "model": "gpt-4o",
  "messages": [
    { "role": "user", "content": "What is the capital of France?" }
  ]
}
HTTP/1.1 200 OK
x-llmux-cache: hit

{
  "choices": [
    {
      "finish_reason": "stop",
      "index": 0,
      "message": {
        "content": "llmux demo · echo provider.\nRouted to echo/echo-nano (tier 1, task: simple_text). No external call was made — this is a synthetic response so you can see routing and the dashboard end-to-end.\n\nYour message: \"What is the capital of France?\"",
        "role": "assistant"
      }
    }
  ],
  "id": "chatcmpl-echo",
  "model": "echo-nano",
  "object": "chat.completion",
  "usage": {
    "completion_tokens": 61,
    "prompt_tokens": 8,
    "total_tokens": 69
  }
}
  • cache
  • x-llmux-cache

Forced model via header

examples/forced-model.jsonDetail: Forced model via header →
POST /v1/chat/completions
content-type: application/json
x-llmux-model: echo-ultra

{
  "messages": [
    { "role": "user", "content": "Hello" }
  ]
}
HTTP/1.1 200 OK
x-llmux-echo: 1

{
  "choices": [
    {
      "finish_reason": "stop",
      "index": 0,
      "message": {
        "content": "llmux demo · echo provider.\nRouted to echo/echo-ultra (tier 5, task: simple_text). No external call was made — this is a synthetic response so you can see routing and the dashboard end-to-end.\n\nYour message: \"Hello\"",
        "role": "assistant"
      }
    }
  ],
  "id": "chatcmpl-echo",
  "model": "echo-ultra",
  "object": "chat.completion",
  "usage": {
    "completion_tokens": 55,
    "prompt_tokens": 2,
    "total_tokens": 57
  }
}
  • override
  • x-llmux-model