Showcase
Each example shows real output produced by llmux itself, from files in the project repository. Input on the left, output on the right.
Simple question → tier 1
examples/simple-question.jsonDetail: Simple question → tier 1 →POST /v1/chat/completions
content-type: application/json
{
"model": "gpt-4o",
"messages": [
{ "role": "user", "content": "What is the capital of France?" }
]
}HTTP/1.1 200 OK
x-llmux-echo: 1
{
"choices": [
{
"finish_reason": "stop",
"index": 0,
"message": {
"content": "llmux demo · echo provider.\nRouted to echo/echo-nano (tier 1, task: simple_text). No external call was made — this is a synthetic response so you can see routing and the dashboard end-to-end.\n\nYour message: \"What is the capital of France?\"",
"role": "assistant"
}
}
],
"id": "chatcmpl-echo",
"model": "echo-nano",
"object": "chat.completion",
"usage": {
"completion_tokens": 61,
"prompt_tokens": 8,
"total_tokens": 69
}
}Code review → tier 3
examples/code-review.jsonDetail: Code review → tier 3 →POST /v1/chat/completions
content-type: application/json
{
"model": "gpt-4o",
"messages": [
{ "role": "user", "content": "Review this function for bugs: fn add(a: i32, b: i32) -> i32 { a - b }" }
]
}HTTP/1.1 200 OK
x-llmux-echo: 1
{
"choices": [
{
"finish_reason": "stop",
"index": 0,
"message": {
"content": "llmux demo · echo provider.\nRouted to echo/echo-base (tier 3, task: code_review). No external call was made — this is a synthetic response so you can see routing and the dashboard end-to-end.\n\nYour message: \"Review this function for bugs: fn add(a: i32, b: i32) -> i32 { a - b }\"",
"role": "assistant"
}
}
],
"id": "chatcmpl-echo",
"model": "echo-base",
"object": "chat.completion",
"usage": {
"completion_tokens": 71,
"prompt_tokens": 18,
"total_tokens": 89
}
}Architecture question → tier 4
examples/architecture.jsonDetail: Architecture question → tier 4 →POST /v1/chat/completions
content-type: application/json
{
"model": "gpt-4o",
"messages": [
{ "role": "user", "content": "Explain the architecture trade-offs between a monolith and microservices." }
]
}HTTP/1.1 200 OK
x-llmux-echo: 1
{
"choices": [
{
"finish_reason": "stop",
"index": 0,
"message": {
"content": "llmux demo · echo provider.\nRouted to echo/echo-pro (tier 4, task: architecture). No external call was made — this is a synthetic response so you can see routing and the dashboard end-to-end.\n\nYour message: \"Explain the architecture trade-offs between a monolith and microservices.\"",
"role": "assistant"
}
}
],
"id": "chatcmpl-echo",
"model": "echo-pro",
"object": "chat.completion",
"usage": {
"completion_tokens": 72,
"prompt_tokens": 19,
"total_tokens": 91
}
}Secret in the prompt → local only
examples/private-key.jsonDetail: Secret in the prompt → local only →POST /v1/chat/completions
content-type: application/json
{
"model": "gpt-4o",
"messages": [
{ "role": "user", "content": "Does this look valid? -----BEGIN RSA PRIVATE KEY----- MIIEow..." }
]
}HTTP/1.1 200 OK
x-llmux-echo: 1
{
"choices": [
{
"finish_reason": "stop",
"index": 0,
"message": {
"content": "llmux demo · echo provider.\nRouted to echo/echo-nano (tier 1, task: private_sensitive). No external call was made — this is a synthetic response so you can see routing and the dashboard end-to-end.\n\nYour message: \"Does this look valid? -----BEGIN RSA PRIVATE KEY----- MIIEow...\"",
"role": "assistant"
}
}
],
"id": "chatcmpl-echo",
"model": "echo-nano",
"object": "chat.completion",
"usage": {
"completion_tokens": 71,
"prompt_tokens": 16,
"total_tokens": 87
}
}Repeated request → cache hit
examples/cache-hit.jsonDetail: Repeated request → cache hit →POST /v1/chat/completions
content-type: application/json
{
"model": "gpt-4o",
"messages": [
{ "role": "user", "content": "What is the capital of France?" }
]
}HTTP/1.1 200 OK
x-llmux-cache: hit
{
"choices": [
{
"finish_reason": "stop",
"index": 0,
"message": {
"content": "llmux demo · echo provider.\nRouted to echo/echo-nano (tier 1, task: simple_text). No external call was made — this is a synthetic response so you can see routing and the dashboard end-to-end.\n\nYour message: \"What is the capital of France?\"",
"role": "assistant"
}
}
],
"id": "chatcmpl-echo",
"model": "echo-nano",
"object": "chat.completion",
"usage": {
"completion_tokens": 61,
"prompt_tokens": 8,
"total_tokens": 69
}
}Forced model via header
examples/forced-model.jsonDetail: Forced model via header →POST /v1/chat/completions
content-type: application/json
x-llmux-model: echo-ultra
{
"messages": [
{ "role": "user", "content": "Hello" }
]
}HTTP/1.1 200 OK
x-llmux-echo: 1
{
"choices": [
{
"finish_reason": "stop",
"index": 0,
"message": {
"content": "llmux demo · echo provider.\nRouted to echo/echo-ultra (tier 5, task: simple_text). No external call was made — this is a synthetic response so you can see routing and the dashboard end-to-end.\n\nYour message: \"Hello\"",
"role": "assistant"
}
}
],
"id": "chatcmpl-echo",
"model": "echo-ultra",
"object": "chat.completion",
"usage": {
"completion_tokens": 55,
"prompt_tokens": 2,
"total_tokens": 57
}
}