TypedLM

CASOON Open Source

LLM calls with a type signature.

TypedLM turns a language model call into a Rust function: a struct goes in, a validated Rust type comes out. The same contract lets you measure the program on a labelled dataset.

cargo add typedlm --git https://github.com/casoon/typedlm
MITMSRV 1.85not on crates.io yet
Terminal: with TYPEDLM_MODEL=qwen3:32b, cargo run --example classification prints 'escalate: TicketClassification { urgency: High, category: Technical }' and then 'TicketClassification { urgency: High, category: Billing } — 1 attempt(s), 318 tokens, model qwen3:32b'.
output strategies, chosen per provider
4
direct dependencies without default features
4
async runtime required by the core
0
confidence interval on every evaluation score
95 %

What it does

  1. Contracts instead of prompts

    Derive TypedLm on the input struct and name an ordinary output type. Doc comments become the instructions, the type becomes the JSON Schema.

  2. Every answer validated

    Extraction tolerates code fences and trailing commas, the schema check lists every violation with its path, your own rules run last. Invalid answers go back to the model with that list.

  3. Measured, not eyeballed

    Datasets in JSON Lines with partial labels, exact-match and field accuracy, and a report with confidence interval, repair and failure rates, latency and tokens.

  4. Any OpenAI-compatible endpoint

    OpenAI, local models via Ollama, or a router in front of other vendors. Schemas are rewritten for each provider; retries for 429 and 5xx are built in.

The contractRust
/// Classify a customer support ticket by urgency and category.
#[derive(TypedLm, Serialize)]
#[lm(output = TicketClassification)]
struct ClassifyTicket {
    text: String,
}

#[derive(Debug, Serialize, Deserialize, JsonSchema)]
struct TicketClassification {
    urgency: Urgency,   // enum Urgency { Low, Medium, High }
    category: Category, // enum Category { Billing, Technical, Shipping, Other }
}

Generated at build time

All examples →
Terminal: cargo run --example evaluation prints the report for ClassifyTicket with model qwen3:32b on 8 examples: exact match 87.5 % (95 % CI 52.9 % – 97.8 %), category 85.7 % (6/7), urgency 100.0 % (5/5), valid 100.0 %, repaired 0.0 %, failed 0.0 %, latency p50 96248 ms, p95 121426 ms, tokens 1099 in / 1925 out.

Recorded with qwen3:32b on a local Ollama (cargo run --example evaluation). Eight examples give an interval from 53 to 98 percent — the report says so.

examples/recorded/evaluation.txt

Quickstart

From contract to running call. The full walkthrough is in the documentation.

  1. Add the crate from GitHub.
  2. Describe the call as an input struct with #[derive(TypedLm)] and an output type.
  3. Bind it to a provider with Program::new and call run.
main.rsRust
use typedlm::http::OpenAiCompatible;
use typedlm::prelude::*;

let classify = Program::<ClassifyTicket, _>::new(OpenAiCompatible::ollama("qwen3:32b"))
    .temperature(0.0);

let ticket = classify.run("Production database is down since 9:00").await?;
match ticket.urgency {
    Urgency::High => escalate(ticket).await?,
    _ => queue(ticket).await?,
}