Updates
New benchmarks, results and guides, newest first.
Subscribe to the feed Add it to a feed reader to hear when new results are published.
Support-urgency benchmark: five models on 600 held-out cases
A second benchmark asks a yes/no question: is this support message urgent? It reports each model’s probability on 600 held-out messages, the threshold each model needs, what another model’s threshold costs, errors by kind of message, calibration and response times, with every per-case answer to download.
Support-routing benchmark: OpenAI Decisions API results
OpenAI’s Decisions API measured on the same 100 held-out cases and instructions as the other models. Accuracy with 95% intervals, response times, calibration and every per-case answer with its probabilities are in the results, with a measured profile for the model.
OpenAI Decisions API in the catalog and playground
OpenAI’s Decisions API, in public beta on GPT-6 Luna, is listed with its input price and decision types. The playground builds its request for all three recipes, with integration code to run in your own application.
Support-routing benchmark: probability calibration
Per-case probability distributions for Jev, Clef and Clef Flash on the 100 held-out cases, with expected calibration error, Brier scores and reliability bins. The results with probabilities are available to download.
Guide: what is a decision model?
A plain definition of a decision model, the three question types it answers, how it differs from a general LLM, a trained classifier and a rules engine, and why it is not the same thing as DMN.
Support-routing benchmark: four models on 100 held-out cases
Jev, Clef, Clef Flash and a structured-output LLM baseline measured on the same frozen cases, with 95% intervals, per-label results, response times and estimated cost. Every case and result is available to download under CC BY 4.0.
Case explorer: each model’s recorded answer on every held-out case
Search and filter the 100 held-out support messages and see what each model answered, whether it was right and how long it took.
Guide: decision model or structured-output LLM for bounded routing?
What the benchmark shows about choosing between a purpose-built decision model and a general LLM with structured output: where they differ, where they do not, and how to run your own comparison.
Stay informed or ask for help
Checking sign-in status…