Architecture guide
Decision model or structured-output LLM for bounded routing?
In this comparison all four models returned a valid label for every case. Operational success was 98% (93–99) for the structured-output LLM and 92% (85–96), 92% (85–96) and 87% (79–92) for the three decision models. The intervals describe uncertainty in each model's observed rate. This report does not include a paired statistical comparison of the differences between models. In this run, the decision models answered in 138 to 271 ms at the median, against 1,325 ms.
Why we ran this
Decision models are built for bounded decisions such as routing a support message: one typed answer from a fixed set, returned quickly. A general LLM with structured output can be asked for the same answer. We wanted to see both approaches on the same cases, under the same instructions and scoring, so we ran three decision models and one structured-output LLM on 100 held-out support messages and published every answer.
The measured difference
Operational success (95% interval)
- Clef
- 92% (85–96)
- Clef Flash
- 87% (79–92)
- GPT-6 LunaStructured-output baseline
- 98% (93–99)
- Jev
- 92% (85–96)
Median response time
- Clef
- 271 ms
- Clef Flash
- 195 ms
- GPT-6 LunaStructured-output baseline
- 1,325 ms
- Jev
- 138 ms
Evidence noteQuality, validity, and latency are independently measured from the same frozen cases and instructions. Estimated cost multiplies provider-reported token counts by list prices last compared on October 9, 2026; it is not a billed or measured cost. Values in parentheses are the 95% interval at n=100.
| Approach | Model | Valid | Correct / attempted | Operational success | Median latency | Estimated cost / 1,000 tickets |
|---|---|---|---|---|---|---|
| Decision model | Jevjev-1.13.0 | 100/100 | 92/100 | 92% (85–96) | 138 ms | ≈ $0.020input tokens only |
| Decision model | Clefclef | 100/100 | 92/100 | 92% (85–96) | 271 ms | ≈ $0.071input tokens only |
| Decision model | Clef Flashclef-flash | 100/100 | 87/100 | 87% (79–92) | 195 ms | ≈ $0.027input tokens only |
| Structured-output LLM baseline | GPT-6 Lunagpt-6-luna | 100/100 | 98/100 | 98% (93–99) | 1325 ms | ≈ $0.040input and output tokens |
Observed quality and latency on this routing task
Every model returned a valid label for every case, so validity did not separate the approaches in this run. GPT-6 Luna had the highest point estimate for operational success, 98% (93–99), against 92% (85–96) for Jev, 92% (85–96) for Clef, and 87% (79–92) for Clef Flash. On the headline equal-label macro-F1, GPT-6 Luna scored 98% (95–100), compared with 86% (79–92) for Clef Flash. The intervals describe uncertainty in each model's observed rate. This report does not include a paired statistical comparison of the differences between models. The largest observed difference was on ambiguous messages: GPT-6 Luna recognised 18 of the 20 clarification cases, against 14, 14, and 10, which is a small sample.
Latency is where the approaches differ clearly. The decision models had median latencies of 138 to 271 ms. GPT-6 Luna had a median of 1,325 ms and a 95th percentile of 2,476 ms. Decision models also return a probability for each option, which a structured-output LLM does not.
The same model behind a decisions endpoint
OpenAI’s Decisions API asks gpt-6-luna, the model used as the structured-output baseline here, through a dedicated endpoint that takes a typed question and returns one answer with a probability for each option. On the same 100 cases and instructions it returned a valid label for 100 of 100, with operational success of 94% (88–97) and a median response time of 111 ms, against 98% (93–99) and 1,325 ms for the structured-output request. The intervals describe uncertainty in each model's observed rate. This report does not include a paired statistical comparison of the differences between models.
It recognised 20 of the 20 clarification cases, and 6 of its 6 mistakes were requests with a responsible team that it sent for clarification.
Give a reasoning model room to answer
Reasoning tokens count against a structured-output LLM's output cap. The routing answer here is a short JSON object, yet GPT-6 Luna's responses used a median of 35 output tokens and as many as 101. A cap sized for the answer alone ends such a response before the label, and the result looks like invalid output.
If you evaluate a structured-output LLM, budget output tokens for reasoning as well as the answer, and record why a response failed, not only that it failed. A truncated response and a wrong label call for different fixes.
Choose a decision model when
- Your action space is a stable choice, score, or probability.
- Your code benefits from typed answers and per-option probabilities.
- Response time is an acceptance criterion for the routing step.
- You can build a focused labeled evaluation for the exact decision.
Keep a general LLM in the test when
- Quality on ambiguous or mixed-intent messages matters more than response time.
- The task still requires explanation, transformation, or broad language generation.
- The label policy changes frequently or depends on long, nuanced instructions.
- You need one model to perform both the decision and a separately evaluated generative step.
- Your production harness has schema validation, an output budget that covers reasoning, and fallback and retry controls.
A practical evaluation sequence
- Freeze the decision.Write allowed labels, boundaries, and a safe ambiguity path.
- Label representative cases.Include difficult and mixed-intent examples, not only clean demos.
- Run the same harness.Use the same cases, instructions, retry policy, and denominators.
- Count operational failures.Track invalid, truncated, refusal, and timeout outcomes separately, with the cause of each.
- Choose against thresholds.Decide using your latency, quality, privacy, and cost requirements, not a headline ranking.
Limits of this evidence
This comparison is a single run over balanced fictional support messages, with concurrency one, no retries, and no cache. It does not measure production label prevalence, multilingual behavior, long-context degradation, or total application cost. The observed ordering may change on another task or a future model revision.