DecisionModelHub

Architecture guide

Decision model or structured-output LLM for bounded routing?

In this comparison all four models returned a valid label for every case. Operational success was 98% (93–99) for the structured-output LLM and 92% (85–96), 92% (85–96) and 87% (79–92) for the three decision models. The intervals describe uncertainty in each model's observed rate. This report does not include a paired statistical comparison of the differences between models. In this run, the decision models answered in 138 to 271 ms at the median, against 1,325 ms.

Task5-label support routing
Cases100 held-out
Models4 measured
Run dateOctober 9, 2026

Why we ran this

Decision models are built for bounded decisions such as routing a support message: one typed answer from a fixed set, returned quickly. A general LLM with structured output can be asked for the same answer. We wanted to see both approaches on the same cases, under the same instructions and scoring, so we ran three decision models and one structured-output LLM on 100 held-out support messages and published every answer.

The measured difference

Operational success (95% interval)

Clef
92% (85–96)
Clef Flash
87% (79–92)
GPT-6 LunaStructured-output baseline
98% (93–99)
Jev
92% (85–96)

Median response time

Clef
271 ms
Clef Flash
195 ms
GPT-6 LunaStructured-output baseline
1,325 ms
Jev
138 ms
Operational success with its 95% interval, and median response time, for each model on the same 100 held-out cases. Run October 9, 2026.

Evidence noteQuality, validity, and latency are independently measured from the same frozen cases and instructions. Estimated cost multiplies provider-reported token counts by list prices last compared on October 9, 2026; it is not a billed or measured cost. Values in parentheses are the 95% interval at n=100.

ApproachModelValidCorrect / attemptedOperational successMedian latencyEstimated cost / 1,000 tickets
Decision modelJevjev-1.13.0100/10092/10092% (85–96)138 ms≈ $0.020input tokens only
Decision modelClefclef100/10092/10092% (85–96)271 ms≈ $0.071input tokens only
Decision modelClef Flashclef-flash100/10087/10087% (79–92)195 ms≈ $0.027input tokens only
Structured-output LLM baselineGPT-6 Lunagpt-6-luna100/10098/10098% (93–99)1325 ms≈ $0.040input and output tokens

Observed quality and latency on this routing task

Every model returned a valid label for every case, so validity did not separate the approaches in this run. GPT-6 Luna had the highest point estimate for operational success, 98% (93–99), against 92% (85–96) for Jev, 92% (85–96) for Clef, and 87% (79–92) for Clef Flash. On the headline equal-label macro-F1, GPT-6 Luna scored 98% (95–100), compared with 86% (79–92) for Clef Flash. The intervals describe uncertainty in each model's observed rate. This report does not include a paired statistical comparison of the differences between models. The largest observed difference was on ambiguous messages: GPT-6 Luna recognised 18 of the 20 clarification cases, against 14, 14, and 10, which is a small sample.

Latency is where the approaches differ clearly. The decision models had median latencies of 138 to 271 ms. GPT-6 Luna had a median of 1,325 ms and a 95th percentile of 2,476 ms. Decision models also return a probability for each option, which a structured-output LLM does not.

The same model behind a decisions endpoint

OpenAI’s Decisions API asks gpt-6-luna, the model used as the structured-output baseline here, through a dedicated endpoint that takes a typed question and returns one answer with a probability for each option. On the same 100 cases and instructions it returned a valid label for 100 of 100, with operational success of 94% (88–97) and a median response time of 111 ms, against 98% (93–99) and 1,325 ms for the structured-output request. The intervals describe uncertainty in each model's observed rate. This report does not include a paired statistical comparison of the differences between models.

It recognised 20 of the 20 clarification cases, and 6 of its 6 mistakes were requests with a responsible team that it sent for clarification.

Give a reasoning model room to answer

Reasoning tokens count against a structured-output LLM's output cap. The routing answer here is a short JSON object, yet GPT-6 Luna's responses used a median of 35 output tokens and as many as 101. A cap sized for the answer alone ends such a response before the label, and the result looks like invalid output.

If you evaluate a structured-output LLM, budget output tokens for reasoning as well as the answer, and record why a response failed, not only that it failed. A truncated response and a wrong label call for different fixes.

Choose a decision model when

  • Your action space is a stable choice, score, or probability.
  • Your code benefits from typed answers and per-option probabilities.
  • Response time is an acceptance criterion for the routing step.
  • You can build a focused labeled evaluation for the exact decision.

Keep a general LLM in the test when

  • Quality on ambiguous or mixed-intent messages matters more than response time.
  • The task still requires explanation, transformation, or broad language generation.
  • The label policy changes frequently or depends on long, nuanced instructions.
  • You need one model to perform both the decision and a separately evaluated generative step.
  • Your production harness has schema validation, an output budget that covers reasoning, and fallback and retry controls.

A practical evaluation sequence

  1. Freeze the decision.Write allowed labels, boundaries, and a safe ambiguity path.
  2. Label representative cases.Include difficult and mixed-intent examples, not only clean demos.
  3. Run the same harness.Use the same cases, instructions, retry policy, and denominators.
  4. Count operational failures.Track invalid, truncated, refusal, and timeout outcomes separately, with the cause of each.
  5. Choose against thresholds.Decide using your latency, quality, privacy, and cost requirements, not a headline ranking.

Limits of this evidence

This comparison is a single run over balanced fictional support messages, with concurrency one, no retries, and no cache. It does not measure production label prevalence, multilingual behavior, long-context degradation, or total application cost. The observed ordering may change on another task or a future model revision.