DecisionModelHub

How the support-routing benchmark was run

The cases, reviews, run conditions and scoring behind the published results, and how to evaluate on your own tickets.

How this benchmark was run

The task

Each case is one English support message. A model must choose exactly one of 5 labels: billing, technical, account, sales or needs_clarification. A written label policy decides the expected answer. It routes the immediate blocker before the eventual goal, so an upgrade page that crashes is a technical issue, not a sales one.

Cases and authoring

125 fictional support messages were written for this benchmark, 25 for each label. The scenarios were drafted with AI assistance under the label policy. No private customer conversations, playground traffic or unlicensed third-party datasets were used.

Review

A separate AI reviewer checked the drafts, and its findings were resolved first. Two independent human reviewers then reviewed all 125 cases against the label policy and agreed with the labels and rationales. The AI review is disclosed but does not count as one of the human reviews. Reviewer identities are kept private.

Development and held-out cases

25 development cases could guide the instructions. The 100 held-out cases were kept out of that process: 20 per label, made up of 40 clear, 30 boundary, 20 ambiguous and 10 adversarial messages. Each held-out case was compared with the development, catalog, pilot and policy examples by exact and semantic matching. Held-out cases that resembled one were rewritten and their labels rechecked. The held-out cases are now published, so they are no longer a private test set.

Run conditions

Every model answered the same 100 held-out cases with the same instructions on October 9, 2026. The runner sent one request at a time, with no retries and caching disabled. Each provider alias was resolved to a model revision at run time, and the revision is recorded with the results.

Scoring

An answer is correct only when it is a valid label that matches the expected one. Operational success keeps every attempted case in the denominator, so a failed or invalid response counts as a miss. Accuracy and operational success are shown with a 95% Wilson interval. The headline metric is macro-F1 with equal weight for each label, shown with a 95% bootstrap interval from 2,000 resamples of the cases. The intervals describe each model on its own; the report does not include a paired statistical comparison between models.

Latency and cost

Latency is the complete request time observed in that run, not a provider guarantee. Cost per 1,000 tickets is an estimate: the token counts each provider reported, multiplied by list prices checked on October 9, 2026. It is not a billed amount.

Check it yourself

Every held-out message, expected label, model answer and response time is in the case explorer. The results and the run manifest, with the dataset hash and model revisions, can be downloaded under CC BY 4.0.

Evaluate on your own tickets

One benchmark on fictional messages cannot tell you how a model handles your traffic. The steps below build the evidence that can.

Examples are a starting point.

Each sample case has an authored expected answer and rationale. Example mode shows these annotations for the selected models without calling their APIs. These examples explain the schema; they do not measure model performance.

No dataset yet? Build the first 30–50 cases.

Begin with realistic fictional tickets written with a developer or support operator. Cover ordinary requests, ambiguous messages, mixed intent, typos, and instructions embedded inside the customer text. Record an expected answer and a short reason for each.

  • Keep early examples for schema development.
  • Set aside a separate test set before tuning instructions.
  • Have a second person review labels and resolve disagreements.
  • Add redacted, authorized real cases as you learn about the workflow.
  • Keep uncertain cases and mistakes in the dataset.

Compare the same task under the same conditions.

Use the same state, allowed answers, and instructions for each model. Record the exact model version, dataset version, run date, provider errors, and timeout policy. Repeat calls to study variability. Report a denominator and sample size alongside accuracy.

Separate offline evaluation from production review.

An expected-label mismatch is useful while evaluating a labeled test set, but the expected answer does not exist for a new production message. Production review must use observable signals such as a fallback route, a missing or invalid output, disagreement between models, or a sensitive destination team.

Keep confidence and probability separate.

An option’s probability is not automatically the provider’s confidence field. Preserve the provider’s original values. For a binary probability question, avoid inventing a separate confidence value. A confidence threshold should be chosen from observed errors on your own test set.

Measure real latency and cost.

Time the complete server-to-provider round trip. Report medians and tail latency over repeated requests. Use returned token usage and the dated provider price to estimate cost, and clearly label missing usage. Input price alone does not prove which model is cheaper for your task.

Let mistakes inform the product.

Evaluate incorrect routes, abstentions, and review load together. A small improvement matters only if the workflow benefits. Avoid publishing a universal winner from a small support dataset.