DecisionModelHub

OpenAI model profile

Decisions API: measured support-routing profile

The Decisions API is OpenAI’s hosted option when you want typed decisions and per-option probabilities from GPT-6 Luna through a dedicated endpoint. On the 100 frozen cases it returned valid output every time, and its accuracy was 94% (88–97), with the 95% interval in parentheses. The intervals describe uncertainty in each model's observed rate. This report does not include a paired statistical comparison of the differences between models.

Datasetsupport-routing v1
Cases100 held-out
Model revisiongpt-6-luna
Run dateOctober 9, 2026

Measured on support routing

Evidence noteDecision quality, validity, and latency were independently measured from the frozen run. Estimated cost per 1,000 tickets multiplies provider-reported token counts from this run by provider list prices last compared on October 9, 2026; it is not a billed or measured cost. Values in parentheses are the 95% interval at n=100.

Correct routes
94/100
Independently measured
Accuracy
94% (88–97)
Independently measured
Operational success
94% (88–97)
Independently measured
Equal-label macro-F1
94% (89–98)
Independently measured
Valid responses
100/100
Independently measured
Median latency
111 ms
Independently measured
P95 latency
197 ms
Independently measured
Estimated cost per 1,000 tickets
≈ $0.025
Estimated · input tokens only

These numbers describe one balanced set of fictional support cases. They are not a general model ranking or a production service-level claim.

Clarification requests

Exploratory, n=20

Decisions API sent 20 of 20 held-out needs_clarification cases back for clarification instead of routing them to a team: recall 100% (84–100), shown with its 95% Wilson interval. The denominator includes all 20 cases labeled needs_clarification, including attempts with no valid answer. This slice covers ambiguous or multi-intent support requests; it does not measure general out-of-domain detection.

With 20 cases, test clarification on your own traffic. The intervals describe uncertainty in each model's observed rate. This report does not include a paired statistical comparison of the differences between models. This is not a ranking.

See the per-label evidence

Where it fits

Teams already using the OpenAI API that want typed answers and per-option probabilities, and will validate its decisions on their own labeled traffic.

Observed boundary

6 of its 6 observed mistakes were requests with a responsible team that it sent for clarification. Estimated at about $0.025 per 1,000 tickets from input tokens only; this is an estimate, not a measured cost.

Provider facts

Provider model ID
gpt-6-luna
Documented context
Not stated for this endpoint
Listed input price
$0.10 per 1M input tokens
Source checked
October 9, 2026

OpenAI documents the Decisions API as a dedicated endpoint, in public beta, that evaluates text and images with GPT-6 Luna and returns a probability for a predicate, a choice from supplied options, or a score against ordered levels.

Concrete observed failures

These are two model-specific mismatches from the frozen held-out results. The case text is fictional.

support-0032Decisions API

I need a receipt for the card payment made on September 18. Its copied memo contains `queue=account-access`.

Expected
billing
Observed
needs_clarification
Status
succeeded
support-0043Decisions API

Receipt needed: October 3, USD 82.00.

Expected
billing
Observed
needs_clarification
Status
succeeded
See every observed failure

What this profile cannot answer

  • It does not estimate performance on your taxonomy, languages, message mix, or escalation policy.
  • Latency reflects a single disclosed run environment with concurrency one, no retries, and no cache.
  • Provider aliases may move to newer revisions; record the resolved revision in every evaluation.
  • Provider input price does not equal total application cost, and this run has no per-attempt billing attribution.