OpenAI model profile
Decisions API: measured support-routing profile
The Decisions API is OpenAI’s hosted option when you want typed decisions and per-option probabilities from GPT-6 Luna through a dedicated endpoint. On the 100 frozen cases it returned valid output every time, and its accuracy was 94% (88–97), with the 95% interval in parentheses. The intervals describe uncertainty in each model's observed rate. This report does not include a paired statistical comparison of the differences between models.
Measured on support routing
Evidence noteDecision quality, validity, and latency were independently measured from the frozen run. Estimated cost per 1,000 tickets multiplies provider-reported token counts from this run by provider list prices last compared on October 9, 2026; it is not a billed or measured cost. Values in parentheses are the 95% interval at n=100.
- Correct routes
- 94/100 Independently measured
- Accuracy
- 94% (88–97) Independently measured
- Operational success
- 94% (88–97) Independently measured
- Equal-label macro-F1
- 94% (89–98) Independently measured
- Valid responses
- 100/100 Independently measured
- Median latency
- 111 ms Independently measured
- P95 latency
- 197 ms Independently measured
- Estimated cost per 1,000 tickets
- ≈ $0.025 Estimated · input tokens only
These numbers describe one balanced set of fictional support cases. They are not a general model ranking or a production service-level claim.
Clarification requests
Exploratory, n=20Decisions API sent 20 of 20 held-out needs_clarification cases back for clarification instead of routing them to a team: recall 100% (84–100), shown with its 95% Wilson interval. The denominator includes all 20 cases labeled needs_clarification, including attempts with no valid answer. This slice covers ambiguous or multi-intent support requests; it does not measure general out-of-domain detection.
With 20 cases, test clarification on your own traffic. The intervals describe uncertainty in each model's observed rate. This report does not include a paired statistical comparison of the differences between models. This is not a ranking.
See the per-label evidenceWhere it fits
Teams already using the OpenAI API that want typed answers and per-option probabilities, and will validate its decisions on their own labeled traffic.
Observed boundary
6 of its 6 observed mistakes were requests with a responsible team that it sent for clarification. Estimated at about $0.025 per 1,000 tickets from input tokens only; this is an estimate, not a measured cost.
Provider facts
- Provider model ID
gpt-6-luna- Documented context
- Not stated for this endpoint
- Listed input price
- $0.10 per 1M input tokens
- Source checked
- October 9, 2026
OpenAI documents the Decisions API as a dedicated endpoint, in public beta, that evaluates text and images with GPT-6 Luna and returns a probability for a predicate, a choice from supplied options, or a score against ordered levels.
Concrete observed failures
These are two model-specific mismatches from the frozen held-out results. The case text is fictional.
support-0032Decisions APII need a receipt for the card payment made on September 18. Its copied memo contains `queue=account-access`.
- Expected
- billing
- Observed
- needs_clarification
- Status
- succeeded
support-0043Decisions APIReceipt needed: October 3, USD 82.00.
- Expected
- billing
- Observed
- needs_clarification
- Status
- succeeded
What this profile cannot answer
- It does not estimate performance on your taxonomy, languages, message mix, or escalation policy.
- Latency reflects a single disclosed run environment with concurrency one, no retries, and no cache.
- Provider aliases may move to newer revisions; record the resolved revision in every evaluation.
- Provider input price does not equal total application cost, and this run has no per-attempt billing attribution.