Cloudflare model profile
Clef decision model: measured support-routing profile
Clef is a practical hosted option when you want Cloudflare Workers AI and the larger Clef model. On the 100 frozen cases it returned valid output every time, and its accuracy was 92% (85–96), with the 95% interval in parentheses. The intervals describe uncertainty in each model's observed rate. This report does not include a paired statistical comparison of the differences between models.
Measured on support routing
Evidence noteDecision quality, validity, and latency were independently measured from the frozen run. Estimated cost per 1,000 tickets multiplies provider-reported token counts from this run by provider list prices last compared on October 9, 2026; it is not a billed or measured cost. Values in parentheses are the 95% interval at n=100.
- Correct routes
- 92/100 Independently measured
- Accuracy
- 92% (85–96) Independently measured
- Operational success
- 92% (85–96) Independently measured
- Equal-label macro-F1
- 92% (86–97) Independently measured
- Valid responses
- 100/100 Independently measured
- Median latency
- 271 ms Independently measured
- P95 latency
- 434 ms Independently measured
- Estimated cost per 1,000 tickets
- ≈ $0.071 Estimated · input tokens only
These numbers describe one balanced set of fictional support cases. They are not a general model ranking or a production service-level claim.
Clarification requests
Exploratory, n=20Clef sent 14 of 20 held-out needs_clarification cases back for clarification instead of routing them to a team: recall 70% (48–85), shown with its 95% Wilson interval. The denominator includes all 20 cases labeled needs_clarification, including attempts with no valid answer. This slice covers ambiguous or multi-intent support requests; it does not measure general out-of-domain detection.
With 20 cases, test clarification on your own traffic. The intervals describe uncertainty in each model's observed rate. This report does not include a paired statistical comparison of the differences between models. This is not a ranking.
See the per-label evidenceWhere it fits
Teams already using Cloudflare or comparing the larger Clef model with Clef Flash under one API shape.
Observed boundary
6 of its 8 observed mistakes were clarification cases routed to a specific team. Estimated at about $0.071 per 1,000 tickets from input tokens only; this is an estimate, not a measured cost.
Provider facts
- Provider model ID
@cf/cloudflare/clef- Documented context
- 65,536 tokens
- Listed input price
- $0.24 per 1M input tokens
- Source checked
- October 9, 2026
Cloudflare documents Clef as a hosted 27B multimodal decision model for typed Noul, Choice, and Score questions, with text, structured-data, image, and video state.
Concrete observed failures
These are two model-specific mismatches from the frozen held-out results. The case text is fictional.
support-0071ClefPlease deactivate the contractor's workspace account at the end of today.
- Expected
- account
- Observed
- needs_clarification
- Status
- succeeded
support-0100Clef50 seats, EU, annual—quote?
- Expected
- sales
- Observed
- needs_clarification
- Status
- succeeded
What this profile cannot answer
- It does not estimate performance on your taxonomy, languages, message mix, or escalation policy.
- Latency reflects a single disclosed run environment with concurrency one, no retries, and no cache.
- Provider aliases may move to newer revisions; record the resolved revision in every evaluation.
- Provider input price does not equal total application cost, and this run has no per-attempt billing attribution.