DecisionModelHub

Cloudflare model profile

Clef decision model: measured support-routing profile

Clef is a practical hosted option when you want Cloudflare Workers AI and the larger Clef model. On the 100 frozen cases it returned valid output every time, and its accuracy was 92% (85–96), with the 95% interval in parentheses. The intervals describe uncertainty in each model's observed rate. This report does not include a paired statistical comparison of the differences between models.

Datasetsupport-routing v1
Cases100 held-out
Model revisionclef
Run dateOctober 9, 2026

Measured on support routing

Evidence noteDecision quality, validity, and latency were independently measured from the frozen run. Estimated cost per 1,000 tickets multiplies provider-reported token counts from this run by provider list prices last compared on October 9, 2026; it is not a billed or measured cost. Values in parentheses are the 95% interval at n=100.

Correct routes
92/100
Independently measured
Accuracy
92% (85–96)
Independently measured
Operational success
92% (85–96)
Independently measured
Equal-label macro-F1
92% (86–97)
Independently measured
Valid responses
100/100
Independently measured
Median latency
271 ms
Independently measured
P95 latency
434 ms
Independently measured
Estimated cost per 1,000 tickets
≈ $0.071
Estimated · input tokens only

These numbers describe one balanced set of fictional support cases. They are not a general model ranking or a production service-level claim.

Clarification requests

Exploratory, n=20

Clef sent 14 of 20 held-out needs_clarification cases back for clarification instead of routing them to a team: recall 70% (48–85), shown with its 95% Wilson interval. The denominator includes all 20 cases labeled needs_clarification, including attempts with no valid answer. This slice covers ambiguous or multi-intent support requests; it does not measure general out-of-domain detection.

With 20 cases, test clarification on your own traffic. The intervals describe uncertainty in each model's observed rate. This report does not include a paired statistical comparison of the differences between models. This is not a ranking.

See the per-label evidence

Where it fits

Teams already using Cloudflare or comparing the larger Clef model with Clef Flash under one API shape.

Observed boundary

6 of its 8 observed mistakes were clarification cases routed to a specific team. Estimated at about $0.071 per 1,000 tickets from input tokens only; this is an estimate, not a measured cost.

Provider facts

Provider model ID
@cf/cloudflare/clef
Documented context
65,536 tokens
Listed input price
$0.24 per 1M input tokens
Source checked
October 9, 2026

Cloudflare documents Clef as a hosted 27B multimodal decision model for typed Noul, Choice, and Score questions, with text, structured-data, image, and video state.

Concrete observed failures

These are two model-specific mismatches from the frozen held-out results. The case text is fictional.

support-0071Clef

Please deactivate the contractor's workspace account at the end of today.

Expected
account
Observed
needs_clarification
Status
succeeded
support-0100Clef

50 seats, EU, annual—quote?

Expected
sales
Observed
needs_clarification
Status
succeeded
See every observed failure

What this profile cannot answer

  • It does not estimate performance on your taxonomy, languages, message mix, or escalation policy.
  • Latency reflects a single disclosed run environment with concurrency one, no retries, and no cache.
  • Provider aliases may move to newer revisions; record the resolved revision in every evaluation.
  • Provider input price does not equal total application cost, and this run has no per-attempt billing attribution.