DecisionModelHub

Cloudflare model profile

Clef Flash decision model: measured support-routing profile

Clef Flash is the lower listed input-price Cloudflare starting point and had a 195 ms median in this run. On the 100 frozen cases it returned valid output every time, and its accuracy was 87% (79–92), with the 95% interval in parentheses. The intervals describe uncertainty in each model's observed rate. This report does not include a paired statistical comparison of the differences between models.

Datasetsupport-routing v1
Cases100 held-out
Model revisionclef-flash
Run dateOctober 9, 2026

Measured on support routing

Evidence noteDecision quality, validity, and latency were independently measured from the frozen run. Estimated cost per 1,000 tickets multiplies provider-reported token counts from this run by provider list prices last compared on October 9, 2026; it is not a billed or measured cost. Values in parentheses are the 95% interval at n=100.

Correct routes
87/100
Independently measured
Accuracy
87% (79–92)
Independently measured
Operational success
87% (79–92)
Independently measured
Equal-label macro-F1
86% (79–92)
Independently measured
Valid responses
100/100
Independently measured
Median latency
195 ms
Independently measured
P95 latency
292 ms
Independently measured
Estimated cost per 1,000 tickets
≈ $0.027
Estimated · input tokens only

These numbers describe one balanced set of fictional support cases. They are not a general model ranking or a production service-level claim.

Clarification requests

Exploratory, n=20

Clef Flash sent 10 of 20 held-out needs_clarification cases back for clarification instead of routing them to a team: recall 50% (30–70), shown with its 95% Wilson interval. The denominator includes all 20 cases labeled needs_clarification, including attempts with no valid answer. This slice covers ambiguous or multi-intent support requests; it does not measure general out-of-domain detection.

With 20 cases, test clarification on your own traffic. The intervals describe uncertainty in each model's observed rate. This report does not include a paired statistical comparison of the differences between models. This is not a ranking.

See the per-label evidence

Where it fits

Teams that prioritize the lower provider-listed input price and will test whether its task quality meets their routing threshold.

Observed boundary

10 of its 13 observed mistakes were clarification cases routed to a specific team. Estimated at about $0.027 per 1,000 tickets from input tokens only; this is an estimate, not a measured cost.

Provider facts

Provider model ID
@cf/cloudflare/clef-flash
Documented context
65,536 tokens
Listed input price
$0.09 per 1M input tokens
Source checked
October 9, 2026

Cloudflare documents Clef Flash as a fast hosted 9B multimodal decision model for typed Noul, Choice, and Score questions, with text, structured-data, image, and video state.

Concrete observed failures

These are two model-specific mismatches from the frozen held-out results. The case text is fictional.

support-0057Clef Flash

I can sign in and view billing, but downloading invoice INV-905 returns a server error.

Expected
technical
Observed
billing
Status
succeeded
support-0096Clef Flash

For a future 40-seat subscription, can we pay annually by invoice?

Expected
sales
Observed
billing
Status
succeeded
See every observed failure

What this profile cannot answer

  • It does not estimate performance on your taxonomy, languages, message mix, or escalation policy.
  • Latency reflects a single disclosed run environment with concurrency one, no retries, and no cache.
  • Provider aliases may move to newer revisions; record the resolved revision in every evaluation.
  • Provider input price does not equal total application cost, and this run has no per-attempt billing attribution.