Cloudflare decision models
Clef vs Clef Flash for support routing
In this 100-case support-routing benchmark, Clef and Clef Flash scored 92% (86–97) and 86% (79–92) on the headline equal-label macro-F1, with 95% intervals in parentheses. The intervals describe uncertainty in each model's observed rate. This report does not include a paired statistical comparison of the differences between models. Clef Flash had lower median and p95 latency. Both returned valid output for all 100 cases.
Side-by-side evidence
Evidence noteQuality and latency values are independently measured. Listed prices are provider list prices last compared with the providers' pages on October 9, 2026. Estimated cost per 1,000 tickets multiplies provider-reported token counts from the October 9, 2026 run by those prices; it is not a billed or measured cost. Values in parentheses are the 95% interval at n=100.
| Measure | Clefclef | Clef Flashclef-flash | Evidence |
|---|---|---|---|
| Correct routes | 92/100 | 87/100 | Independently measured |
| Accuracy | 92% (85–96) | 87% (79–92) | Independently measured |
| Operational success | 92% (85–96) | 87% (79–92) | Independently measured |
| Equal-label macro-F1 | 92% (86–97) | 86% (79–92) | Independently measured |
| Median latency | 271 ms | 195 ms | Independently measured |
| P95 latency | 434 ms | 292 ms | Independently measured |
| Invalid responses | 0% | 0% | Independently measured |
| Listed input price | $0.24 per 1M input tokens | $0.09 per 1M input tokens | Provider reported |
| Estimated cost per 1,000 tickets | ≈ $0.071 (input tokens only) | ≈ $0.027 (input tokens only) | Estimated |
When Clef may fit
Prefer Clef when your own evaluation shows a quality difference that justifies its higher listed input price; validate that tradeoff on your own traffic.
Read the Clef profileWhen Clef Flash may fit
Prefer Clef Flash when the lower listed input price matters and your own evaluation shows its decisions clear your acceptance threshold.
Read the Clef Flash profileRepresentative mistakes
These are independently measured mismatches on fictional held-out cases, not provider-reported examples.
support-0071ClefPlease deactivate the contractor's workspace account at the end of today.
- Expected
- account
- Observed
- needs_clarification
- Status
- succeeded
support-0100Clef50 seats, EU, annual—quote?
- Expected
- sales
- Observed
- needs_clarification
- Status
- succeeded
support-0057Clef FlashI can sign in and view billing, but downloading invoice INV-905 returns a server error.
- Expected
- technical
- Observed
- billing
- Status
- succeeded
support-0096Clef FlashFor a future 40-seat subscription, can we pay annually by invoice?
- Expected
- sales
- Observed
- billing
- Status
- succeeded
Interpret the difference carefully
The result covers one balanced support-routing dataset; neither latency nor quality is a service-level or cross-task guarantee.
- All cases are fictional and evenly distributed across five labels.
- Production traffic may have a different label mix. The intervals describe uncertainty in each model's observed rate. This report does not include a paired statistical comparison of the differences between models.
- Latency is an observed result, not a provider guarantee.
- No benchmark-attributed dollar cost is available for either model.