Task-specific results with frozen inputs, disclosed failures, and downloadable evidence.
FROZEN DATASET · V1
Support routing
100 held-out cases
Compare Jev, Clef, Clef Flash, OpenAI’s Decisions API and an OpenAI structured-output baseline on the same fictional support-routing cases and instructions.
Accuracy, equal-label macro-F1, invalid/error rates, and latency
Observed mismatches and provider failures before interpretation
Frozen JSON, CSV, manifest, model revisions, and run conditions