DecisionModelHub

Benchmarks you can inspect.

Task-specific results with frozen inputs, disclosed failures, and downloadable evidence.

FROZEN DATASET · V1

Support routing

100 held-out cases

Compare Jev, Clef, Clef Flash, OpenAI’s Decisions API and an OpenAI structured-output baseline on the same fictional support-routing cases and instructions.

  • Accuracy, equal-label macro-F1, invalid/error rates, and latency
  • Observed mismatches and provider failures before interpretation
  • Frozen JSON, CSV, manifest, model revisions, and run conditions

FROZEN DATASET · V1

Support urgency

600 held-out cases

Compare the same five models on a yes/no question, is this support message urgent, and see where each model’s probability puts the line.

  • Missed urgent messages and false alarms at the 0.5 threshold
  • The threshold each model needs, and what another model’s threshold costs
  • Every recorded answer and probability, with frozen JSON, CSV and manifest