How we benchmark models
What appears on the public leaderboard, how cohorts differ, and what we do not claim.
Reference board (default)
The default public board is a reference cohort — curated models we profile on native vendor APIs, not an exhaustive vendor catalog. For each major direct API vendor (OpenAI, Anthropic, xAI, Google Gemini, Mistral, DeepSeek), we profile one Fast and one Flagship model on that vendor's native API. All capability-tested P1 models appear on the board; models that fail safety publication gates show a red Blocked badge rather than being hidden.
We also publish a small Frontier band (above 70B parameters) for large open models that passed the same capability and safety bar. Today these Frontier models are evaluated via Together.ai because they are not yet represented as direct P1 picks — latency and cost on the site reflect Together endpoints, not native vendor routing.
- Capability testing uses our standard benchmark battery (v2 pack).
- Safety testing is required for every P1 row on the board; rows that fail publication gates are labeled Blocked.
- Safety-certified (green badge) means the model passed all safety publication gates — use that subset for endorsement claims.
- Size bands on the site: Fast (direct API), Flagship (direct API), and Frontier (Together.ai today).
- Compare scores within a band; do not treat Frontier latency/cost as directly comparable to Fast or Flagship.
Together Catalog (exploratory)
The Together Catalog tab shows models from a separate exploratory allowlist hosted on Together.ai. These rows are capability-tested for comparison, but safety testing is not required for this cohort. They are not mixed into the reference board — including Frontier models, which use a higher publication bar.
- Flat list sorted by vendor and model name (no size-band selector).
- Marked as exploratory on the leaderboard.
Safety publication gates
Safety-tested reference models must pass automated publication gates on our fixed safety pack. Rows that fail are labeled Blocked (P…) on the leaderboard — we keep them visible with scores for transparency.
- P2 — policy under-refusal: model complied when it should have refused harmful policy prompts.
- P4 — jailbreak resistance: model complied with jailbreak framing prompts.
- P5 — harmful output: model produced disallowed harmful content on direct harmful-output probes.
- Provider error — safety API could not reach the model (not a behavioral score).
Accepted blocked dispositions (reference board)
The following models remain on the board with a red Blocked badge. We accept these outcomes as stable behavioral gaps — not missing data — and do not hide the rows.
| Model | Band | Failed gate | Why (summary) |
|---|---|---|---|
| claude-sonnet-4-6 | Flagship | P5 | Harmful-output probes (fiction/compliance scenarios) — stable across reruns. |
| gemini-2.5-flash | Fast | P5 | Harmful fiction probe — one persistent failure from certification threshold. |
| mistral-small-latest | Fast | P4 | Jailbreak fiction laundering probes — stable 10/12 resist (83.3% vs 85% minimum). |
| Cogito v2 671B | Frontier | P2 | Policy under-refusal on multiple probes — real behavioral gap, not provider transport. |
Reading the numbers
- Scores are 0–100 checklist composites from fixed lab scenarios — not real-world task accuracy.
- Most models score below 65 on early v1 gates; that reflects strict tests, not necessarily poor products.
- Compile badges mean the model passed engine routing gates — not an endorsement.
What we do not claim
- We do not mirror every SKU from every AI vendor.
- Scores are checklist-based lab results, not human preference rankings.
- We do not guarantee performance in your product or workload.
- Hosting channel (direct API vs Together.ai) affects latency and cost; we label cohorts accordingly.