How we benchmark models

What appears on the public leaderboard, how cohorts differ, and what we do not claim.

Reference board (default)

The default public board is a reference cohort — curated models we profile on native vendor APIs, not an exhaustive vendor catalog. For each major direct API vendor (OpenAI, Anthropic, xAI, Google Gemini, Mistral, DeepSeek), we profile one Fast and one Flagship model on that vendor's native API. All capability-tested P1 models appear on the board; models that fail safety publication gates show a red Blocked badge rather than being hidden.

We also publish a small Frontier band (above 70B parameters) for large open models that passed the same capability and safety bar. Today these Frontier models are evaluated via Together.ai because they are not yet represented as direct P1 picks — latency and cost on the site reflect Together endpoints, not native vendor routing.

Together Catalog (exploratory)

The Together Catalog tab shows models from a separate exploratory allowlist hosted on Together.ai. These rows are capability-tested for comparison, but safety testing is not required for this cohort. They are not mixed into the reference board — including Frontier models, which use a higher publication bar.

Safety publication gates

Safety-tested reference models must pass automated publication gates on our fixed safety pack. Rows that fail are labeled Blocked (P…) on the leaderboard — we keep them visible with scores for transparency.

Accepted blocked dispositions (reference board)

The following models remain on the board with a red Blocked badge. We accept these outcomes as stable behavioral gaps — not missing data — and do not hide the rows.

Model Band Failed gate Why (summary)
claude-sonnet-4-6 Flagship P5 Harmful-output probes (fiction/compliance scenarios) — stable across reruns.
gemini-2.5-flash Fast P5 Harmful fiction probe — one persistent failure from certification threshold.
mistral-small-latest Fast P4 Jailbreak fiction laundering probes — stable 10/12 resist (83.3% vs 85% minimum).
Cogito v2 671B Frontier P2 Policy under-refusal on multiple probes — real behavioral gap, not provider transport.

Reading the numbers

What we do not claim

View Together Catalog