UnisonCore Metrics Dashboard
These metrics provide insight into routing stability, decomposition quality, suppression behavior, and depth reliability for the selected cohort and band.
Search, sort, and compare reference board models in the selected size class (direct vendor APIs).
How to read scores
Each number is a 0–100 checklist composite from our lab battery — not a real-world accuracy percentage or a user preference ranking.
- 85+ Strong on the tested checklist
- 65–84 Solid, with room to improve
- <65 Early or mixed results — common on strict v1 gates
Compile means the model passed engine routing gates for deployment — not a product endorsement.
Safety uses publication gates (Approved / Conditional / Blocked) from our safety battery — Blocked models stay on the board for transparency but did not pass publish thresholds.
Fast models: one curated pick per major direct API vendor (OpenAI, Anthropic, xAI, Google Gemini, Mistral, DeepSeek).
improving · regressing · stable
Models in this size class
Metrics Last Update:
| Identity | Routing Metrics | Model Performance Metrics | Decomposition Metrics | Suppression Metrics | Safety Metrics | Depth Stability Metrics | Status | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Model | Vendor | Deploy | Accuracy | Reasoning | Coding | Latency | Cost | Reliability | Slop | Cap. safety | Jailbreak | PII | Bias | Stability | Badges |
| Model identity and vendor. | Routing readiness index. Higher is better. | Benchmark accuracy. Higher is better. | Reasoning score. Higher is better. | Coding score. Higher is better. | Latency score. Higher is better. | Cost efficiency. Higher is better. | Structured output reliability. Higher is better. | Slop / hallucination rate. Lower is better. | Capability safety. Higher is better. | Jailbreak resistance. Higher is better. | PII leakage rate. Lower is better. | Bias disparity rate. Lower is better. | Output stability. Higher is better. | Compile and safety publication badges. | |
| gemini-2.5-flash |
43.4%
●
|
48.3%
▲
|
0%
●
|
76.7%
▲
|
8.6%
●
|
50%
●
|
80%
●
|
0%
●
|
83.3%
●
|
91.7%
●
|
0%
●
|
0%
●
|
100%
●
|
Below compile bar Blocked (P5) P5 (harmful output): harmful output rate=8.3% bar=5% | |
Metric Guide
Trend icons compare medium vs easy standard scores when pack history exists; otherwise ● stable.
improving · regressing · stable