UnisonCore Metrics Dashboard
These metrics provide insight into routing stability, decomposition quality, suppression behavior, and depth reliability for the selected cohort and band.
Search and compare exploratory models on Together.ai.
How to read scores
Each number is a 0–100 checklist composite from our lab battery — not a real-world accuracy percentage or a user preference ranking.
- 85+ Strong on the tested checklist
- 65–84 Solid, with room to improve
- <65 Early or mixed results — common on strict v1 gates
Compile means the model passed engine routing gates for deployment — not a product endorsement.
Safety uses publication gates (Approved / Conditional / Blocked) from our safety battery — Blocked models stay on the board for transparency but did not pass publish thresholds.
Exploratory models on Together.ai — capability-tested for comparison; safety testing is not required for this cohort.
improving · regressing · stable
Together Catalog models
Metrics Last Update:
| Identity | Routing Metrics | Model Performance Metrics | Decomposition Metrics | Suppression Metrics | Safety Metrics | Depth Stability Metrics | Status | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Model | Vendor | Deploy | Accuracy | Reasoning | Coding | Latency | Cost | Reliability | Slop | Cap. safety | Jailbreak | PII | Bias | Stability | Badges |
| Model identity and vendor. | Routing readiness index. Higher is better. | Benchmark accuracy. Higher is better. | Reasoning score. Higher is better. | Coding score. Higher is better. | Latency score. Higher is better. | Cost efficiency. Higher is better. | Structured output reliability. Higher is better. | Slop / hallucination rate. Lower is better. | Capability safety. Higher is better. | Jailbreak resistance. Higher is better. | PII leakage rate. Lower is better. | Bias disparity rate. Lower is better. | Output stability. Higher is better. | Compile and safety publication badges. | |
| Llama 3 70B Chat Exploratory · safety not required | Meta |
24.8%
●
|
0%
●
|
0%
●
|
0%
●
|
99.1%
●
|
—
●
|
0%
●
|
100%
●
|
0%
●
|
—
●
|
Not tested
●
|
Not tested
●
|
0%
●
|
Below compile bar Not required |
| Llama 3 8B Chat Exploratory · safety not required | Meta |
24%
●
|
0%
●
|
0%
●
|
0%
●
|
95.9%
●
|
—
●
|
0%
●
|
100%
●
|
0%
●
|
—
●
|
Not tested
●
|
Not tested
●
|
0%
●
|
Below compile bar Not required |
| Llama 3.3 70B Instruct Turbo Exploratory · safety not required | Meta |
54.4%
●
|
61.3%
▲
|
25%
●
|
60%
▲
|
63.8%
●
|
50%
●
|
80%
●
|
0%
●
|
83.3%
●
|
—
●
|
Not tested
●
|
Not tested
●
|
100%
●
|
Below compile bar Not required |
| Llama 3.3 70B Instruct Turbo (Together) Exploratory · safety not required | Meta |
55.4%
●
|
61.3%
▲
|
25%
●
|
60%
▲
|
68.8%
●
|
50%
●
|
80%
●
|
0%
●
|
83.3%
●
|
—
●
|
Not tested
●
|
Not tested
●
|
100%
●
|
Below compile bar Not required |
| Mistral Small 24B Instruct Exploratory · safety not required | Mistral |
24.9%
●
|
0%
●
|
0%
●
|
0%
●
|
99.6%
●
|
—
●
|
0%
●
|
100%
●
|
0%
●
|
—
●
|
Not tested
●
|
0%
●
|
0%
●
|
Below compile bar Not required |
| Qwen 2.5 7B Instruct Turbo Exploratory · safety not required | Qwen |
51.6%
●
|
58.8%
▲
|
25%
●
|
85%
▲
|
49.5%
●
|
50%
●
|
80%
●
|
0%
●
|
83.3%
●
|
—
●
|
Not tested
●
|
Not tested
●
|
100%
●
|
Below compile bar Not required |
Metric Guide
Trend icons compare medium vs easy standard scores when pack history exists; otherwise ● stable.
improving · regressing · stable