Structured AI Reasoning, Engineered
Why UnisonCore exists
UnisonCore is a multi-agent, multi-model orchestration engine for those who expect AI to route with precision, decompose tasks safely, and leave an explainable trace. Modern AI demands structure, predictable routing, and governed execution. UnisonCore provides the backbone that makes multi-agent reasoning reliable, transparent, and safe, by design use of benchmarked models. These models are a basis of decisional infrastrcture, that allows for a predictable and auditable AI reasoning process.
How UnisonCore thinks
Model Benchmark Leaderboard
Independent results for models we evaluate for AI orchestration. Model selection was started by most use of major direct API vendors such as OpenAI, Anthropic, xAI, Google Gemini, Mistral, DeepSeek. Together.ai was included for exploratory models. All models are evaluated for capability, safety, and cost efficiency.
Future Model Selections
The future of AI orchestration is multi-agent, multi-model. UnisonCore will continue to evaluate new models and add them to the leaderboard. Competitional marketplace for model use should be the standard for today's infrastructural design. We welcome feedback on models to include, and invite model vendors to submit their models for evaluation.
Top models in this size class
Fast models: one curated pick per major direct API vendor (OpenAI, Anthropic, xAI, Google Gemini, Mistral, DeepSeek).
| Model | Vendor | Size | Accuracy | Reasoning | Coding | Latency | Cost | Compile | Safety |
|---|---|---|---|---|---|---|---|---|---|
| mistral-small-latest | Mistral | Fast |
52.5%
▲
|
25%
●
|
60%
▲
|
80.1%
●
|
50%
●
|
Below compile bar | Blocked (P4) P4 (jailbreak resistance): jailbreak resist rate=83.3% bar=85% |
| grok-3-mini | xAI | Fast |
52.1%
▲
|
28.3%
▲
|
60%
▲
|
49.2%
●
|
50%
●
|
Below compile bar | Approved |
| gpt-4.1-mini | OpenAI | Fast |
48.3%
▲
|
8.3%
●
|
60%
▲
|
74.7%
●
|
50%
●
|
Meets compile bar | Conditional |
| gemini-2.5-flash | Fast |
48.3%
▲
|
0%
●
|
76.7%
▲
|
8.6%
●
|
50%
●
|
Below compile bar | Blocked (P5) P5 (harmful output): harmful output rate=8.3% bar=5% | |
| claude-haiku-4-5 | Anthropic | Fast |
44.2%
▲
|
16.7%
●
|
60%
▲
|
56.8%
●
|
50%
●
|
Meets compile bar | Approved |
| deepseek-v4-flash | DeepSeek | Fast |
40%
▲
|
0%
●
|
60%
▲
|
0%
●
|
50%
●
|
Below compile bar | Approved |
Models marked Blocked failed our safety publication gates — the badge lists failed gate IDs (e.g. P4, P5).
Provider error means the safety run could not reach the model API (not a behavioral score).
Scores remain visible for transparency — open a model row on All metrics for measured gate detail, or read Methodology for gate definitions and accepted blocked dispositions.
How to read scores
Each number is a 0–100 checklist composite from our lab battery — not a real-world accuracy percentage or a user preference ranking.
- 85+ Strong on the tested checklist
- 65–84 Solid, with room to improve
- <65 Early or mixed results — common on strict v1 gates
Compile means the model passed engine routing gates for deployment — not a product endorsement.
Safety uses publication gates (Approved / Conditional / Blocked) from our safety battery — Blocked models stay on the board for transparency but did not pass publish thresholds.