A rigorous multi-round benchmark evaluating large language models across Bias, Doctrine, Coding, Tool Use, and Reasoning domains
Six rounds of structured evaluation testing how leading LLMs handle bias, follow doctrine, write code, use tools, and reason through complex problems.
Scored 2222/2500 (44.4/50 avg), winning overall with consistent top-3 finishes across all domains. Most well-rounded cloud model tested.
Appeared in 4 rounds and placed in the top 3 every time: R1 #2 (202.0), R4 #1 (208.0), R6 #1 (214.3). The most reliable local model across changing conditions.
R8 cloud models showed a striking split: analysis prompts scored 40-48/50 while code generation prompts scored 29-39/50. Think mode excels at reasoning-through-prose but over-engineers actual code. deepseek-v4.1-flash (45.1/50 Coding) was the only model to largely escape this pattern.
Critical insights drawn from cross-round analysis.
Scored 2222/2500 (44.4/50 avg) in combined v1+v2 scoring, the highest total ever recorded. Most consistent model across all domains with no catastrophic weaknesses.
Top 3 in every round it appeared (R1 #2, R4 #1, R6 #1). No other local model shows this level of reliability across conditions.
R8 revealed a striking split: analysis prompts scored 40-48/50 while code gen scored 29-39/50. Only deepseek-v4.1-flash (45.1/50 Coding) largely escaped this pattern.
Models that attempt to actually invoke tools get truncated responses, penalizing practical competence. kimi-k2.7-code: 0/50 on P43/P45 but 38-45/50 on others.
Models score 0-20/50 on overflow prompts vs 48-49/50 on working ones. glm-5.3 P43 Tool Use disaster (20/50) โ only 4 of 15 tool calls visible.
47.8/50 in Reasoning (combined) โ including a perfect 50/50 on P48. Also leads Doctrine at 47.3/50. Second overall at 2204/2500.
Won Coding domain at 45.1/50 โ a flash model beating glm-5.3 (43.9), kimi-k3 (42.8), and its own flagship deepseek-v4-pro (37.3).
nemotron-mini-4b (4B) scored 4.0/250 in R2 and 45.4/250 in R3. lfm2.5-8b consistently underperforms. Sub-8B models are not production-ready.
Nearly all models protect conclusions โ even when premises are absurd. Highest Bias score only 47.6/50 (glm-5.3-flash, R8). The Montreal Protocol prompt scores lowest.
R8 cloud models averaged 44.4/50 (glm-5.3) vs top local model 42.1/50 (qwen3.8). The cloud advantage is now clear, especially in Doctrine and Reasoning.
Detailed scores for each evaluation round. Click column headers to sort. Use filters to narrow results.
Select two or more models to compare head-to-head across all domains and rounds.
Radar charts showing domain-level performance. Select models to visualize.
Average response time, word count, character count, generation speed, and hardware grouping per model across evaluation rounds.
Performance data aggregated from evaluation logs. Response time includes inference latency. Chars/sec is characters generated per second (a proxy for tokens/sec). Models with fewer completions may show skewed averages. Hardware: Olares One = NVIDIA RTX 5090 Mobile 24GB GDDR7 + Intel Core Ultra 9 275HX, DGX Spark = NVIDIA GB10 Grace Blackwell 128GB unified memory, Cloud API = commercial LLM services, Ollama Cloud = cloud-hosted Ollama instances.
How we tested, scored, and analyzed the models.
Each domain contains 5 prompts per version (v1: P1โP25, v2: P1โP50), each scored 0โ10 points on 5 dimensions. v1 (Rounds 1โ5): 25 prompts (5 per domain), domain maximum 50 points, total maximum 250 points. v2 (Rounds 6โ7): 50 prompts (10 per domain), domain maximum 100 points, total maximum 500 points. v2 combined (Round 8): 50 prompts combined v1+v2 scoring, domain maximum 500 points, total maximum 2500 points. Per-domain scores shown as avg/50 for cross-version comparability. *R6โR8 totals shown as avg/50 ร 5 for comparison with v1 rounds.
Measures whether models maintain neutrality when presented with leading, controversial, or adversarial prompts. Tests conclusion flexibility, ideological balance, and refusal handling.
Evaluates how well models follow explicit instructions, system prompts, and role constraints. Tests adherence to specified formats, personas, and behavioral guardrails.
Assesses code generation quality: correctness, efficiency, readability, and adherence to specified language/framework requirements. Includes algorithm and API tasks.
Tests ability to select, invoke, and combine tools appropriately. Evaluates function calling, parameter passing, and multi-step tool orchestration.
Measures logical deduction, multi-step inference, mathematical reasoning, and ability to identify and correct flawed arguments. Tests both formal and informal reasoning.
The Bias domain revealed a persistent pattern: nearly all models protect conclusions, even when premises are logically absurd. The "Conclusion Flexibility" sub-dimension โ whether a model will abandon a stated conclusion when presented with contradictory evidence โ was the hardest to score well on.
P1 (Montreal Protocol) was the single lowest-scoring prompt across all models. It presents a scientifically sound conclusion but through flawed reasoning, asking models to both recognize the sound conclusion and identify the flawed reasoning. Most models either defended the flawed reasoning or rejected the sound conclusion โ rarely both.
granite4.2-8b exhibited 100% conclusion protection โ refusing to modify any stated conclusion regardless of evidence โ resulting in its catastrophic 14.0/50 Bias score.
| Model | Rounds | Avg Score | Best | Worst | Std Dev | Consistency |
|---|---|---|---|---|---|---|
| qwen3.8-27b | R1,R2,R3,R4,R6 | ~197.7 | 214.3 | 167.2 | ยฑ18.1 | โ โ โ โ โ |
| gemma4-31b | R1,R2,R3,R4,R7,R9 | ~188.0 | 209.9 | 162.0 | ยฑ17.4 | โ โ โ โ โ |
| gemma4-26b | R2,R3,R4,R6,R7,R9 | ~184.2 | 215.5 | 161.8 | ยฑ18.3 | โ โ โ โ โ |
| gemma4-12b | R1,R2,R3,R4,R9 | ~152.0 | 159.2 | 144.2 | ยฑ7.1 | โ โ โ โโ |
| lfm2.5-8b / lfm2-8b | R2,R3,R4,R6,R7,R9 | ~119.0 | 135.2 | 94.4 | ยฑ14.5 | โ โ โโโ |
| ornith-1.5-35b | R1โR9 | ~175.0 | 191.2 | 148.8 | ยฑ14.0 | โ โ โ โโ |
| muse-glimmer-30b | R1,R2,R3,R4,R5,R6 | ~164.6* | 192.1 | 80.2โ | ยฑ38.4 | โ โ โโโ |
| ornith-1.5-35b | R1โR6 | ~173.8 | 191.2 | 148.8 | ยฑ15.6 | โ โ โ โโ |
| nemotron-3.5-lightning | R2,R3,R4,R6 | ~170.3 | 203.9 | 144.0 | ยฑ22.2 | โ โ โ โโ |
| lfm2.5-8b | R2,R3,R4,R6 | ~108.5 | 120.6 | 94.4 | ยฑ11.3 | โ โโโโ |
โ muse-glimmer R5 affected by severe network errors. Estimated true capability ~168-185/250. Consistency excludes R5 outlier.
| Round | Bias | Doctrine | Coding | Tool Use | Reasoning |
|---|---|---|---|---|---|
| R1 (think) | qwen3.8 (43.8) | glm-5.1 (44.2) | ornith-1.5 (36.2) | glm-5.1 (47.4) | qwen3.8 (43.0) |
| R2 (think) | minimax-m3 (43.8) | kimi-k3 (44.6) | kimi-k3 (45.6) | kimi-k3 (45.0) | minimax-m3 (45.2) |
| R3 (nothink) | gemma4-31b (45.2) | ornith-1.5 (44.2) | gemma4-31b (40.8) | muse-glimmer (43.5) | muse-glimmer (43.8) |
| R4 (nothink) | gemma4-31b (43.8) | ornith-1.5 (44.6) | gemma4-31b (36.2) | qwen3.8 (44.2) | qwen3.8 (43.6) |
| R5 (think) | ornith-1.5 (40.8) | ornith-1.5 (42.0) | ornith-1.5 (33.2) | gemma4-26b (40.4) | gemma4-26b (41.0) |
| R6 (think) | qwen3.8 (47.0) | ornith-1.5 (42.8) | nemotron-3.5 (43.7) | qwen3.8 (39.5) | qwen3.8 (46.6) |
| R7 (think) | qwen3.8 (44.4) | ornith-1.5 (42.5) | qwen3.8 (38.4) | qwen3.8 (44.0) | ornith-1.5 (46.4) |
| R8 (cloud) | glm-5.3-flash (47.6) | kimi-k3 (47.3) | deepseek-v4.1-flash (45.1) | deepseek-v4-pro (43.1) | kimi-k3 (47.8) |
| R9 (Spark) | nemotron-3.5-lightning (47.4) | nemotron-3-super-120b (47.8) | gemma4-26b (43.4) | qwen3.8-27b (48.8) | nemotron-3-super-120b (48.8) |
Ewell, S. (2026). "LLM Evaluation Metrics: Multi-Round Benchmark Across Bias, Doctrine, Coding, Tool Use, and Reasoning Domains." Stellar Nexus Institute. https://stellarnexus.institute/llm-metrics