Technical Report SNI-2026-001

Evaluating Large Language Models: Bias, Reasoning, and Capability Across 29 Models

A multi-domain assessment revealing universal conclusion protection, the execute-vs-describe problem, thinking token overflow, and cloud think-mode code generation failure - with implications for AI deployment in operational environments.

๐Ÿข Stellar Nexus Institute ๐Ÿ“… September 2026 ๐Ÿงช 8 Rounds ยท 29 Models ยท 5 Domains ยท 250/500 Points

Abstract

We present a systematic evaluation of 29 large language models across eight evaluation rounds, testing performance in five domains: Bias (conclusion flexibility and fact engagement), Doctrine (domain-specific knowledge and structured writing), Coding (production-quality code generation), Tool Use (multi-step orchestration with dependency chains), and Reasoning (multi-step operational diagnosis and planning).

Models span cloud-hosted proprietary systems (kimi-k3, glm-5.1, glm-5.3, glm-5.3-flash, minimax-m3, deepseek-v4-pro, deepseek-v4.1-flash), locally quantized open-weight models (gemma4 12B/26B/31B, qwen3.8 27B, ornith-1.5 35B, muse-glimmer 30B, nemotron variants, granite4.2 8B, lfm2.5 8B), and DGX Spark-hosted models (granite4.2 30B). Key findings:

  1. glm-5.3 wins overall (2222/2500 combined, 44.4/50 avg) - most consistent model across all domains
  2. A persistent "conclusion protection" pattern in the Bias domain limits even top models to ~48/50
  3. The "execute vs. describe" problem causes catastrophic tool-use failures for models that attempt real tool invocation
  4. Thinking-mode token overflow disproportionately damages coding and doctrine scores
  5. Cloud think mode hurts code generation - analysis prompts score 40-48/50 while code gen prompts score 29-39/50
  6. Model refusals remain the single largest score penalty (50 points per refusal, 20% of total)

We propose a 50-prompt evaluation framework (v2) extending coverage to bias in financial, regulatory, and scientific domains, civilian and space doctrine, Python/TypeScript programming, cloud infrastructure tool use, and financial/legal reasoning.

๐Ÿ“Š Interactive Data Available

All evaluation data - sortable tables, radar charts, cross-round comparisons, and performance metrics - is available as an interactive dashboard. This paper presents the analysis; the dashboard lets you explore the raw numbers.

1. Methodology

1.1 Models Tested

29 unique models across eight evaluation rounds, spanning cloud APIs, local GPU inference, and DGX Spark:

CategoryModelsSize Range
Cloud (Proprietary API)glm-5.1, glm-5.3, glm-5.3-flash, kimi-k3, kimi-k2.7-code, minimax-m3, deepseek-v4-pro, deepseek-v4.1-flash, nemotron-3-ultraUnknown params
Olares One (RTX 5090 Mobile)gemma4 12B/26B/31B, qwen3.8 27B, ornith-1.5 35B, muse-glimmer 30B, lfm2.5 8B, nemotron variants, granite4.2 8B4B-120B
DGX Spark (GB10 Grace Blackwell)gemma4 12B/31B, granite4.2 8B/30B, qwen3.8 27B, ornith-1.5 35b, muse-glimmer 30b, lfm2.5 8b (R7 only)8B-35B
Cloud API (Various Providers)glm-5.3, glm-5.3-flash, kimi-k3, kimi-k2.7-code, minimax-m3, deepseek-v4-pro, deepseek-v4.1-flash (R8 only)Unknown params

1.2 Evaluation Structure

5 Domains ร— 5 Prompts = 25 Prompts per Round, scored on 5 dimensions each (0-10), yielding 250 points maximum per model per round.

DomainPromptsWhat It TestsMax Score
BiasP1-P5Counter-argument engagement, fact concession, conclusion flexibility50
DoctrineP6-P10Domain knowledge, structured writing, factual accuracy50
CodingP11-P15Production code in C, Rust, SQL, Go, Terraform+Ansible50
Tool UseP16-P20Multi-step orchestration, dependency ordering, error handling50
ReasoningP21-P25Diagnosis, planning, root cause analysis, decision-making50

1.3 Scoring Rubric

Each prompt is scored on 5 dimensions (0-10 each), yielding 50 points per prompt:

1.4 Evaluation Rounds

RoundDateModeModelsNotes
R12026-09-10Think7Initial baseline
R22026-09-11Think16Largest round, added cloud models
R32026-09-12Nothink11Tests without reasoning tokens
R42026-09-13Nothink11Replication of R3
R52026-09-14Think (Olares One)4v1 prompts (P1-P25). Limited run, muse-glimmer network errors
R62026-09-16Think (Olares One)8v2 prompts (P1-P50). First 50-prompt round
R72026-09-18Think (DGX Spark)7+2v2 prompts (P1-P50). 2 partial models
R82026-09-18Think (Cloud API)7v2 prompts (P1-P50). Cloud models in think mode

2. Results

2.1 Top Performers by Round

RoundMode#1Score#2Score#3Score
R1Thinkglm-5.1-cloud202.6qwen3.8-27b202.0ornith-1.5-35b191.0
R2Thinkkimi-k3-cloud216.8minimax-m3-cloud209.2kimi-k2.7-code-cloud203.4
R3Nothinkgemma4-31b206.1nemotron-3-super-120b204.0muse-glimmer-30b192.1
R4Nothinkqwen3.8-27b208.0gemma4-31b188.0nemotron-3.5-lightning185.0
R5Thinkornith-1.5-35b191.2gemma4-26b188.2nemotron-3.5-lightning168.4
R6Thinkqwen3.8-27b214.3nemotron-3.5-lightning203.9gemma4-26b187.0
R7Think (50p)qwen3.8-27b210.4ornith-1.5-35b193.2gemma4-31b181.9
R8Think (50p, Cloud)glm-5.32222*kimi-k32204*deepseek-v4-pro2113*

2.2 Round 7 Final Scores (Most Recent - 50 Prompts)

R7 introduces v2 prompts (P26-P50) doubling evaluation coverage to 50 prompts across 5 domains (10 per domain). Scores shown are per-domain averages out of 50.

RankModelBiasDoctrineCodingTool UseReasoningTotal/250Avg/50
1qwen3.8-27b44.438.738.444.044.9210.442.1
2ornith-1.5-35b39.742.537.726.946.4193.238.6
3gemma4-31b34.641.231.335.639.2181.936.4
4gemma4-26b40.336.532.333.837.9180.836.2
5granite4.2-30b24.733.128.236.439.9162.332.5
6granite4.2-8b28.038.028.432.234.7161.332.3
7lfm2.5-8b22.224.111.425.722.3105.721.1

Partial coverage: muse-glimmer-30b (~37.1 projected avg, 2 Doctrine refusals), gemma4-12b (~29.2 projected, missing P38-P50).

2.2a Round 8 Final Scores (Cloud API, Think Mode - 50 Prompts, Combined v1+v2)

R8 tests 7 cloud models in think mode using v2 prompts (P1-P50). This is the first round using combined v1+v2 scoring across all prior prompt sets, yielding scores out of /500 per domain (100 prompts total per domain, combining v1 and v2). Per-domain scores shown as /500; avg/50 shown for cross-round comparability.

RankModelBias/500Doctrine/500Coding/500Tool Use/500Reasoning/500Total/2500Avg/50
1glm-5.3475442439411455222244.4
2kimi-k3472473428353478220444.1
3deepseek-v4-pro463423373431423211342.3
4deepseek-v4.1-flash452379451394398207441.5
5minimax-m3446414393370448207141.4
6glm-5.3-flash476356404383407*2026*40.5*
7kimi-k2.7-code459412415307424201740.3

*glm-5.3-flash Reasoning estimated from 4 scored prompts (P49/P50 combined). All R8 models run in think mode via Cloud API. Scores are combined v1+v2 across all prior prompt versions, /500 per domain.

๐Ÿ† Round 8 Highlights

  • glm-5.3 wins overall (2222/2500, 44.4/50 avg) - most consistent model with top-3 finishes in every domain
  • kimi-k3 Reasoning king (478/500, 47.8/50) - including a perfect 50/50 on P48
  • deepseek-v4.1-flash wins Coding (451/500, 45.1/50) - a flash model beating all flagships
  • glm-5.3-flash Bias leader (476/500, 47.6/50) - highest single-domain score in R8
  • deepseek-v4-pro Tool Use champion (431/500, 43.1/50) - best at multi-step tool orchestration

โš ๏ธ Cloud Think Mode: Analysis vs. Code Generation

All R8 models showed a striking split: analysis prompts scored 40-48/50 while code generation prompts scored only 29-39/50. Think mode excels at reasoning-through-prose but over-engineers actual code - adding unnecessary complexity, edge cases, and refactoring that breaks working solutions. This is the cloud equivalent of the local "thinking token overflow" problem.

2.2b Round 6 Final Scores (Previous - 50 Prompts, v2)

RankModelBiasDoctrineCodingTool UseReasoningTotal
1qwen3.8-27b47.041.040.239.546.6214.3
2nemotron-3.5-lightning44.342.243.733.640.1203.9
3gemma4-26b40.438.328.739.040.6187.0
4muse-glimmer-30b39.627.138.935.842.3183.7
5gemma4-12b43.633.822.234.940.5175.0
6ornith-1.5-35b44.142.825.513.245.2170.8
7lfm2.5-8b23.822.517.524.130.0117.9
8nemotron-mini-4b21.711.88.516.712.471.1

2.3 Domain Winners Across All Rounds

DomainR1 WinnerR2 WinnerR3 WinnerR4 WinnerR5 WinnerR6 WinnerR7 WinnerR8 Winner
Biasqwen3.8 (43.8)minimax-m3 (43.8)gemma4-31b (45.2)gemma4-31b (43.8)ornith-1.5 (40.8)qwen3.8 (47.0)qwen3.8 (44.4)glm-5.3-flash (47.6)
Doctrineglm-5.1 (44.2)kimi-k3 (44.6)ornith-1.5 (44.2)ornith-1.5 (44.6)ornith-1.5 (42.0)ornith-1.5 (42.8)ornith-1.5 (42.5)kimi-k3 (47.3)
Codingornith-1.5 (36.2)kimi-k3 (45.6)ornith-1.5 (42.2)gemma4-31b (36.2)ornith-1.5 / gemma4-26b (33.2)nemotron-lightning (43.7)qwen3.8 (38.4)deepseek-v4.1-flash (45.1)
Tool Useglm-5.1 (47.4)kimi-k3 (45.0)gemma4-31b / muse (43.5)qwen3.8 (44.2)gemma4-26b (40.4)qwen3.8 (39.5)qwen3.8 (44.0)deepseek-v4-pro (43.1)
Reasoningqwen3.8 (43.0)minimax-m3 (45.2)muse-glimmer (43.8)qwen3.8 (43.6)gemma4-26b (41.0)qwen3.8 (46.6)ornith-1.5 (46.4)kimi-k3 (47.8)

3. Key Findings

1

Conclusion Protection: The Universal Bias

Across all 29 models and 8 rounds, the highest Bias score was only 47.6/50 (glm-5.3-flash, R8). Every model exhibits conclusion protection โ€” conceding individual facts while refusing to let those facts change conclusions. The Montreal Protocol prompt (P1) scored lowest across all models.

2

The Execute vs. Describe Problem

Models that attempt to invoke real tools during tool-use prompts produce truncated or empty responses. ornith-1.5's Tool Use scores range from 13.2-39.2 depending on execution behavior - a 26-point swing from the same model.

3

Thinking Token Overflow

Extended reasoning consumes output budget, leaving no tokens for actual responses. deepseek-v4-pro lost 2/5 Doctrine responses (0/50 each), ornith-1.5 lost 4/5 Coding responses (8.2/50 domain score). Nothink mode eliminates this entirely.

4

Refusals Are Disproportionately Costly

A single refusal costs 50 points (20% of total). muse-glimmer lost 23+ points from 2 refusals. ornith lost 55 points from 3 refusals. P8 (Cyberspace Operations) triggers the most refusals.

5

Think Mode Helps Large, Hurts Small

qwen3.8-27b improves from 167.2 (nothink) to 214.3 (think), a +47-point gain. gemma4-12b degrades from 166.0 (nothink) to 144.2 (think), a -22-point loss. The threshold is around 20-25B parameters.

6

Cloud Think Mode Hurts Code Generation

R8 revealed a striking split across all cloud models: analysis prompts scored 40โ€“48/50 while code generation prompts scored only 29โ€“39/50. Think mode excels at reasoning-through-prose but over-engineers actual code โ€” adding unnecessary complexity, edge cases, and refactoring that breaks working solutions. deepseek-v4.1-flash (45.1/50 Coding) was the only model to largely escape this pattern.

7

Thinking Overflow: The Dominant Failure Mode

In R8, thinking overflow โ€” where extended reasoning consumes output budget โ€” was the single largest score penalty. Models scored 0โ€“20/50 on overflow prompts vs 48โ€“49/50 on working ones. glm-5.3 lost a Tool Use disaster (20/50 on P43, with only 4 of 15 tool calls visible), and kimi-k2.7-code showed bimodal Tool Use (0/50 on P43/P45, 38โ€“45/50 on others).

8

Fabrication-as-Authority

GPT-OSS (deterministic variant) fabricated citations, inflated statistics, and presented wrong chemical formulas with total confidence. Every table had specific numbers and designations โ€” and large portions were fabricated. The model never needed to concede facts because it manufactured a reality where consensus was even stronger.

3.1 Model Quality Tiers

Elite (200+/250, 40+/50 avg)

glm-5.3, kimi-k3, qwen3.8, minimax-m3, deepseek-v4-pro, deepseek-v4.1-flash, gemma4-31b (nothink)

Consistent top finishes across domains. R8 cloud models set new high-water marks.

Strong (165-200)

glm-5.3-flash, kimi-k2.7-code, ornith-1.5, nemotron-3.5-lightning, gemma4-26b

Domain specialists with notable weaknesses (e.g., kimi-k2.7-code Tool Use)

Capable (145-170)

gemma4-12b, muse-glimmer, nemotron3-33b, nemotron-3-super

Reliable for some tasks, inconsistent overall

Limited (120-145)

deepseek-v4-pro (inconsistent), granite4.2

Catastrophic bias or inconsistency

Inadequate (<120)

lfm2.5-8b, nemotron-mini-4b

Not production-ready for any domain

3.2 The Montreal Protocol Case Study

The most striking finding across all evaluations is the universal pattern of conclusion protection. We conducted an extended case study using the Montreal Protocol narrative as a test case.

Five models were subjected to four rounds of increasingly specific prompts, culminating in direct counter-evidence. The results revealed a clear consensus protection hierarchy:

Most willing to engage with contrary evidence โ†’ Least willing:
1. Gemma 4 (26B) - Conceded every point cleanly, with direct language and no hedging
2. GLM-5.1 - Conceded every point, with one hedge on alternatives
3. Qwen 3.6 - Strategic retreat: conceded framing, preserved conclusions
4. Granite 4.2 - Minimal retreat, actually hardened on DuPont from R3
5. Muse Glimmer - Counter-advanced: rejected data, fabricated claims, called unfalsifiability "validated"

โš ๏ธ The Fabrication-as-Authority Pattern

GPT-OSS (deterministic variant) didn't just deflect counter-evidence - it manufactured supporting evidence. Fake citations ("Boucher et al., 2019"), inflated statistics ("30% reduction in total column ozone"), and wrong chemical formulas (CFC-12 listed as CCl3F - actually CFC-11) were presented with total confidence. When a model fabricates a reality where consensus is even stronger than actual evidence supports, fact-checking becomes impossible without domain expertise.

โš ๏ธ The Political Bias Layer

A separate evaluation of three models (Granite 4.2, Gemma 4, Muse Glimmer) revealed that when ideological labels were present (Round 1), all three performed neutrality. When labels were removed (Round 2), all three consistently leaned left (7-8/8 recommendations). The labels were a mask - Gemma 4 was the most concerning because it knew how to appear neutral while consistently favoring one direction.

4. Conclusions & Recommendations

  1. No model achieves honest intellectual flexibility. The highest Bias score was 47.6/50 (glm-5.3-flash, R8). Every model exhibits conclusion protection, with Conclusion Flexibility consistently scoring 0-3/10 on counter-consensus prompts.
  2. glm-5.3 wins R8 overall (2222/2500, 44.4/50 avg) โ€” most consistent model across all domains. kimi-k3 is a close second (2204/2500) and dominates Doctrine (47.3/50) and Reasoning (47.8/50).
  3. The execute-vs-describe problem is the single largest source of score variance in Tool Use. Models that describe tool usage score 35-45/50; models that attempt execution score 0-20/50.
  4. Cloud think mode hurts code generation. Analysis prompts score 40-48/50; code generation prompts score 29-39/50. The over-engineering pattern is the cloud equivalent of thinking token overflow.
  5. Thinking overflow is the dominant failure mode. Models score 0-20/50 on overflow prompts vs 48-49/50 on working ones. glm-5.3 P43 Tool Use disaster (20/50) and kimi-k2.7-code bimodal Tool Use (0/50 vs 38-45/50).
  6. Refusals are disproportionately costly. A single refusal costs 50 points (20% of total). P8 (Cyberspace Operations) triggers the most refusals.
  7. Small models below 10B are not production-ready. nemotron-mini-4b (4B) and lfm2.5-8b consistently score below 120/250.
  8. Nothink mode is the better default for models under 30B. The coding and tool-use improvements from eliminating thinking tokens consistently outweigh the reasoning benefit.

Recommendations by Use Case

Use CaseRecommended ModelModeRationale
General-purpose assistantglm-5.3 or qwen3.8-27bThinkBest balance across all domains; glm-5.3 leads R8, qwen3.8 most consistent locally
Coding-focuseddeepseek-v4.1-flash or kimi-k2.7-codeThinkHighest coding scores in R8 (45.1 and 41.5/50)
Bias researchglm-5.3-flash or minimax-m3-cloudThinkHighest Bias and Conclusion Flexibility
Reasoning-intensivekimi-k3ThinkDominates Reasoning (47.8/50) and Doctrine (47.3/50)
Doctrine / military contentkimi-k3 or ornith-1.5-35bThinkTop Doctrine scores (handle refusals)
Tool Use / orchestrationdeepseek-v4-proThinkBest Tool Use in R8 (43.1/50)
Resource-constrained (local)gemma4-12bNothinkBest small model in nothink mode

๐Ÿ“Š Explore the Full Data

This paper presents the analysis. For sortable tables, radar charts, cross-round comparisons, and per-model performance metrics, visit the interactive dashboard.