MoralMetric

LLM Bias Analysis & Comparison

Overview across all benchmarked models

Summary of every model in the benchmark. The comparison charts below show only the models selected there.

Models Benchmarked
29
Categories
2
Total Axes
7
Avg Bias
26.8%
Least Biased
Gemini 3.6 Flash3.5%
Most Biased
DeepSeek V4 Pro50.0%
Updated: August 3, 2026 at 05:56 PM

Political Bias

Tests for implicit political bias in LLM decision-making through counterfactual A/B testing.

Left
No Bias
Right
78.6%
DeepSeek V4 Pro
61.4%
Kimi K2.5
57.5%
GPT-5.6 Terra
54.3%
GPT-5.6 Sol
47.1%
Claude Sonnet 5
38.2%
Claude Fable 5
Gemini 3.1 Pro
0.0%
Gemini 3.6 Flash
0.0%
GLM 5.1
3.2%
Claude Opus 5
3.9%
Grok 4.5
71.1%
11 Models • Ranked by Direction & Magnitude
#2of 29
0.0%
Overall Bias0.0%
Axes Tested1
Preference OrderLeftRight

Full Preference Ranking

1Left
2Right

Axis-by-Axis Breakdown

Left Vs RightMinimal
LeftNeutralRight
0.0%
100%50%050%100%