← Back to feed
2026-07-13alignmentdatacode

Inside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge Bias

Zixiang Xu, Sixian Li, Huaxing Liu, Xiang Wang, Shuai Li, Zirui Song, Xiuying Chen

PDF preview for Inside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge Bias
Read on arXiv →

Key claim

Bias in LLMs can be controlled via hidden state geometry.

In plain English

Bias in LLM scoring is often studied by changing inputs and observing score changes, but this paper highlights that biases also exist in the hidden states of the models. Current methods may overlook this representation-level perspective, which can provide deeper insights into bias behavior. The authors show that by analyzing the geometry of hidden states, they can predict and control scoring biases more effectively than traditional input-output methods. Builders might find this approach useful for developing more reliable and fair LLM applications.

Novelty
8.0/10

Introduces a novel representation-level account of scoring bias in LLMs.

Reliability
7.5/10

Findings are supported by multiple judges and benchmarks, though details on methodology are limited.

Deep reliability assessment

The methodology supports the claim that biases in LLM judges can be represented as geometric structures in hidden states, allowing for causal manipulation. However, the operational utility of these findings on unseen benchmarks may be overclaimed without extensive validation.

Reproducibility

Yes, the project page is available at https://xzx34.github.io/unfair-judge/, which likely includes code and datasets.

Key figure

Figure 1 likely illustrates the geometric displacement of biased inputs along a low-dimensional subspace in the activation manifold.

Codelink
xzx34.github.io/unfair-judgeOfficial