Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Observers on Shared Endpoints
Haoyaun Zhu, Jie Zhang
Read on arXiv →Key claim
Model evaluations are less reliable than previously assumed.
In plain English
Imagine you're developing a language model that needs to consistently generate accurate responses based on user queries. You might assume that if you ask the same question today and tomorrow, the model will give you the same answer. However, this paper reveals that this assumption often fails in practice. In two large-scale audits involving nearly 53,000 requests, the consistency of rankings for identical queries was significantly lower than expected, with Spearman correlations falling short of the required thresholds. This inconsistency can arise from various factors, including biases in how labels are interpreted and noise in the model's output that can lead to different rankings even for the same input. These issues are termed 'measurement instrument failures' and highlight the fragility of relying on model names as stable indicators of performance. To address these challenges, the authors propose a structured approach that includes a three-level snapshot-identity ladder and a set of design rules for better evaluation practices. They emphasize the importance of measuring a model's performance before making any assumptions about its reliability. This work shifts the focus from simply trusting model outputs to critically assessing the conditions under which those outputs are generated, which is crucial for anyone building applications that depend on consistent and reliable AI behavior.
The paper uncovers significant issues with the reliability of language model evaluations, challenging existing assumptions.
The study is based on extensive audits and preregistered campaigns, providing solid empirical evidence.
Deep reliability assessment
The methodology supports the claim that language model observers on shared endpoints are unreliable due to infrastructure-induced variability, but it overclaims by suggesting that no current solutions can address this issue effectively.
Reproducibility
no
Key figure
The paper does not provide a specific figure or architectural diagram description.
