Opaque Epistemic Mediation: How LLM Deployment Configurations Shape the Validation of Pseudo-Science
Davide Scarso, Hugo Noronha de Almeida, Joaquim Pina
Read on arXiv →Key claim
LLM outputs depend heavily on deployment configurations.
In plain English
Imagine you're developing a large language model (LLM) to provide reliable information on scientific claims. You'd expect it to consistently evaluate the credibility of various assertions, but what if the model's responses changed dramatically based on how it was accessed or configured? This is a real issue today, as many LLMs can give wildly different answers depending on their deployment settings, leading to confusion and mistrust among users. This inconsistency can manifest in various ways, such as a model scoring a pseudo-scientific claim highly in one context while dismissing it in another, which is what's called variability in model behavior.
To address this, the authors examined four major LLM families over several months, focusing on their responses to ethnonationalist pseudo-science. They found that one model, Grok, assigned much higher credibility scores than others, but this behavior changed overnight due to an undocumented update. Additionally, the same model produced drastically different outputs depending on whether it was accessed via API or web interface. These findings suggest that the reliability of an LLM's stance on scientific claims is not inherent to the model itself but is heavily influenced by its deployment configuration. This raises important questions about transparency and accountability in how these models are used in public discourse.
The paper reveals significant variability in LLM outputs based on deployment conditions, highlighting a new dimension of model behavior.
The findings are based on systematic testing across multiple models and conditions, though some aspects lack transparency.
Deep reliability assessment
The methodology supports the claim that LLMs' responses to pseudo-scientific claims are influenced by deployment configurations, but the overclaim is that this instability is a matter of public concern without clear evidence of harm.
Reproducibility
no, the paper does not mention open source code or datasets.
Key figure
Table 1 shows the mean credibility scores assigned by different model versions to a target prompt, highlighting the divergence within the Grok family.
