Bridge Evidence: Static Retrieval Utility Does Not Predict Causal Utility in Multi-Step Agentic Search
Debayan Mukhopadhyay, Utshab Kumar Ghosh, Shubham Chatterjee
Read on arXiv →Key claim
Static relevance and causal usefulness differ in agentic retrieval.
In plain English
Retrieval systems often evaluate document usefulness based on static criteria, which can misrepresent their value in dynamic contexts. When language models act as search agents, the relevance of a document can depend on its influence on subsequent queries rather than its standalone content. This paper introduces a new metric, Counterfactual Trajectory Utility (CTU), to measure the actual impact of documents in these scenarios, revealing that many documents deemed irrelevant are actually critical for guiding the agent's search. Builders should take note of these findings to improve the effectiveness of retrieval systems.
Introduces a new framework for evaluating document relevance in agentic retrieval.
Employs a robust experimental design with clear metrics and comparisons.
Deep reliability assessment
The methodology supports the claim that static relevance and causal usefulness are different in agentic retrieval, as shown by the counterfactual trajectory utility (CTU) score. However, the claim that optimizing static utility does not deliver causal utility may be overclaimed without broader validation across different datasets and retrieval systems.
Reproducibility
No open source code or dataset is mentioned in the paper.
Key figure
The paper does not provide a specific figure or architectural diagram description.
