TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents
Leitian Tao, Baolin Peng, Wenlin Yao, Tao Ge, Hao Cheng, Mike Hang Wang, Jianfeng Gao, Sharon Li
Read on arXiv →Key claim
TRACE improves long-horizon agent performance through dense credit assignment.
In plain English
Multi-turn agents face challenges in credit assignment due to sparse and misleading outcome rewards, especially in complex tasks with many tool interactions. Current methods often fail to recognize useful actions that contribute to long-term goals. TRACE addresses this by providing a dense credit-assignment approach that improves reward assignment at each tool-call boundary. Builders might find this method valuable as it enhances agent performance without the need for extensive pre-training or live data.
TRACE introduces a novel dense credit-assignment method for reinforcement learning.
The results are supported by improvements on specific benchmarks and clear methodology.
Deep reliability assessment
The methodology supports improved credit assignment in long-horizon tasks, but the generalization claims across different environments and languages may be overclaimed without extensive cross-validation.
Reproducibility
No open source code or dataset is mentioned in the paper.
Key figure
Figure 1 illustrates credit assignment at tool-call boundaries in a search trajectory, showing how early actions can add task-relevant evidence even if the final answer is incorrect.
