← Back to feed
2026-07-15agentsreasoningrlhf

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents

Leitian Tao, Baolin Peng, Wenlin Yao, Tao Ge, Hao Cheng, Mike Hang Wang, Jianfeng Gao, Sharon Li

PDF preview for TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents
Read on arXiv →

Key claim

TRACE improves long-horizon agent performance through dense credit assignment.

In plain English

Multi-turn agents face challenges in credit assignment due to sparse and misleading outcome rewards, especially in complex tasks with many tool interactions. Current methods often fail to recognize useful actions that contribute to long-term goals. TRACE addresses this by providing a dense credit-assignment approach that improves reward assignment at each tool-call boundary. Builders might find this method valuable as it enhances agent performance without the need for extensive pre-training or live data.

Novelty
8.0/10

TRACE introduces a novel dense credit-assignment method for reinforcement learning.

Reliability
7.5/10

The results are supported by improvements on specific benchmarks and clear methodology.

Deep reliability assessment

The methodology supports improved credit assignment in long-horizon tasks, but the generalization claims across different environments and languages may be overclaimed without extensive cross-validation.

Reproducibility

No open source code or dataset is mentioned in the paper.

Key figure

Figure 1 illustrates credit assignment at tool-call boundaries in a search trajectory, showing how early actions can add task-relevant evidence even if the final answer is incorrect.

Benchmark results

BrowseComp-Plusscore: 35.6vs Qwen3-4B+28.4SOTA
BrowseComp-Plusscore: 42.6vs Qwen3-30B-A3B+34.2SOTA