← Back to feed
2026-06-30· Yongkangagentsreasoningscalingrlhf

GR2 Technical Report

Yufei Li, Zaiwei Zhang, Mingfu Liang, Kavosh Asadi, Jay Xu, Jimmy Kim, Chongyang Bai, Jieyi Zhang, Hongye Xie, Prachi Agrawal, Dian Yu, Tianyi Chen, Jean-Pascal Billaud, Garret Buell, YK, Zhu, Sachin Patil, Brooke Bian, Zhou Fang, Kevin Huang, Shiva Sudanagunta, Yuzhen Huang, Emma Lu, Chris O'Brien, Yang Song, Lihong Li, Jacob Tao, Zhicheng Zhu, Chao Li, Gaoxiang Liu, Neil Wu, Zhongyin Hu, Li Han, Loki Chen, Ming Lei, Greg Rehm, Siyuan Song, Tianwei Zhang, Li Li, Ketan Singh, Yavuz Yetim, Ilyas Atishev, Satendra Gera, Ashkan Sadeghi, Rachel Yan, Nikko Mizutani, Shuaiwen Wang, Song Yang, Zhijing Li, Jiang Liu, Mengying Sun, Fei Tian, Xiaohan Wei, Chonglin Sun, Parish Aggarwal, Kaushik Rangadurai, Zhi Hua, Frank Shyu, Ruchit Sharma, Liyuan Li, Shike Mei, Wenlin Chen, Santanu Kolay, Ben Schulte, Deepak Chandra, Adam, Song, Sandeep Pandey, Xi Liu, Hamed Firooz, Luke Simon

PDF preview for GR2 Technical Report
Read on arXiv →

Key claim

GR2 improves re-ranking performance by over 18% on key metrics.

In plain English

Imagine you're building a recommendation system that needs to show users the most relevant items from a massive catalog. The challenge is that the final step of re-ranking — deciding which items to display after initial filtering — is crucial for keeping users engaged. However, many existing systems focus on earlier stages like retrieval and ranking, leaving re-ranking underexplored. This can lead to missed opportunities where the displayed items don't resonate with users, ultimately hurting engagement. This is what's called a re-ranking gap.

Current methods often use large language models (LLMs) in a zero-shot or fine-tuning manner, which doesn't fully leverage their reasoning capabilities. Additionally, many catalogs use non-semantic identifiers that LLMs can't easily understand, complicating the process. This is where GR2 comes in. It combines several innovative techniques: it trains on unique semantic IDs, distills reasoning from a stronger model, and employs reinforcement learning with verifiable rewards tailored for re-ranking.

What sets GR2 apart is its ability to effectively handle the unique challenges of re-ranking in a way that previous methods have not. It not only improves the relevance of displayed items but also addresses the critical issue of reward design, ensuring that LLMs don't exploit biases in the data. In practical terms, this means that if you're building a recommendation system, using GR2 could lead to a significant boost in user engagement metrics, making it a compelling choice for industrial applications.

Novelty
8.0/10

The paper introduces a novel framework for re-ranking in recommendation systems that leverages LLMs and reinforcement learning, addressing significant gaps in current approaches.

Reliability
8.0/10

The claims are supported by strong experimental results on industrial-scale traffic, demonstrating clear improvements over legacy methods.

Deep reliability assessment

The methodology supports the claim that GR2 improves re-ranking performance with specific metrics, but the generalizability to different industrial settings may be overclaimed without broader testing.

Reproducibility

No open source code or dataset is mentioned in the paper.

Key figure

Figure 1 illustrates the three-stage training pipeline of GR2, highlighting mid-training on semantic IDs, reasoning-trace distillation, and reinforcement learning with verifiable rewards.

Benchmark results

~industrial-scale trafficR@1: 18.7vs legacy baselines+18.7%SOTA
~industrial-scale trafficR@3: 7.1vs legacy baselines+7.1%SOTA
~industrial-scale trafficN@3: 9.6vs legacy baselines+9.6%SOTA
GR2 Technical Report — Frontier Papers