Score: 3

Reinforced Preference Optimization for Recommendation

Published: October 14, 2025 | arXiv ID: 2510.12211v1

By: Junfei Tan , Yuxin Chen , An Zhang and more

BigTech Affiliations: Alibaba

Potential Business Impact:

Makes movie suggestions better by learning from mistakes.

Business Areas:

A/B Testing Data and Analytics

Recent breakthroughs in large language models (LLMs) have fundamentally shifted recommender systems from discriminative to generative paradigms, where user behavior modeling is achieved by generating target items conditioned on historical interactions. Yet current generative recommenders still suffer from two core limitations: the lack of high-quality negative modeling and the reliance on implicit rewards. Reinforcement learning with verifiable rewards (RLVR) offers a natural solution by enabling on-policy sampling of harder negatives and grounding optimization in explicit reward signals. However, applying RLVR to generative recommenders remains non-trivial. Its unique generation space often leads to invalid or repetitive items that undermine sampling efficiency, and ranking supervision is sparse since most items receive identical zero rewards. To address these challenges, we propose Reinforced Preference Optimization for Recommendation (ReRe), a reinforcement-based paradigm tailored to LLM-based recommenders, an important direction in generative recommendation. ReRe incorporates constrained beam search to improve sampling efficiency and diversify hard negatives, while augmenting rule-based accuracy rewards with auxiliary ranking rewards for finer-grained supervision. Extensive experiments on three real-world datasets demonstrate that ReRe consistently outperforms both traditional and LLM-based recommenders in ranking performance. Further analysis shows that ReRe not only enhances performance across both base and SFT-initialized models but also generalizes robustly across different backbone families and scales. Beyond empirical gains, we systematically investigate the design space of RLVR in recommendation across generation, sampling strategy, reward modeling, and optimization algorithm, offering insights for future research.

Beyond Static LLM Policies: Imitation-Enhanced Reinforcement Learning for Recommendation

Information Retrieval

Makes movie suggestions faster and smarter.

15 Oct 2025 2

90%

The Reasoning Boundary Paradox: How Reinforcement Learning Constrains Language Models

Artificial Intelligence

Fixes AI reasoning errors by focusing on hard problems.

2 Oct 2025 1

89%

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle

Computation and Language

Teaches computers to think and follow instructions better.

20 Sep 2025 1

View PDF Login to Bookmark

Country of Origin

🇸🇬 🇨🇳 Singapore, China

Repos / Data Links

github.com

Page Count

23 pages

Reinforced Preference Optimization for Recommendation

Makes movie suggestions better by learning from mistakes.

Technical Abstract

Beyond Static LLM Policies: Imitation-Enhanced Reinforcement Learning for Recommendation

The Reasoning Boundary Paradox: How Reinforcement Learning Constrains Language Models

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle