Score: 0

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL

Published: August 26, 2025 | arXiv ID: 2508.18588v1

By: Jingkai He , Tianjian Li , Erhu Feng and more

Potential Business Impact:

Makes AI learn faster and use computers better.

Business Areas:

Natural Language Processing Artificial Intelligence, Data and Analytics, Software

With the rapid advancement of large language models (LLMs), reinforcement learning (RL) has emerged as a pivotal methodology for enhancing the reasoning capabilities of LLMs. Unlike traditional pre-training approaches, RL encompasses multiple stages: rollout, reward, and training, which necessitates collaboration among various worker types. However, current RL systems continue to grapple with substantial GPU underutilization, due to two primary factors: (1) The rollout stage dominates the overall RL process due to test-time scaling; (2) Imbalances in rollout lengths (within the same batch) result in GPU bubbles. While prior solutions like asynchronous execution and truncation offer partial relief, they may compromise training accuracy for efficiency. Our key insight stems from a previously overlooked observation: rollout responses exhibit remarkable similarity across adjacent training epochs. Based on the insight, we introduce RhymeRL, an LLM RL system designed to accelerate RL training with two key innovations. First, to enhance rollout generation, we present HistoSpec, a speculative decoding inference engine that utilizes the similarity of historical rollout token sequences to obtain accurate drafts. Second, to tackle rollout bubbles, we introduce HistoPipe, a two-tier scheduling strategy that leverages the similarity of historical rollout distributions to balance workload among rollout workers. We have evaluated RhymeRL within a real production environment, demonstrating scalability from dozens to thousands of GPUs. Experimental results demonstrate that RhymeRL achieves a 2.6x performance improvement over existing methods, without compromising accuracy or modifying the RL paradigm.

Shuffle-R1: Efficient RL framework for Multimodal Large Language Models via Data-centric Dynamic Shuffle

Machine Learning (CS)

Makes AI better at learning from mistakes.

7 Aug 2025 1

89%

Shuffle-R1: Efficient RL framework for Multimodal Large Language Models via Data-centric Dynamic Shuffle

Machine Learning (CS)

Teaches AI to learn better and faster.

7 Aug 2025 2

89%

PipelineRL: Faster On-policy Reinforcement Learning for Long Sequence Generatio

Machine Learning (CS)

Trains AI faster and smarter using new methods.

23 Sep 2025 1

View PDF Login to Bookmark

Page Count

15 pages

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL

Makes AI learn faster and use computers better.

Technical Abstract

Shuffle-R1: Efficient RL framework for Multimodal Large Language Models via Data-centric Dynamic Shuffle

Shuffle-R1: Efficient RL framework for Multimodal Large Language Models via Data-centric Dynamic Shuffle

PipelineRL: Faster On-policy Reinforcement Learning for Long Sequence Generatio