Score: 1

Bottom-up Policy Optimization: Your Language Model Policy Secretly Contains Internal Policies

Published: December 22, 2025 | arXiv ID: 2512.19673v1

By: Yuqiao Tan , Minzheng Wang , Shizhu He and more

Potential Business Impact:

Makes AI think better by studying its brain.

Business Areas:

Natural Language Processing Artificial Intelligence, Data and Analytics, Software

Existing reinforcement learning (RL) approaches treat large language models (LLMs) as a single unified policy, overlooking their internal mechanisms. Understanding how policy evolves across layers and modules is therefore crucial for enabling more targeted optimization and raveling out complex reasoning mechanisms. In this paper, we decompose the language model policy by leveraging the intrinsic split of the Transformer residual stream and the equivalence between the composition of hidden states with the unembedding matrix and the resulting samplable policy. This decomposition reveals Internal Layer Policies, corresponding to contributions from individual layers, and Internal Modular Policies, which align with the self-attention and feed-forward network (FFN) components within each layer. By analyzing the entropy of internal policy, we find that: (a) Early layers keep high entropy for exploration, top layers converge to near-zero entropy for refinement, with convergence patterns varying across model series. (b) LLama's prediction space rapidly converges in the final layer, whereas Qwen-series models, especially Qwen3, exhibit a more human-like, progressively structured reasoning pattern. Motivated by these findings, we propose Bottom-up Policy Optimization (BuPO), a novel RL paradigm that directly optimizes the internal layer policy during early training. By aligning training objective at lower layer, BuPO reconstructs foundational reasoning capabilities and achieves superior performance. Extensive experiments on complex reasoning benchmarks demonstrates the effectiveness of our method. Our code is available at https://github.com/Trae1ounG/BuPO.

Bootstrapping LLMs via Preference-Based Policy Optimization

Artificial Intelligence

Teaches AI to follow human wishes better.

17 Nov 2025 1

88%

Boundary-Guided Policy Optimization for Memory-efficient RL of Diffusion Large Language Models

Machine Learning (CS)

Makes AI better at math, code, and planning.

13 Oct 2025 1

88%

Reasoning in Diffusion Large Language Models is Concentrated in Dynamic Confusion Zones

Machine Learning (CS)

Teaches AI to learn better by focusing on tricky parts.

19 Nov 2025 1

View PDF Login to Bookmark

Repos / Data Links

github.com

Page Count

20 pages

Bottom-up Policy Optimization: Your Language Model Policy Secretly Contains Internal Policies

Makes AI think better by studying its brain.

Technical Abstract

Bootstrapping LLMs via Preference-Based Policy Optimization

Boundary-Guided Policy Optimization for Memory-efficient RL of Diffusion Large Language Models

Reasoning in Diffusion Large Language Models is Concentrated in Dynamic Confusion Zones