Score: 0

Failure-Aware RL: Reliable Offline-to-Online Reinforcement Learning with Self-Recovery for Real-World Manipulation

Published: January 12, 2026 | arXiv ID: 2601.07821v1

By: Huanyu Li , Kun Lei , Sheng Zang and more

Potential Business Impact:

Robots learn to avoid mistakes and work better.

Business Areas:

Industrial Automation Manufacturing, Science and Engineering

Post-training algorithms based on deep reinforcement learning can push the limits of robotic models for specific objectives, such as generalizability, accuracy, and robustness. However, Intervention-requiring Failures (IR Failures) (e.g., a robot spilling water or breaking fragile glass) during real-world exploration happen inevitably, hindering the practical deployment of such a paradigm. To tackle this, we introduce Failure-Aware Offline-to-Online Reinforcement Learning (FARL), a new paradigm minimizing failures during real-world reinforcement learning. We create FailureBench, a benchmark that incorporates common failure scenarios requiring human intervention, and propose an algorithm that integrates a world-model-based safety critic and a recovery policy trained offline to prevent failures during online exploration. Extensive simulation and real-world experiments demonstrate the effectiveness of FARL in significantly reducing IR Failures while improving performance and generalization during online reinforcement learning post-training. FARL reduces IR Failures by 73.1% while elevating performance by 11.3% on average during real-world RL post-training. Videos and code are available at https://failure-aware-rl.github.io.

Feasibility-Guided Fair Adaptive Offline Reinforcement Learning for Medicaid Care Management

Machine Learning (CS)

Makes AI fairer and safer for everyone.

11 Sep 2025 1

88%

FailSafe: Reasoning and Recovery from Failures in Vision-Language-Action Models

Robotics

Robots learn to fix their own mistakes.

2 Oct 2025 1

88%

Offline Multi-Agent Reinforcement Learning for 6G Communications: Fundamentals, Applications and Future Directions

Multiagent Systems

Teaches AI to control many devices safely.

1 Jan 2026 0

View PDF Login to Bookmark

Page Count

9 pages

Failure-Aware RL: Reliable Offline-to-Online Reinforcement Learning with Self-Recovery for Real-World Manipulation

Robots learn to avoid mistakes and work better.

Technical Abstract

Feasibility-Guided Fair Adaptive Offline Reinforcement Learning for Medicaid Care Management

FailSafe: Reasoning and Recovery from Failures in Vision-Language-Action Models

Offline Multi-Agent Reinforcement Learning for 6G Communications: Fundamentals, Applications and Future Directions