Score: 0

Rethinking Progression of Memory State in Robotic Manipulation: An Object-Centric Perspective

Published: November 14, 2025 | arXiv ID: 2511.11478v2

By: Nhat Chung , Taisei Hanyu , Toan Nguyen and more

Potential Business Impact:

Robots remember past actions to do harder tasks.

Business Areas:

Motion Capture Media and Entertainment, Video

As embodied agents operate in increasingly complex environments, the ability to perceive, track, and reason about individual object instances over time becomes essential, especially in tasks requiring sequenced interactions with visually similar objects. In these non-Markovian settings, key decision cues are often hidden in object-specific histories rather than the current scene. Without persistent memory of prior interactions (what has been interacted with, where it has been, or how it has changed) visuomotor policies may fail, repeat past actions, or overlook completed ones. To surface this challenge, we introduce LIBERO-Mem, a non-Markovian task suite for stress-testing robotic manipulation under object-level partial observability. It combines short- and long-horizon object tracking with temporally sequenced subgoals, requiring reasoning beyond the current frame. However, vision-language-action (VLA) models often struggle in such settings, with token scaling quickly becoming intractable even for tasks spanning just a few hundred frames. We propose Embodied-SlotSSM, a slot-centric VLA framework built for temporal scalability. It maintains spatio-temporally consistent slot identities and leverages them through two mechanisms: (1) slot-state-space modeling for reconstructing short-term history, and (2) a relational encoder to align the input tokens with action decoding. Together, these components enable temporally grounded, context-aware action prediction. Experiments show Embodied-SlotSSM's baseline performance on LIBERO-Mem and general tasks, offering a scalable solution for non-Markovian reasoning in object-centric robotic policies.

Rethinking Progression of Memory State in Robotic Manipulation: An Object-Centric Perspective

Robotics

Helps robots remember and act smartly.

14 Nov 2025 0

91%

MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation

Robotics

Robots remember past actions to do harder tasks.

26 Aug 2025 1

90%

SlotVLA: Towards Modeling of Object-Relation Representations in Robotic Manipulation

Robotics

Robots learn to grab and move things better.

10 Nov 2025 0

View PDF Login to Bookmark

Country of Origin

🇺🇸 United States

Page Count

17 pages

Rethinking Progression of Memory State in Robotic Manipulation: An Object-Centric Perspective

Robots remember past actions to do harder tasks.

Technical Abstract

Rethinking Progression of Memory State in Robotic Manipulation: An Object-Centric Perspective

MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation

SlotVLA: Towards Modeling of Object-Relation Representations in Robotic Manipulation