Divide-and-RememberRecursive Action-Relevant Memory for Long-Horizon VLA Policies

Anonymous authors
Preprint, 2026
Paper (coming soon) Code (coming soon)
Memory is a question of what to remember
Memory is a question of what to remember. (a) A history-dependent task: “watch the video carefully, then move the cube to the target in the same manner as before”. Two videos show different manners (pick-and-place versus peg push), yet at time step t the robot sees the same observation. A memoryless policy cannot recall how the cube was moved. (b) We view memory as an optimisation problem: find the memory mt = π“œ(ht), selected from the full history, that carries the key fact the current observation lacks (here, the grasp pose at t−3), so that the policy knows how to move the cube.

Abstract

Vision–language–action (VLA) models struggle on history-dependent manipulation tasks, where the current observation alone does not determine the action, and the policy needs a memory of the history. Existing memory methods decide what to remember by design, for example, keeping frames with large pixel changes, and show inconsistent gains across tasks. We view what to remember as an optimisation problem. From the POMDP formulation of imitation learning, we show that the optimal memory maximises the conditional mutual information I(at; mt | ot) between the action and the memory given the current observation. Intuitively, this means preserving the action-relevant information in the history that is not already contained in the current observation. Based on our analysis, we propose Divide-and-Remember (D&R), a recursive memory method that learns a memory function mt = π“œ(ht) and scales to long contexts while staying compute-light. It involves two strategies: (1) the selection over the full history is divided recursively into subproblems of top-K selection over 2K tokens, so that fixed-size, lightweight selectors learned end-to-end support an unbounded history; (2) all recursion blocks share one selector, which captures the selection rule common to every block and keeps the method efficient. On RoboMME, a benchmark of 16 long-horizon manipulation tasks that require remembering when, where, what, and how to act, D&R achieves a state-of-the-art average success rate with consistent gains across all four suites under a budget of only 64 tokens; real-robot experiments show the same gain.

Simulation Rollout Demos

Front and wrist camera of the RoboMME evaluation episodes. The same episode id is the same environment across methods; pick the episode and the two methods to compare. All methods share the π0.5 backbone; the memoryless π0.5 sees only the current observation.

Task 1.1: BinFillTask Goal: put one red cube into the bin, then press the button to stop
Task 1.2: PickXtimesTask Goal: pick up the green cube and place it on the target, then press the button to stop
Task 1.3: SwingXtimesTask Goal: pick up the green cube, move it to the top of the right-side target, then move it to the top of the left-side target, repeating this back-and-forth motion two times, finally press the button to stop
Task 1.4: StopCubeTask Goal: press the button to stop the cube just as it reaches the target for the fourth time
Task 2.1: VideoUnmaskTask Goal: watch the video carefully, then pick up the container hiding the green cube
Task 2.2: ButtonUnmaskTask Goal: first press the button, then pick up the container hiding the green cube
Task 2.3: VideoUnmaskSwapTask Goal: watch the video carefully, then pick up the container hiding the green cube, finally pick up another container hiding the blue cube
Task 2.4: ButtonUnmaskSwapTask Goal: first press both buttons on the table, then pick up the container hiding the blue cube, finally pick up another container hiding the green cube
Task 3.1: PickHighlightTask Goal: first press the button, then pick up all cubes that have been highlighteted with white areas on the table
Task 3.2: VideoRepickTask Goal: watch the video carefully, then repeatedly pick up and put down the same block that was previously picked up for three times, finally put it down and press the button to stop
Task 3.3: VideoPlaceButtonTask Goal: watch the video carefully, then place the green cube on the target right after the button was pressed
Task 3.4: VideoPlaceOrderTask Goal: watch the video carefully, then place the green cube on the second target it was previously placed on
Task 4.1: MoveCubeTask Goal: watch the video carefully, then move the cube to the target in the same manner as before
Task 4.2: InsertPegTask Goal: watch the video carefully, then grasp the same end of the same peg you've picked before and insert it into the same side of the box
Task 4.3: PatternLockTask Goal: watch the video carefully, then use the stick attached to the robot to retrace the same pattern
Task 4.4: RouteStickTask Goal: watch the video carefully, then use the stick attached to the robot to navigate around the sticks on the table, following the same path

Real Robot Rollout Demos

Real-robot trials on a Franka Panda, filmed by an external camera, together with the two input views, the front camera (also used as memory) and the wrist camera (left to right), in real time. For each task, we evaluate 10 trials.

PutBottlesCounting
Task Goal:
π0.5 (no memory)
FrameSamp
D&R (ours)
TrackCubePermanence
Task Goal:
π0.5 (no memory)
FrameSamp
D&R (ours)
RepickCubeReference
Task Goal:
π0.5 (no memory)
FrameSamp
D&R (ours)
DrawPatternImitation
Task Goal:
π0.5 (no memory)
FrameSamp
D&R (ours)

BibTeX

@misc{anonymous2026divide,
  title  = {Divide-and-Remember: Recursive Action-Relevant Memory for Long-Horizon VLA Policies},
  author = {Anonymous},
  year   = {2026},
  note   = {Preprint}
}