Recovering the View: Benchmarking Physical Active Vision for Occlusion Recovery in Robotic Manipulation

Kaijun Luo1, Yudi Huang2, Qijun Zhong1, Xinshuai Song1, Yang Liu1,4†, Liang Lin1,3,4
1 Sun Yat-sen University 2 University of Electronic Science and Technology of China 3 Pengcheng Laboratory 4 X-Era AI Lab

† Corresponding author

Overview of physical active vision, the bimanual benchmark, and the A-FAR recovery policy
Overview of our motivation, benchmark, and approach. (a) We study physical active vision as closed-loop recovery from external occlusion: using a single active camera as the sole visual sensor, the robot actively changes its viewpoint to recover task-relevant visibility and continue manipulation. (b) We introduce BAVO-Bench, a bimanual active-vision manipulation benchmark that systematically introduces controlled external occlusions to study viewpoint recovery, together with a VR teleoperation pipeline for collecting active-view demonstrations. (c) We present A-FAR, an active-vision policy that canonicalizes moving-view observations in robot-centric 3D and distills D4RT's 4D relational priors, enabling stable, future-aware control under active viewpoint changes.

Abstract

Physical active vision allows robots to change their viewpoint when task-relevant observations become unreliable, yet existing manipulation benchmarks provide limited support for studying how policies recover from occlusion during execution. We introduce BAVO-Bench (Bimanual Active Vision under Occlusion), a bimanual active-vision benchmark that systematically controls external visibility through Clean, Stage Occlusion, and Random-time Occlusion conditions, enabling evaluation of both manipulation performance and active visual recovery. Building on this setting, we present A-FAR (Active Future-Aware Recovery), an active-vision policy for joint viewpoint and manipulation control. A-FAR represents moving-camera observations in a unified robot-centric 3D frame and distills relational structure together with its future evolution from a pretrained 4D model, providing the policy with future-aware geometric guidance without requiring future observations at deployment. Experiments across multiple bimanual manipulation tasks show that A-FAR improves robustness to both structured and temporally shifted occlusions while maintaining strong performance under clean observations.

Benchmark

Five bimanual manipulation tasks and examples of stage and random-time occlusion
BAVO-Bench overview. Top: Five multi-stage bimanual manipulation tasks. Bottom: Representative rollouts under the two occlusion settings. Red-bordered frames mark occlusion events, and the timelines illustrate recovery and manipulation behaviors. Both settings occlude task-relevant targets; Stage Occlusion aligns interventions with predefined task transitions, whereas Random-time Occlusion samples intervention times independently of these transitions.

A-FAR

A-FAR architecture with robot-centric 3D observations, a training-only 4D teacher, and a future-aware action policy
Overview of A-FAR. Current RGB-D observations are canonicalized into a robot-base point cloud and encoded as current point tokens Ht. A state-conditioned WorldQueryFormer predicts future relational tokens ĤtF, which are fused with the current representation for ManiFlow action generation. During training, a frozen D4RT teacher queries the same source points at current and future target times to construct relational targets that supervise the current structure (ℒcur) and its temporal evolution (ℒevo). The teacher branch and future frames are removed at deployment.

Simulation Rollouts

Explore the five BAVO-Bench tasks under clean observations, stage-aligned occlusion, and randomly timed occlusion. Each video shows the third-person scene and the active-camera view side by side.

Task
Visibility condition

Cube Handoff

Handoff a cube between arms and place it on the target plate.

Random-time Occlusion introduces visibility disruptions at sampled times during task execution.

Physical Deployment

Four frames showing a physical robot manipulation sequence
Real-world deployment of A-FAR. Representative rollout under an externally introduced occlusion. From left to right, task-relevant visibility is disrupted during execution, the active camera changes viewpoint to recover the scene, and manipulation proceeds after the target region becomes observable again. The sequence demonstrates closed-loop active-view behavior on our physical manipulation platform.

Physical Robot Rollouts

Watch three physical robot tasks with or without an externally introduced occlusion.

Rollout

Rack Cleaning

A physical robot rollout with an externally introduced occlusion.

BibTeX

@misc{luo2026recoveringviewbenchmarkingphysical,
  title={Recovering the View: Benchmarking Physical Active Vision for Occlusion Recovery in Robotic Manipulation},
  author={Kaijun Luo and Yudi Huang and Qijun Zhong and Xinshuai Song and Yang Liu and Liang Lin},
  year={2026},
  eprint={2609.37292},
  archivePrefix={arXiv},
  primaryClass={cs.RO},
  url={https://arxiv.org/abs/2609.37292},
}