Figure 1 – Experimental setup for autonomous dump pocket cleaning using imitation learning. Top left: teleoperated dataset (162 demonstrations). Top right: physical testbed with dual RGB cameras (top-view, wrist-mounted), dump pocket, and SO-ARM100 platform (follower and leader arms). Bottom: imitation learning pipeline receiving observations (camera streams, task prompt*, proprioceptive state), processing through policy network, and executing actions in closed loop. *Task prompt: “use the scoop-like end effector to gather material from the dump pocket and deposit it into the primary crusher at the center”, used only by VLA models (π0.5, SmolVLA) for language conditioning; ACT and Diffusion operate on vision and proprioception only. of the dump pocket, and deposits the material. This cycle repeats continuously to maximize removal rate. Success criteria: We define success operationally as (i) fully autonomous operation without human intervention during the 10-minute session, (ii) no safety-limit violations, and (iii) measurable material removal. All methods start from the same initial state (3 kg material mass), and the primary performance metric is total material removed (g/10 min). 3.2 Observations and Action Space We define the observation ot as the concatenation of two RGB streams and proprio- ception: • Wrist camera: 2MP camera (30 FPS) mounted on the follower arm end-effector (bucket), providing close-range perspective of material and manipulation area. • Overhead camera: Webcam providing global context view of the dump pocket and
RkJQdWJsaXNoZXIy MTM0Mzk2