arm workspace. • Proprioception: Joint angles (6-DoF from STS3215 servos) and gripper state, syn- chronized with camera timestamps at 30 Hz. • Task prompt (VLA only): Natural language instruction: “use the scoop-like end effector to gather material from the dump pocket and deposit it into the primary crusher at the center”, used by π0.5 and SmolVLA for multimodal conditioning; ACT and Diffusion do not use language input. The action at is a vector of joint-position targets qtarget predicted by the policy and executed by the low-level servo controller. Actions are produced at 30 Hz control frequency. ACT uses action chunking with horizon H = 100 steps; π0.5, SmolVLA, and Diffusion Policy operate in receding-horizon mode with model-specific temporal windows as configured in the training library [9]. Table 1 summarizes the platform specifications and dataset statistics. Table 1 – Platform and dataset summary. Platform Dataset SO-ARM100, 6-DoF, bucket 0–18 g 162 eps (129/33 split), 73,944 frames Sub-$250, leader-follower 30 FPS, mean 15.18 s/ep (range 10.4–26.6 s) Dual RGB (wrist + overhead) 40.99 min total, expert rate 62.04 g/min 3.3 Models The control is governed by four imitation learning paradigms: ACT (Action Chunk- ing with Transformers) uses a CVAE to predict action chunks trained from scratch with an L1 + λ KL objective [7, 8]; π0.5 (the ∼3.7B-parameter VLA Hugging Face checkpoint con- figuration evaluated in this study) employs hierarchical inference, with high-level subtask prediction followed by flow-matching action generation, pretrained on heterogeneous multi-source data (household manipulation, cross-embodiment, web data) [12, 11]; SmolVLA (450M-parameter compact VLA) uses a flow-matching Action Expert pretrained on 481 open-source community manipulation datasets [13]; and Diffusion Policy formulates action generation as conditional denoising diffusion via an MSE noiseprediction objective with iterative denoising at inference [14]. 3.4 Experimental Protocol Baseline (Expert Teleoperation): Expert human teleoperation serves as the gold standard. Across 162 demonstration episodes (40.99 min total), the expert achieved 62.04 g/min, yielding an estimated 620 g over a 10-minute session, the performance upper bound under current hardware constraints. Evaluation protocol: Each policy was evaluated in a single 10-minute continuous autonomous session (3 kg material mass, fixed lighting). The primary metric is total material
RkJQdWJsaXNoZXIy MTM0Mzk2