Track 1: AI and Data-Driven Decision Making

web) may provide richer priors than SmolVLA’s 481 community manipulation datasets, neither of which includes mining excavation data. These factors are consistent with the observed ranking, but isolating their individual effects requires replicated testing and targeted ablations. Deployment barriers: (i) Material variability: bentonite testbed versus hetero- geneous ore (dust to boulders); (ii) Scale: testbed 1000× smaller than industrial hoppers (0.016 vs 10–50 m3); (iii) Environment: dust can degrade LiDAR-based perception in mining environments [17], while poor lighting and vibration remain additional operational concerns; (iv) Control responsiveness: Diffusion Policy exhibited visibly slower reactive be- havior compared to other methods, consistent with the iterative denoising inference process requiring multiple forward passes per action; and (v) Data acquisition: scaling to industrial embodiments requires new demonstrations on the target hardware; the algorithmic workflow may transfer, but final deployment data must be collected at the relevant scale and hardware. This lab-to-mine gap is the primary limitation. Task-state ambiguity and visual activation: ACT, SmolVLA, and Diffusion Policy exhibited hesitation at rollout start: the initial observation closely resembles the terminal state (same rest pose, similar material distribution), creating visual ambiguity between “task just started” and “task finished.” π0.5 did not exhibit this, likely due to language conditioning. A practical mitigation is a visual activation cue (e.g., a color indicator at rollout start) to help vision-only policies distinguish task states without architectural changes. Dataset scale: The 129-episode training split is modest by IL standards; scaling to several hundred demonstrations with greater material-state variability would likely im- prove performance across all architectures and remains a prerequisite before drawing firm conclusions about their relative upper bounds. 6. CONCLUSIONS AND IMPLICATIONS FOR INDUSTRY This work introduces an experimental testbed and open benchmark for evaluating imitation learning architectures on mining excavation tasks under controlled laboratory conditions. In a single preliminary 10-minute evaluation session, π0.5 removes 404 g (65% of the expert teleoperation estimate) while ACT, SmolVLA, and Diffusion Policy remove 169, 113, and 57 g respectively. A notable observation is that ACT, trained from scratch on 162 demonstrations, outperforms SmolVLA despite its pretraining on community manipulation datasets in this single-session benchmark. This raises the hypothesis that for narrow single-task settings where the target domain differs from pretraining data, learning from scratch may remain competitive with fine-tuning a pretrained model. These preliminary results establish an experimental foundation and open dataset to support future research on autonomous tasks for the mining industry. The primary practical implication is safety: these results demonstrate laboratory feasibility of autonomous material-removal behavior in a simplified dump-pocket analogue, addressing the engulfment risk that manual clearing poses to workers in confined hoppers and crusher areas [4, 5, 2]. Substantial barriers remain before industrial deployment: the testbed uses bentonite rather than heterogeneous ore, operates at 1000× smaller scale (0.016

RkJQdWJsaXNoZXIy MTM0Mzk2