Track 1: AI and Data-Driven Decision Making

Table 2 – Benchmark results: total material removed in a single 10-minute autonomous session. Expert baseline estimated from demonstration data (62.04 g/min × 10 min). All learned policies evaluated under identical initial conditions (3 kg bentonite, fixed lighting). Given N=1 session per model, statistical aggregation (mean±std) is not applicable; values represent single observed measurements. Model Remove d (g / 10 min) Rat e (g/mi n) of Exper t (est.) Expert Teleoperation (est.) ≈620 62.0 100 π0.5 404 40.4 65 ACT 169 16.9 27 SmolVLA 113 11.3 18 Diffusion Policy 57 5.7 9 removal performance (404 g in 10 minutes, 40.4 g/min), reaching approximately 65% of the expert teleoperation estimate (620 g). ACT (169 g, 27% of expert) ranks second despite having no pretrained representations, trained solely on 162 demonstration episodes. SmolVLA (113 g, 18% of expert) ranks third, removing less material than ACT despite being pretrained on 481 open-source community manipulation datasets. This single-session result raises the hypothesis that for narrow single-task settings where the target domain differs from pretraining data, learning from scratch may remain competitive with finetuning a pretrained model. Diffusion Policy (57 g, 9% of expert) removes the least material, consistent with the slower reactive behavior observed during rollout. 5. DISCUSSION The benchmark demonstrates laboratory feasibility of learned policies for au- tonomous material removal in a simplified dump-pocket analogue. A key interpretive hypothesis is that pretraining benefit may depend on domain proximity and model capacity: the π0.5 Hugging Face checkpoint configuration evaluated in this study (3.7B parameters, broadly pretrained across diverse robot embodiments) achieves the highest observed removal rate in this novel setting, while ACT trained from scratch outperforms SmolVLA (450M parameters, pretrained on 481 community manipulation datasets) in the singlesession benchmark. This result suggests a hypothesis for future replicated testing: in narrow single- task settings where the target domain differs from pretraining data, learning from scratch may remain competitive with fine-tuning a pretrained model. Results reflect simplified labo- ratory conditions (bentonite, fixed lighting, N=1 session per model) that may not stress-test reliability differences emerging under operational challenges. Model performance analysis: π0.5’s leading observed performance may be related to several factors: (i) scale: the evaluated ∼3.7B-parameter Hugging Face checkpoint configuration was trained for 25,774 steps, the most compute-intensive configuration in this study; (ii) hierarchical reasoning: it decouples high-level subtask prediction from lowlevel action generation, which may support adaptation to varying material states; and (iii) pretraining breadth: its heterogeneous multi-source data (household, cross-embodiment,

RkJQdWJsaXNoZXIy MTM0Mzk2