removed, measured by weighing the outlet collection tray before and after each session. Given N=1 session per model, we report observed total removed mass and equivalent removal rate (g/min) without statistical aggregation. Offline training: All models were trained on the 129-episode dataset using an open-source IL library [9]. Best checkpoints selected by validation loss: ACT at step 9205 (val loss 0.2746), π0.5 at step 25774 (0.0745), SmolVLA at step 14728 (0.0254), Diffusion at step 7252 (0.0141). VLA models (π0.5, SmolVLA) leverage pretrained backbones [12, 13]; ACT and Diffusion trained from scratch. Since ACT optimizes a composite objective (L1 + λ KL) while other policies use MSE/flow-matching losses, Figure 2(a) plots ACT as (val/loss)2 for visual scale comparability (monotone transformation; checkpoint ranking preserved); this panel is a convergence diagnostic, not a cross-architecture ranking. 4. RESULTS, OUTCOMES, OR PERFORMANCE Results are summarized in Figure 2 and Table 2. All models were evaluated in a single 10-minute session under identical initial conditions (3 kg bentonite, fixed lighting, indoor laboratory setting). Figure 2 – Benchmark results and offline training metrics. (a) Validation loss trajectories; ACT is plotted as (val/loss)2 for visual scale comparability (see Section 3 for justification). Best checkpoints (stars): ACT 0.0754 @ 9,205 steps; π0.5 0.0745 @ 25,774 steps; SmolVLA 0.0254 @ 14,728 steps; Diffusion 0.0141 @ 7,252 steps. This panel is a convergence diagnostic; cross-architecture comparison uses real-robot results only. (b) Total material removed (g) in a single 10-minute autonomous session. Expert baseline (≈620 g) estimated from demonstration data (62.0 g/min, 2542 g over 40.99 min). π0.5 removes 404 g, followed by ACT (169 g), SmolVLA (113 g), and Diffusion Policy (57 g). Note on statistical power: This evaluation reports single observed values per model (N=1 session); statistical aggregation is not applicable. Results represent exploratory observations pending systematic replication. Failure observations: Qualitative rollout observations included partial scoops, wall adhesion, bucket slip, and occasional hesitation; these were not quantified in this preliminary evaluation. Key findings: Under controlled laboratory conditions, π0.5 achieves the highest
RkJQdWJsaXNoZXIy MTM0Mzk2