autoregressive structure is present. At the 10-minute horizon, the autoregressive (model 3) retained R² = 0.90 and MAE = 2.88 µm, confirming robust generalization across both forecast steps. The autoregressive model (Model 3) was selected as the surrogate environment for offline RL training on the basis of this analysis. Its lower input dimensionality (nine features versus seventeen for the hybrid) reduces the state space complexity for the agent without sacrificing predictive fidelity, as the near-identical RMSE values at both horizons (3.53 µm vs. 3.56 µm at t+1; 4.82 µm vs. 4.82 µm at t+2) confirm. This selection reflects a broader design principle: model complexity should be justified by measurable performance gains, not by the availability of additional features The Differentiable Actor-Critic (DAC) agent trained on historical transitions demonstrated that model-based offline policy learning through a frozen LSTM surrogate is technically feasible. The mean discounted return improved from −32.83 at iteration 1 to +1.94 at iteration 300, an absolute gain of +34.77 units. The sign change from negative to positive return has a direct physical interpretation: a positive return under the reward function is only achievable when the policy spends a meaningful fraction of rollout steps within the ±10 µm tolerance band around the 160 µm setpoint. This transition first occurred at iteration 200, indicating that the policy had learned the direction in action space that moves P80 toward the target from an initial deviation of approximately 23.4 µm. Gradient flow from the reward signal through the LSTM to the actor was numerically verified, confirming that the surrogate provides a valid and informative training signal for policy optimization without any plant interaction. Beyond the specific quantitative results, this work establishes the viability of converting a high-fidelity autoregressive LSTM into a differentiable surrogate environment for offline RL, enabling gradient-based policy learning from historical plant data without interrupting operations. Finally, this work provides quantitative evidence that an autoregressive LSTM achieves the predictive accuracy required for P80 forecasting in an industrial grinding circuit, and that this surrogate can serve as a differentiable training environment for offline reinforcement learning. Together, these two components establish a technically validated foundation for shifting grinding circuit management from reactive manual adjustment to predictive, data-driven control, with the trained DAC providing directionally correct policy learning as a first-stage result and more advanced offline agents representing a well-defined path toward deployment-ready closed-loop operation ACKNOWLEDGEMENTS We gratefully acknowledge Pueblo Viejo and Cerro Verde mine sites, where we had the opportunity to work professionally and gain invaluable practical experience and technical insight into industrial grinding operations. This hands-on experience was fundamental in shaping the applied perspective of this research. We also sincerely thank San Pablo Catholic University and the professors of the Master’s Program in Data Science for their academic guidance, knowledge, and continuous support, which were essential to the successful development of this work.
RkJQdWJsaXNoZXIy MTM0Mzk2