Track 1: AI and Data-Driven Decision Making

control by deploying Long Short-Term Memory (LSTM) networks as short-horizon, multi-step P80 forecasting models trained on historical multivariate sensor data. The LSTM captures the dynamic, nonlinear behavior of the grinding process and serves as a surrogate environment for offline reinforcement learning (RL), enabling an agent to identify optimal adjustments of controllable process variables without interrupting plant operations. Together, the LSTM–RL framework provides a foundation for adaptive, resilient grinding circuit control. 2. STATE OF THE ART Grinding is widely recognized as a core operation in mineral processing plants, and the P80 has long served as the critical product size indicator directly linking circuit performance to downstream metallurgical outcomes. Hodouin (2011) established the foundational taxonomy of process variables (manipulated, disturbance, internal state, and controlled) that underpins the variable selection and control formulation adopted in this work. As machine learning became integrated into predictive grinding models, Both and Dimitrakopoulos (2021) demonstrated that neural networks can capture nonlinear relationships between mill power, ore hardness, and throughput, with P80 used as a predictor variable. Their formulation addresses the inverse of the problem treated here, where P80 is the output to be forecast from operational inputs and its focus on long-term performance rather than short-horizon dynamics underscores the need for approaches designed specifically for real-time P80 prediction. Saldaña et al. (2023) extended this line of work to SAG mill optimization using Random Forest, XGBoost, and neural networks, achieving up to R² = 0.89. However, these non-temporal models operate on process snapshots and are structurally incapable of exploiting the strong autocorrelation of P80 at short sampling intervals, a fundamental limitation that motivates the adoption of recurrent neural networks. Zeb et al. (2024) demonstrated that LSTM networks achieve accurate short-horizon forecasts of dynamic metallurgical variables (specifically flotation concentrate grade) in multivariate industrial time-series over windows up to 30 minutes. The temporal dependencies governing concentrate grade are structurally analogous to those of P80 in a closed circuit, providing direct methodological justification for the autoregressive LSTM design adopted here, in which the lagged P80 value is included as an additional input feature. Quintanilla et al. (2024) developed the closest existing architecture to the present framework: a digital twin coupling a recurrent neural network surrogate with a real-time control system for a closed SAG mill circuit. The key distinctions are that their control policy is rule-based rather than learned, P80 is not the primary forecasting target, and the control problem is not formulated as a sequential decision-making task. Nian et al. (2020) established reinforcement learning as a principled paradigm for industrial process control, identifying online exploration cost as the central practical barrier and motivating offline RL from historical data. Moerland et al. (2023) further demonstrated that differentiable learned surrogates can serve as transition models for policy gradient training, enabling backpropagation through the surrogate without any real plant interaction. While prior works have demonstrated the value of machine learning and intelligent optimization for improving mill performance, no previous study has addressed the combined problem of forecasting P80 as a dynamic autoregressive time-series variable using LSTM networks and using the resulting surrogate model as the environment for an offline RL control agent trained from historical plant data. In this context, Long Short-Term Memory networks are adopted due to their ability to capture temporal dependencies and nonlinear behavior inherent to grinding

RkJQdWJsaXNoZXIy MTM0Mzk2