Track 1: AI and Data-Driven Decision Making

LSTM-Reinforcement Learning Hybrid Approach for P80 Forecasting and Control in Grinding Circuits *C.A. Gomez Paredes1, Y.V. Gomez Sucasaca2 1Master’s in Data Science Graduate, Universidad Católica San Pablo, Arequipa, Peru, (cesar.gomez.paredes@ucsp.edu.pe) 2Master’s Student in Data Science, Universidad Católica San Pablo, Arequipa, Peru (yahaira.gomez@ucsp.edu.pe) ABSTRACT Comminution in ore grinding circuits represents a critical, energy-intensive stage where product size distribution P80 instability compromises downstream recovery, equipment wear, and energy efficiency. Conventional control strategies are predominantly reactive and struggle to manage the nonlinear, time-dependent, and stochastic behavior induced by ore variability and fluctuating operating conditions. Aligned with the conference theme of Artificial Intelligence and Data-Driven Decision Making, this research proposes a hybrid framework that integrates Long Short-Term Memory (LSTM) networks as a high-fidelity surrogate environment with offline Reinforcement Learning (RL) to enable predictive and P80 control. The methodology utilizes historical operational data aggregated at a 5-minute temporal resolution. To approximate mill residence time dynamics, a sliding window of six timesteps (30minute lookback) was employed to forecast P80 over a multi-step horizon. Four LSTM architectures were systematically evaluated under a scaled-capacity protocol: a baseline exogenous model, a feature-engineered model incorporating derived variables from metallurgical principles, an autoregressive model with temporal lag analysis, and a hybrid configuration. The hybrid architecture demonstrated superior predictive performance, achieving a coefficient of determination R2 of 0.95 and a Root Mean Square Error (RMSE) of 3.5 microns. This validated LSTM serves as a differentiable transition model within an offline Actor-Critic architecture, allowing the RL agent to learn optimal parameter adjustments without interrupting plant operations. By integrating the validated LSTM as a differentiable surrogate environment within an offline Actor-Critic architecture, this framework enables the derivation of optimal control laws without interrupting plant operations. This metamorphosis converts the LSTM from a passive temporal predictor into an active environment simulator, facilitating a fundamental shift from reactive to proactive P80 regulation. The deployment of this adaptive system ensures P80 stability within prescribed ranges of tolerance, directly improving downstream metallurgical performance while optimizing operational costs and equipment service life. KEYWORDS LSTM, Reinforcement Learning, Grinding Circuit, Process Control, P80 Forecasting, Mineral Processing, Autoregressive Model, data-driven modeling.

RkJQdWJsaXNoZXIy MTM0Mzk2