0.99 and an actor learning rate of 3×10⁻⁴. The approach does not incorporate explicit mechanisms to mitigate out-of-distribution actions. Table 2 – Differentiable Actor-Critic architecture. Parameter Specification Input layer InputLayer – Shape: (6, 9) → 30 minutes of history (9 features) Hidden Layer 1 LSTM (64 units), activation = 'tanh' Hidden Layer 2 LSTM (32 units), activation = 'tanh' Dense Layer 1 Dense (16 units), activation = 'ReLU' Output Layer Dense (2 units), activation = 'linear' → predicts [P80] 4. RESULTS AND DISCUSSION This section presents the comparative results of the four LSTM-based modeling strategies evaluated for short-term P80 forecasting and subsequent integration with the reinforcement learning (RL) agent in an industrial grinding context. Firstly, Figure 4 illustrates the training and validation behavior of these two top-performing models, highlighting their convergence characteristics and comparative accuracy. Figure 4 – Training Loss Evolution for Autoregressive and Hybrid LSTM Models. Table 3 summarizes the predictive performance obtained for the 5-minute forecasting horizon. All four LSTM-based approaches achieved satisfactory metrics, indicating their ability to capture the short-term dynamics of P80 behavior in an industrial grinding environment. The baseline multivariate model provides a solid reference, while the feature-enhanced and autoregressive formulations show progressive improvements in MAE, R² and RMSE, reflecting the added value of incorporating process-informed features and temporal dependency, respectively. Table 3 – LSTM Model Performance at Different Forecast Horizon Horizon = 1 Horizon = 2
RkJQdWJsaXNoZXIy MTM0Mzk2