research addresses this gap by integrating Lane’s fundamental economic theory with the computational power of Reinforcement Learning, enabling real-time policy adjustments that traditional MILP (Mixed-Integer Linear Programming) models cannot achieve under high uncertainty. 2. METHODOLOGY The methodology begins with the development of a mathematical model based on Lane’s Theory (Lane,1988), which represents the geological and operational constraints of the mine, such as processing and mining capacities and serve like modeled environment. In parallel, a RL agent is trained to interact with this modeled environment. The agent receives rewards based on the economic and operational impact of its decisions, allowing it to gradually improve the production plan through the training process. To better understand, we will briefly explain each finding. First, Kane Lane (Lane,1988) propose the Equation 1. The purpose of the equation is to determine the maximum Net Present Value (NPV) of the mining project and optimal cut‑off grade in the R space. ′ = ∫ max ∈Ω { − 0 ∙ } (1) In his 1988 book Economic Definition of Ore, Lane (Lane,1988) states that the way to solve this problem is through dynamic programming where for each infinitesimal tonnage increment, the integrand selects the largest economic margin across all policies g. Integrating over R yields cumulative NPV sensitivity to incremental cut‑off changes. On the other hand, Richard Sutton and Andrew Barto integrate Bellman equations and Markov Chains to mathematically define the optimal policy in Reinforcement Learning, as shown in Equation 2. This approach aims to teach a machine to perform tasks by training it within a given environment, with specific actions and rewards. The optimal value equals expected immediate reward plus discounted best future return. ∗( , )=∑ ( +1, +1| , ) =1 [ +1 + max +1 ∗( +1, +1)] (2) Based on master’s thesis in Computer Science (Lozano,2024), the mining process was modeled as a Reinforcement Learning (RL) problem. The study demonstrated that Lane’s integral presented in Equation 1 can be addressed by reformulating the mining optimization challenge through RL framework as shown in Equations 3 and 4.
RkJQdWJsaXNoZXIy MTM0Mzk2