Given the natural imbalance between the limited number of known deposits and the much larger area without documented mineralization, class distribution was evaluated before model training. To mitigate algorithmic bias toward the majority class, the dataset oversampling techniques were implemented, in order to preserve geological realism and avoid artificial pattern amplification. Model performance was assessed using k-fold cross-validation to evaluate generalization capacity and reduce overfitting. The dataset was partitioned into k subsets, iteratively training on k−1 folds and validating on the remaining fold. Performance was evaluated using Precision, Recall, and F1score metrics, providing a balanced assessment of classification behavior. Although this procedure improves robustness, it is acknowledged that spatial autocorrelation in geological datasets may inflate predictive metrics. Consequently, results are interpreted as relative indicators of permissivity rather than deterministic predictions of mineral occurrence. Structural influence was represented through Euclidean distance rasters derived from mapped fault systems. A 5 km buffer was assigned to secondary or local-scale structures, while a 20 km buffer was applied to major trans-lithospheric and transverse faults interpreted as lithospheric-scale magma ascent corridors. The larger buffer distance reflects the regional structural architecture associated with Andean porphyry–skarn districts and is intended to represent influence gradients rather than strict geometric thresholds. These values were selected to remain consistent with the regional scale of analysis and to avoid overemphasizing deposit-scale controls. All predictor datasets were rasterized and standardized to a uniform spatial resolution consistent with the provincial scope of the study. The selected cell size balances geological representativeness, data resolution, and computational stability, preventing overfitting to local variability while preserving regional-scale patterns. Prior to model integration, predictor rasters were normalized to ensure comparability among heterogeneous geological variables. 3. DISCUSSION The integration of the Mineral Systems framework with machine learning algorithms produced consistent but not identical results between Random Forest and Artificial Neural Networks. Although both models identified broadly similar permissive domains, spatial variations in prospectivity intensity reflect inherent differences in how each algorithm captures nonlinear relationships and variable interactions. Random Forest tends to emphasize hierarchical feature importance and threshold-based partitioning, whereas Neural Networks model more diffuse nonlinear interactions, leading to variations in spatial probability gradients. Importantly, prospectivity outputs derived from machine learning differ from traditional targeting approaches based solely on remote sensing alteration mapping or predefined metallogenic belts associated with Eocene–Oligocene intrusions. While conventional methods emphasize surface alteration signatures or known metallogenic corridors, the Mineral Systems–AI integration evaluates spatial co-occurrence of permissive geological conditions, potentially identifying domains beyond historically recognized exploration trends. This divergence does not imply
RkJQdWJsaXNoZXIy MTM0Mzk2