plot comparing major-axis measurements in XPL and PPL images (right). 4. DISCUSSION 4.1 Data acquisition, preprocessing, and augmentation Images constitute the input data. Therefore, they must be acquired under controlled and reproducible conditions. Despite their availability, insufficient preparation and standardization limit the development of efficient models in Earth Sciences (Sun, et al., 2022). In this study, images from the BGS and The Open University were used, characterized by high quality, low noise, high resolution, and adequate contrast. Original images can reach sizes of up to 4 GB (The Open University, s.f.), which significantly increases storage requirements and computational demand during model training. For this reason, the images were converted to JPG format, assuming a potential partial loss of information and that moderate to high compression may affect model performance (Ehrlich et al., (2021). Regarding the color system, RGB is known to be sensitive to illumination variations, which has motivated the use of HSI in previous studies (Thompson et al., (2001). However, other studies have reported no substantial differences between RGB and HSI, highlighting the need for further evaluation (Baykan & Yılmaz, 2010). Data augmentation is applied to improve model generalization without altering the fundamental properties of the objects. Considering the constraints inherent to polarized light microscopy (Castro, 2015), controlled photometric and geometric augmentations were applied, including brightness adjustments between -2 and 2%, slight blur (0-1 px), and 90° rotations. These processes do not produce significant adverse effects (Tatar et al., (2025). 4.2 Mineral identification using Detectron2-v0.6 To the author’s knowledge, Detectron2-v0.6 has not previously been applied to mineral segmentation of intrusive igneous rocks; therefore, no reference metrics are available for comparison. The results strongly depend on data quality and model configuration and could be improved through more advanced models, computational capacity, and larger datasets. The comparison between the 3X and 10X datasets indicates that increasing data volume did not lead to significant improvements. The best performance was obtained with the XPL-10X model, showing lower total loss and precision of up to 35% for large objects, whereas PPL-10X yielded the poorest results. Visual evaluation reveals segmentation and overlap issues; however, good performance is observed in gabbros. Although these results were assessed based on model metrics, Vasilev (2019) notes that a common criticism of neural network systems is the interpretation of their outputs. Such models are often regarded as black boxes with logic that is difficult to interpret, leading to reliance on algorithms that are not fully understood and making improvements challenging. Enhancing the explainability of AI models is therefore an important step toward enabling users to understand and trust model outputs. 4.3 The scikit-image image processing library
RkJQdWJsaXNoZXIy MTM0Mzk2