Track 1: AI and Data-Driven Decision Making

Table 2 – Predictor features used for LightGBM forecasting Category Feature Spatial x, y, distance_sensor Hydrological rainfall, rain_lag1 rain_lag2, rain_24_sum, rain_36d_sum, rain_60d_sum, rain_96d_sum Kinematic velocidad_media_3c, vel_change, acceleration Quality sigma_u, sigma_u_rolling_3 Thermal t_avg This feature set provides each LightGBM model with information on both the internal deformation state and external forcing such as seasonal variables. Combining deformation memory with rainfall and temperature helps data-driven models learn how recent kinematics and hydro-meteorological conditions translate into short-term deformation responses (Ge et al., 2024; Zhu et al., 2022). 2.4 Model training, validation, and evaluation Four LightGBM models were developed, one per dam zone, using the corresponding time series and engineered predictors. Each training sample corresponds to a point–epoch pair at acquisition time , with the target defined as the subsequent increment Δ ( + 1) = ( + 1) − ( ). All predictors were constructed using information available up to time to prevent look-ahead bias. Additionally, to assess the contribution of meteorological forcing to forecast performance, parallel model configurations were tested with and without rainfall and temperature variables. 2.4.1 Chronological data partitioning and split selection To preserve the temporal structure of the forecasting problem and avoid temporal leakage, data partitioning was performed chronologically rather than through random sampling. Four candidate train-validation-test splits were assessed. For each split, the earliest observations were used for training, the intermediate observations for hyperparameter tuning and early stopping, and the latest observations for independent testing. The selected split was the one that presented the low testperiod MAE with a limited train-test performance gap, thereby reducing the likelihood of overfitting and improving confidence in model generalisation. 2.4.2 Target scaling and weighted learning To improve numerical sensitivity to small deformation increments, the target variable was expressed in millimetres during training. In addition, we applied sample weighting to reduce under-representation of transient deformation spikes within the optimisation objective, because of into the four areas some points present almost a flat deformation and others show stronger deformation, giving a wide range of variability. Specifically, higher weights were assigned to epochs with larger-magnitude observed increments, so that the learner is penalised more strongly when mis-predicting potentially operationally relevant deviations. Weights were capped to prevent instability from a small number of extreme samples. 2.4.2 Hyperparameter optimization via Optuna Hyperparameters were tuned using Bayesian optimisation in Optuna with a Tree-structured Parzen Estimator sampler. For each zone-specific model, 30 trials were performed to optimise the learning rate (0.01-0.10), number of leaves (20-60), minimum data in leaf (5-20), and L1 regularisation

RkJQdWJsaXNoZXIy MTM0Mzk2