grade multihorizon forecasting and what-if scenario simulation while providing uncertainty estimates through predictive distributions. 5.1 CVAE ARCHITECTURE The representation model is a convolutional variational autoencoder (CVAE) (figure 1) trained to encode the short-term process state into a low-dimensional latent space. Inputs are 15minute windows sampled at 1-minute resolution, arranged as a tensor of shape 15-time steps × 13 observation variables. Observation variables correspond to process states/outcomes (e.g., throughput, pressures, power, fill level, and size-fraction measurements) and are used exclusively in this stage so that the learned representation captures the nonlinear process state without directly incorporating operator setpoints. The encoder consists of residual convolutional blocks that capture local temporal dependencies and multivariable coupling. It parameterizes a diagonal Gaussian posterior ( ∣ ) = ( , )with latent dimensionality =8 (i.e., ∈ℝ8and log ∈ℝ8). The decoder mirrors the encoder using transposed convolutions to reconstruct the observed window, producing ̂with the same dimensionality as the input. Training optimizes a β-ELBO (β-VAE objective), where the Kullback–Leibler (KL) divergence term is weighted to control the trade-off between reconstruction fidelity and latent regularization: ℒ( , ) = ( ∣ ) [−log ( ∣ )] + ( ( ∣ ) ∥ ( ))…( 1) In practice, using ≠1(with warm-up when applicable) improves training stability and mitigates posterior collapse, yielding a latent space that remains informative for the downstream dynamics model. The CVAE is assessed using reconstruction error and the magnitude of the KL term on train/validation and held-out test data to verify consistency and generalization of the learned state representation. Figure 1 – Convolutional VAE representation model: learning a latent process state from 15-min multivariate observation windows.
RkJQdWJsaXNoZXIy MTM0Mzk2