• • The SDTW-IPAM clustering partitions load into double-peak, high-peak, and smooth patterns, with Gap statistics and K-means++ initialization improving stability; this enables targeted forecasting and reduces error for volatile patterns by up to 15% compared to single-model approaches.
• • The Informer model with probabilistic sparse attention and self-attention distillation achieves lower EMAE and ERMSE than Autoformer, FEDformer, and CNN-LSTM-Attention, with R² improvements of 0.05-0.12 on high-volatility loads, directly enhancing scheduling reliability.
• • MIC-based feature selection identifies key nonlinear drivers per cluster, reducing input dimensionality by 30% while maintaining accuracy, which lowers computational overhead for real-time deployment.
• • Validation on Urumqi load data shows the combined model reduces EMAE by 22.4% and ERMSE by 19.8% versus the best baseline, demonstrating industrial viability for grids with high renewable penetration.