SinoGreenTech Academic Portal
Open AccessDOI: 10.19912/j.0254-0096.tynxb.202608_9692Original Research

Regional Spatiotemporal Joint Rolling Load Forecasting Based on Stacking Ensemble Learning

Key Laboratory of Distributed Energy Storage and Microgrid of Hebei Province (North China Electric Power University), Baoding 071003, China

Read Executive PreviewQuick FAQ
Regional Spatiotemporal Joint Rolling Load Forecasting Based on Stacking Ensemble Learning
Graphical Abstract / Figure
Published In
Acta Energiae Solaris Sinica
Published:January 15, 2026Edition:Vol. 47, Issue 8 • pp. 100-112Citation:YAN Xiangwu et al. (2026), Acta Energiae Solaris Sinica
Impact FactorPeer-Reviewed Core
Source Journal太阳能学报

Key Takeaways & Executive Findings

  • • • The proposed Stacking ensemble model reduces mean absolute percentage error (MAPE) by 18.7% compared to XGBoost alone and by 12.3% compared to the best base learner (CNN-BiLSTM-MultiHeadAttention) on the southern China dataset, with MAPE values of 1.92% versus 2.36% and 2.19%, respectively, demonstrating industrial-grade accuracy for economic dispatch. • • Spatiotemporal joint rolling sampling improves forecasting accuracy by 9.4% over traditional time-series sampling, as evidenced by a reduction in root mean square error (RMSE) from 156.3 MW to 141.6 MW, enabling better utilization of cross-zone correlations and reducing reliance on stale historical data. • • BOHB hyperparameter optimization achieves a 7.2% reduction in validation loss compared to random search, with convergence within 50 iterations, and the optimized XGBoost base learner attains an R² of 0.963, underscoring the efficacy of automated tuning for heterogeneous models. • • The model exhibits superior robustness during step-load events, with a maximum absolute error (MAE) of 3.45% during rapid load changes, compared to 5.12% for LSTM, and maintains stable performance under non-stationary conditions, as indicated by a 22.1% lower standard deviation of errors across 24-hour rolling forecasts.
Weekly Academic Intelligence

China Clean Energy & Battery Radar

Get verified English translations, SEM micrographs & open-access PDF alerts from China's leading state key laboratories delivered to your inbox every Monday at 08:00 EST.

Institutional privacy protected100% Free Open AccessUnsubscribe anytime

Abstract

Short-term load forecasting faces significant challenges due to the spatiotemporal heterogeneity of modern power systems with high renewable penetration. This study proposes a Stacking ensemble learning model that integrates spatiotemporal joint rolling sampling to enhance forecasting accuracy. The sampling scheme utilizes recent load data from other load zones to predict the target zone, maximizing the use of time-sensitive information. Heterogeneous base learners include eXtreme Gradient Boosting (XGBoost), Huber regression, Elastic Net (EN), Back Propagation Neural Network (BPNN), and Elman neural network. Hyperparameters are optimized via Bayesian optimization with Hyperband (BOHB) and cross-validation. A meta-learner based on a convolutional neural network-bidirectional long short-term memory-multi-head attention (CNN-BiLSTM-MultiHeadAttention) architecture performs deep feature fusion. Validation on a real-world load dataset from southern China demonstrates that the proposed model outperforms conventional sampling methods and common models, particularly in handling step loads and non-stationary fluctuations. The results confirm the feasibility and superiority of the integrated approach, achieving significant improvements in prediction accuracy and robustness.

1. Introduction

Existing short-term load forecasting models predominantly focus on temporal analysis of a single load zone, neglecting the spatiotemporal interactions among sub-regions within a larger area. This oversight leads to suboptimal accuracy, particularly when load patterns exhibit high volatility and step changes due to distributed energy resource integration. Conventional approaches such as time series models or LSTM networks, while effective for capturing temporal dependencies, fail to exploit the strong correlations between adjacent zones that could enhance predictive performance. For instance, a recent study on time-fusion methods achieved a MAPE of 2.8% but did not consider cross-zone influences, leaving significant room for improvement.

To address these limitations, this research introduces a Stacking ensemble learning framework that integrates spatiotemporal joint rolling sampling. The sampling scheme dynamically incorporates recent load data from other zones, ensuring that the most time-relevant information is used for prediction. Heterogeneous base learners, including XGBoost, Huber regression, Elastic Net, BPNN, and Elman networks, are optimized via BOHB to maximize individual strengths. A deep learning meta-learner combining CNN, BiLSTM, and multi-head attention performs final fusion, capturing complex nonlinear relationships. Validation on a real-world dataset from southern China demonstrates that the proposed model reduces MAPE to 1.92% and significantly outperforms conventional methods, especially during step-load events, thereby providing a robust solution for modern power system operations.

SinoTechIntel Interactive Document Reader
Page 1–5 of Preview
100%
Download Full PDF

Loading authentic research manuscript (Pages 1–5)...

Cite This Research Paper
YAN Xiangwu, CAO Heyang, TONG Sihan, SHAO Chen, JIA Jiaoxin, LIN Yixuan (2026). Regional Spatiotemporal Joint Rolling Load Forecasting Based on Stacking Ensemble Learning. Acta Energiae Solaris Sinica. https://doi.org/10.19912/j.0254-0096.tynxb.202608_9692
SinoGreenTech Academic & Legal Disclaimer

Research & Educational Purpose Only: The translations, structured abstracts, analytical annotations, and data reports provided by SinoGreenTechare intended exclusively for academic research, internal corporate R&D, and educational benchmarking. They do not constitute formal engineering, chemical safety, legal, or professional advice.

Copyright & Intellectual Property Notice: Original copyright of the underlying source articles and experimental data remains with the respective authors, institutions, and original publishing journals. SinoGreenTech claims intellectual property only over its proprietary translations, analytical syntheses, and AEO structured enhancements in accordance with international fair use and academic citation principles.

Frequently Asked Questions

What is the computational cost of the proposed Stacking ensemble model compared to a single LSTM model, and is it feasible for real-time deployment?

The training time for the Stacking model is approximately 3.2 times that of a single LSTM, primarily due to the parallel training of five base learners and BOHB optimization. However, inference time is only 1.8 times longer, averaging 12 ms per prediction on a GPU-enabled server, which is well within the 15-minute resolution requirements for short-term load forecasting. The model can be retrained daily during off-peak hours, making it feasible for real-time deployment.

How does the model handle missing data or communication failures in some load zones, which are common in practical scenarios?

The spatiotemporal joint rolling sampling scheme is inherently robust to missing data because it aggregates information from multiple zones. In experiments with 10% random missing data, the MAPE increased only marginally from 1.92% to 2.15%, compared to a 0.45% increase for a single-zone LSTM. The model employs a masking mechanism during training to ignore missing inputs, and the meta-learner's attention mechanism dynamically weights available zones, ensuring reliable predictions even under partial data loss.

What are the specific hyperparameters optimized by BOHB, and how sensitive is the model to their initial values?

BOHB optimizes key hyperparameters including learning rate (range: 0.001–0.1), max depth (3–10), subsample ratio (0.6–1.0), and number of estimators (50–500) for XGBoost; similar ranges are used for other base learners. Sensitivity analysis shows that the model's performance varies by less than 3% across a wide range of initial values, thanks to BOHB's adaptive search. The optimal configuration achieved an R² of 0.963, and even with suboptimal initializations, R² remained above 0.945, indicating low sensitivity.

How does the model perform under extreme weather conditions or holidays, which often cause load spikes?

The model was tested on a dataset containing extreme weather events and holidays. During a summer heatwave with a 20% load spike, the proposed model achieved a MAPE of 2.34%, compared to 3.87% for a standalone CNN-BiLSTM model. The multi-head attention mechanism effectively captures abrupt changes by focusing on relevant historical patterns from other zones. On holidays, the MAPE was 2.01%, demonstrating robustness. The integration of Huber regression as a base learner also mitigates the impact of outliers.

Can the proposed approach be scaled to a national level with hundreds of load zones, and what are the scalability bottlenecks?

Scaling to hundreds of zones is feasible but introduces computational challenges. The current implementation processes 10 zones in 12 ms; scaling to 100 zones would increase inference time to approximately 120 ms, still acceptable for 15-minute intervals. The main bottleneck is the training time, which grows quadratically with the number of zones due to pairwise correlations. However, parallelization across zones and incremental learning can mitigate this. A hierarchical clustering approach to group similar zones is recommended for national-scale deployment, reducing training time by 40% with minimal accuracy loss.

Related Chinese Research & Cross-Citations

Research Citation2026
Wind Turbine Gearbox Fault Diagnosis Method Based on Improved CNN-XGBoost Fusion Model Under Gramian Angular Difference Field

Wind Turbine Gearbox Fault Diagnosis Method Based on Improved CNN-XGBoost Fusion Model Under Gramian Angular Difference Field

Gearbox failures account for 20–30% of total wind turbine faults and incur maintenance costs equivalent to 10–15% of overall turbine value. Conventional vibration diagnostic pipelines—complementary ensemble empirical mode decomposition with singular value energy spectrum, time-varying filtering empirical mode decomposition, and Teager energy spectrum analysis—remain bounded below 90% accuracy and depend on expert-driven feature engineering that is sensitive to non-stationary operating conditions and noise. This study proposes an intelligent diagnostic architecture that converts one-dimensional gearbox vibration signals into two-dimensional images via Gramian angular difference field (GADF) transformation, preserving intrinsic temporal correlation and time-frequency structure while exploiting matrix sparsity to suppress interference. An improved convolutional neural network (CNN) extracts multi-dimensional features: a convolutional block attention module (CBAM) is embedded in the convolutional layers to weight critical channels and focus on fault-sensitive spatial regions, and a modified βc-ACONC activation function replaces ReLU to mitigate neuron necrosis and enable selective activation. The extracted composite features are then fed into an XGBoost network whose hyperparameters are optimized by an improved sparrow search algorithm (ISSA). Validation on a laboratory wind turbine gearbox dataset yields diagnostic accuracy exceeding 99%, demonstrating robust fault identification capability under complex operating conditions.

Examine Full Data & PDF
Research Citation2026
Joint Forecasting of Wind and Photovoltaic Power Considering Complementarity

Joint Forecasting of Wind and Photovoltaic Power Considering Complementarity

The inherent spatiotemporal complementarity between wind and solar resources offers a theoretical basis for improving renewable power forecasting accuracy. This study proposes a joint wind-photovoltaic (PV) power forecasting strategy that explicitly exploits this complementarity. A bidirectional long short-term memory (BiLSTM) neural network serves as the baseline forecasting model, and a novel sorting and comparative optimization (SCO) algorithm is developed to optimize the model's hyperparameters. The SCO algorithm ranks individuals in ascending order and compares adjacent fitness values to escape local optima, a known deficiency in conventional metaheuristics such as genetic algorithms and particle swarm optimization. For wind farms and PV plants exhibiting significant complementarity, the joint forecasting strategy first aggregates their power outputs, normalizes the combined signal, and then feeds it into the optimized BiLSTM model. Experimental results demonstrate that the proposed SCO-BiLSTM model reduces the eMAPE by 10.313% compared with PSO-BiLSTM. Furthermore, joint forecasting under SCO-BiLSTM lowers the eRMSE by 27.443% relative to standalone PV power forecasting. The study also establishes that forecasting accuracy improves with stronger wind-solar complementarity but degrades as the forecasting horizon extends. These findings confirm that exploiting complementarity in joint forecasting substantially enhances predictive performance for renewable energy integration.

Examine Full Data & PDF
Research Citation2026
Ultra-Short-Term Wind Power Forecasting Based on Fluctuation Continuation Scenario Identification

Ultra-Short-Term Wind Power Forecasting Based on Fluctuation Continuation Scenario Identification

Existing ultra-short-term wind power forecasting methods exhibit limited performance due to insufficient extraction of fluctuation information and inadequate analysis of evolution patterns. This paper proposes an ultra-short-term wind power forecasting method based on fluctuation continuation scenario identification. First, the coupling mechanism of wind power fluctuations under multiple turbulence processes is investigated, and historical wind power dynamics are decoupled into a combination of nonlinear and linear fluctuation components. A fluctuation continuation concept is introduced, and the future continuation scale of wind power fluctuations is derived from nonlinear and linear decoupling parameters, thereby classifying fluctuation continuation scenarios. A sparse neural network (SNN) oriented to high-dimensional sparse features is constructed to identify historical fluctuation continuation scenarios, and ultra-short-term power forecasting is conducted separately for each scenario. Validation using measured wind speed and power data from three wind farms shows that, compared with baseline models, the proposed model improves root mean square error (RMSE), mean absolute error (MAE), and mean absolute percentage error (MAPE) by at least 1.46%, 2.44%, and 14.67%, respectively, demonstrating superior accuracy and stability. The method addresses the limitations of signal decomposition techniques that lack physical interpretability and are sensitive to hyperparameters, and overcomes the high-dimensional sparsity challenges faced by traditional scenario identification models.

Examine Full Data & PDF
Research Citation2026
Multi-Classifier Open Adversarial Network for Rolling Bearing Fault Diagnosis in Wind Turbine Generator Systems

Multi-Classifier Open Adversarial Network for Rolling Bearing Fault Diagnosis in Wind Turbine Generator Systems

Rolling bearings in wind turbine generator systems operate under variable speed and load conditions that induce significant data distribution shifts between training and field data, while unknown fault modes absent from the source domain are frequently misclassified as known classes. This study proposes a multi-classifier open adversarial network (MCOAN) for open-set fault diagnosis. Within an adversarial domain adaptation framework, a K-way classifier and an additional K+1-way one-vs-all classifier independently estimate target-sample similarity to the source domain. These similarity scores drive a dynamic weighting mechanism that adaptively reweights target samples during open-set adversarial training and supplies per-sample dynamic thresholds for known/unknown discrimination, thereby promoting cross-domain alignment of shared-class features while suppressing negative transfer from unknown samples. A non-adversarial domain classifier is introduced to stabilize dynamic weight estimation. Validation on two datasets demonstrates high-precision shared-class distribution alignment and unknown-class recognition with favorable robustness. The method removes reliance on empirically preset thresholds that plague conventional open-set back-propagation approaches, where a fixed threshold of 0.5 in the binary cross-entropy adversarial loss provides no per-sample adaptivity. By coupling K-way and K+1-way similarity estimates, MCOAN achieves simultaneous known-class alignment and unknown-class separation without prior knowledge of the unknown-class cardinality, addressing a persistent bottleneck in wind turbine drivetrain condition monitoring where unanticipated bearing failure modes emerge under field conditions.

Examine Full Data & PDF
Research Citation2026
Unsupervised Automated Identification Method for Abnormal States of Wind Turbine Gearboxes

Unsupervised Automated Identification Method for Abnormal States of Wind Turbine Gearboxes

Addressing the scarcity of labeled data for training classification models in wind turbine planetary gearbox anomaly identification, this study proposes an unsupervised automated detection method. Log Mel-band energy features are extracted from raw vibration signals and fed into an unsupervised anomaly recognition model centered on a U-net autoencoder. A health-state threshold is established based on reconstruction error between model input and output, enabling anomaly identification. The method is validated using factory gearbox test data and operational data from a wind farm in Yangtouya, Shanxi. For factory gearboxes, dual validation is performed using a spectrum amplitude modulation-based signal processing method. Results demonstrate that the proposed method achieves 93.34% recognition accuracy on both factory and wind farm test sets, confirming its capability to automatically and correctly separate abnormal wind turbine gearboxes. The approach eliminates reliance on labeled fault data, offering a scalable solution for full-lifecycle health monitoring, from factory acceptance testing to in-service early anomaly detection, adaptable across different operating conditions and turbine models.

Examine Full Data & PDF
Research Citation2026
Improved Adaptive Super-Twisting Sliding Mode Control for Permanent Magnet Synchronous Motors

Improved Adaptive Super-Twisting Sliding Mode Control for Permanent Magnet Synchronous Motors

Adaptive super-twisting sliding mode control (ASTSMC) for permanent magnet synchronous motors (PMSM) suffers from prolonged convergence and insufficient disturbance rejection under complex operating conditions. This paper proposes an improved ASTSMC incorporating a fixed-time disturbance observer (FTDO). A power term is introduced into the adaptive super-twisting controller to accelerate convergence far from the origin, while the discontinuous sign function is replaced by a continuous h(s) function to mitigate chattering. The FTDO ensures disturbance estimation converges within a fixed time independent of initial states, overcoming the limitations of traditional and finite-time observers. The estimated disturbance is fed forward to the sliding mode controller for compensation, enhancing robustness. The FTDO design is based on an auxiliary state variable z and its derivative, with error dynamics analyzed via Lyapunov stability. Comparative simulations against conventional disturbance observers and sliding mode controllers validate the proposed strategy. The results demonstrate shorter convergence time, improved dynamic performance, and superior disturbance rejection, making the approach suitable for high-performance servo systems. The method addresses the critical need for robust, fast-response control in electric vehicles, rail transit, and aerospace applications where PMSM drives face significant uncertainties and external disturbances.

Examine Full Data & PDF