Key Takeaways & Executive Findings
- •• • The TD3-R_LADRC strategy reduces DC bus voltage fluctuations by optimizing both observer and controller bandwidths via TD3, achieving faster convergence than fixed-parameter LADRC; this directly mitigates the risk of converter shutdown or reduced battery lifetime in microgrids with high renewable penetration. • • The improved LESO estimates the disturbance derivative and applies order reduction to known states, enhancing observation accuracy without increasing system order; this lowers computational burden and parameter tuning complexity, enabling practical deployment on embedded controllers with limited resources. • • Comparative experiments under renewable intermittency and load steps show TD3-R_LADRC outperforms dual-loop PI and conventional LADRC in disturbance rejection and robustness, with reduced settling time and overshoot; this translates to higher power quality and fewer protection trips in commercial energy storage systems. • • The TD3 algorithm addresses Q-value overestimation and local optima issues inherent in DDPG, as evidenced by stable training convergence and improved control performance; this provides a reliable reinforcement learning framework for online parameter adaptation in safety-critical power electronics.
China Clean Energy & Battery Radar
Get verified English translations, SEM micrographs & open-access PDF alerts from China's leading state key laboratories delivered to your inbox every Monday at 08:00 EST.
Abstract
Output voltage fluctuations in DC microgrids arise from renewable generation intermittency, spatiotemporal load variations, and external disturbances. This study proposes a reconstructed linear active disturbance rejection control strategy (TD3-R_LADRC) that integrates a twin delayed deep deterministic policy gradient (TD3) algorithm to enhance the DC bus voltage stabilization capability of battery energy storage interface converters. The improved linear extended state observer (LESO) estimates the derivative of the total disturbance and applies order reduction to known state variables, achieving faster and more accurate tracking and compensation without increasing system order. Frequency-domain performance and stability analyses are conducted for the proposed strategy. The TD3 reinforcement learning algorithm then optimizes the observer bandwidth and controller bandwidth of the improved LADRC, enabling precise observation and rapid convergence. Digital simulations and low-power experiments compare the proposed TD3-R_LADRC against conventional LADRC and dual-loop PI control under various operating conditions. Results demonstrate that TD3-R_LADRC exhibits superior disturbance rejection, stability, and robustness against renewable output uncertainty, load fluctuations, and external disturbances, effectively improving frequency stability control and offering theoretical and engineering value for energy storage converter applications.
1. Introduction
Commercial DC microgrids integrating high shares of solar photovoltaic generation face persistent bus voltage instability due to the stochastic and intermittent nature of renewable resources, compounded by unpredictable load profiles. Conventional dual-loop PI controllers struggle to balance response speed and stability, often exhibiting delayed transient recovery and limited disturbance rejection. While linear active disturbance rejection control (LADRC) offers a model-independent alternative by actively estimating and compensating total disturbances, its performance is constrained by fixed observer and controller bandwidths that cannot adapt to time-varying operating conditions. Prior attempts to enhance LADRC through nonlinear functions or higher-order observers have increased parameter tuning complexity and computational overhead, hindering practical implementation.
This study addresses the parameter rigidity bottleneck by introducing a twin delayed deep deterministic policy gradient (TD3) algorithm to dynamically optimize the bandwidths of a reconstructed LADRC. The proposed TD3-R_LADRC first modifies the linear extended state observer (LESO) to estimate the disturbance derivative and reduce the order of known state variables, improving observation speed and accuracy without elevating system order. The TD3 agent then learns optimal bandwidth parameters through continuous interaction with the converter environment, overcoming the Q-value overestimation and local optima issues of DDPG. Frequency-domain analysis and experimental validation confirm that TD3-R_LADRC achieves superior disturbance rejection, stability, and robustness compared to conventional LADRC and dual-loop PI, offering a practical pathway for enhancing energy storage converter performance in uncertain microgrid conditions.
Loading authentic research manuscript (Pages 1–5)...
MA Youjie, YAN Fengxiang, ZHOU Xuesong, TAO Long, WANG Xinyue, CHEN Yunfei (2026). Improved Linear Active Disturbance Rejection Control of Energy Storage Converters Based on the TD3 Algorithm. Acta Energiae Solaris Sinica. https://doi.org/10.19912/j.0254-0096.tynxb.202608_9662
Research & Educational Purpose Only: The translations, structured abstracts, analytical annotations, and data reports provided by SinoGreenTechare intended exclusively for academic research, internal corporate R&D, and educational benchmarking. They do not constitute formal engineering, chemical safety, legal, or professional advice.
Copyright & Intellectual Property Notice: Original copyright of the underlying source articles and experimental data remains with the respective authors, institutions, and original publishing journals. SinoGreenTech claims intellectual property only over its proprietary translations, analytical syntheses, and AEO structured enhancements in accordance with international fair use and academic citation principles.
Frequently Asked Questions
What specific failure mechanisms in conventional LADRC does the TD3-R_LADRC strategy mitigate under renewable intermittency?
Conventional LADRC suffers from fixed observer and controller bandwidths that cannot adapt to rapid disturbances, leading to sluggish disturbance estimation and residual voltage oscillations. The TD3-R_LADRC dynamically tunes these bandwidths via reinforcement learning, reducing settling time and overshoot. Experimental comparisons show that under step load changes, TD3-R_LADRC achieves faster convergence and lower steady-state error than LADRC, effectively preventing bus voltage excursions that could trigger protective shutdowns.
How does the improved LESO achieve order reduction without increasing system complexity, and what is the impact on computational load?
The improved LESO estimates the derivative of the total disturbance and applies order reduction to known state variables, eliminating the need for higher-order observer structures. This maintains the original system order while enhancing observation accuracy. Consequently, the computational burden remains comparable to standard LESO, making it suitable for real-time implementation on digital signal processors used in commercial converters.
What evidence demonstrates that TD3 outperforms DDPG in optimizing LADRC parameters for this application?
TD3 addresses DDPG's Q-value overestimation by employing twin critics and delayed policy updates, which prevents local optima and stabilizes training. In this study, TD3-R_LADRC exhibited consistent convergence and superior control performance across varying operating conditions, whereas DDPG-based LADRC showed parameter oscillations and suboptimal disturbance rejection. The experimental results confirm that TD3 yields more reliable bandwidth optimization, leading to improved voltage regulation.
What are the scalability bottlenecks for deploying TD3-R_LADRC in multi-converter microgrids?
Scalability is constrained by the need for centralized training of the TD3 agent, which requires communication and coordination among converters. However, the improved LESO's reduced computational load and the TD3's sample efficiency enable distributed implementation with local measurements. Future work should address multi-agent coordination and communication delays to extend the strategy to large-scale microgrids.
How does the proposed control strategy impact the lifetime and cost of battery energy storage systems?
By reducing bus voltage fluctuations and settling time, TD3-R_LADRC minimizes stress on battery cells caused by frequent charge-discharge cycles and overvoltage events. This can extend battery lifetime and reduce replacement costs. Although the TD3 algorithm adds computational overhead, the use of low-power embedded processors keeps implementation costs marginal compared to the benefits of improved reliability and reduced maintenance.
Related Chinese Research & Cross-Citations
Wind Turbine Gearbox Fault Diagnosis Method Based on Improved CNN-XGBoost Fusion Model Under Gramian Angular Difference Field
Gearbox failures account for 20–30% of total wind turbine faults and incur maintenance costs equivalent to 10–15% of overall turbine value. Conventional vibration diagnostic pipelines—complementary ensemble empirical mode decomposition with singular value energy spectrum, time-varying filtering empirical mode decomposition, and Teager energy spectrum analysis—remain bounded below 90% accuracy and depend on expert-driven feature engineering that is sensitive to non-stationary operating conditions and noise. This study proposes an intelligent diagnostic architecture that converts one-dimensional gearbox vibration signals into two-dimensional images via Gramian angular difference field (GADF) transformation, preserving intrinsic temporal correlation and time-frequency structure while exploiting matrix sparsity to suppress interference. An improved convolutional neural network (CNN) extracts multi-dimensional features: a convolutional block attention module (CBAM) is embedded in the convolutional layers to weight critical channels and focus on fault-sensitive spatial regions, and a modified βc-ACONC activation function replaces ReLU to mitigate neuron necrosis and enable selective activation. The extracted composite features are then fed into an XGBoost network whose hyperparameters are optimized by an improved sparrow search algorithm (ISSA). Validation on a laboratory wind turbine gearbox dataset yields diagnostic accuracy exceeding 99%, demonstrating robust fault identification capability under complex operating conditions.
Joint Forecasting of Wind and Photovoltaic Power Considering Complementarity
The inherent spatiotemporal complementarity between wind and solar resources offers a theoretical basis for improving renewable power forecasting accuracy. This study proposes a joint wind-photovoltaic (PV) power forecasting strategy that explicitly exploits this complementarity. A bidirectional long short-term memory (BiLSTM) neural network serves as the baseline forecasting model, and a novel sorting and comparative optimization (SCO) algorithm is developed to optimize the model's hyperparameters. The SCO algorithm ranks individuals in ascending order and compares adjacent fitness values to escape local optima, a known deficiency in conventional metaheuristics such as genetic algorithms and particle swarm optimization. For wind farms and PV plants exhibiting significant complementarity, the joint forecasting strategy first aggregates their power outputs, normalizes the combined signal, and then feeds it into the optimized BiLSTM model. Experimental results demonstrate that the proposed SCO-BiLSTM model reduces the eMAPE by 10.313% compared with PSO-BiLSTM. Furthermore, joint forecasting under SCO-BiLSTM lowers the eRMSE by 27.443% relative to standalone PV power forecasting. The study also establishes that forecasting accuracy improves with stronger wind-solar complementarity but degrades as the forecasting horizon extends. These findings confirm that exploiting complementarity in joint forecasting substantially enhances predictive performance for renewable energy integration.
Ultra-Short-Term Wind Power Forecasting Based on Fluctuation Continuation Scenario Identification
Existing ultra-short-term wind power forecasting methods exhibit limited performance due to insufficient extraction of fluctuation information and inadequate analysis of evolution patterns. This paper proposes an ultra-short-term wind power forecasting method based on fluctuation continuation scenario identification. First, the coupling mechanism of wind power fluctuations under multiple turbulence processes is investigated, and historical wind power dynamics are decoupled into a combination of nonlinear and linear fluctuation components. A fluctuation continuation concept is introduced, and the future continuation scale of wind power fluctuations is derived from nonlinear and linear decoupling parameters, thereby classifying fluctuation continuation scenarios. A sparse neural network (SNN) oriented to high-dimensional sparse features is constructed to identify historical fluctuation continuation scenarios, and ultra-short-term power forecasting is conducted separately for each scenario. Validation using measured wind speed and power data from three wind farms shows that, compared with baseline models, the proposed model improves root mean square error (RMSE), mean absolute error (MAE), and mean absolute percentage error (MAPE) by at least 1.46%, 2.44%, and 14.67%, respectively, demonstrating superior accuracy and stability. The method addresses the limitations of signal decomposition techniques that lack physical interpretability and are sensitive to hyperparameters, and overcomes the high-dimensional sparsity challenges faced by traditional scenario identification models.
Multi-Classifier Open Adversarial Network for Rolling Bearing Fault Diagnosis in Wind Turbine Generator Systems
Rolling bearings in wind turbine generator systems operate under variable speed and load conditions that induce significant data distribution shifts between training and field data, while unknown fault modes absent from the source domain are frequently misclassified as known classes. This study proposes a multi-classifier open adversarial network (MCOAN) for open-set fault diagnosis. Within an adversarial domain adaptation framework, a K-way classifier and an additional K+1-way one-vs-all classifier independently estimate target-sample similarity to the source domain. These similarity scores drive a dynamic weighting mechanism that adaptively reweights target samples during open-set adversarial training and supplies per-sample dynamic thresholds for known/unknown discrimination, thereby promoting cross-domain alignment of shared-class features while suppressing negative transfer from unknown samples. A non-adversarial domain classifier is introduced to stabilize dynamic weight estimation. Validation on two datasets demonstrates high-precision shared-class distribution alignment and unknown-class recognition with favorable robustness. The method removes reliance on empirically preset thresholds that plague conventional open-set back-propagation approaches, where a fixed threshold of 0.5 in the binary cross-entropy adversarial loss provides no per-sample adaptivity. By coupling K-way and K+1-way similarity estimates, MCOAN achieves simultaneous known-class alignment and unknown-class separation without prior knowledge of the unknown-class cardinality, addressing a persistent bottleneck in wind turbine drivetrain condition monitoring where unanticipated bearing failure modes emerge under field conditions.
Unsupervised Automated Identification Method for Abnormal States of Wind Turbine Gearboxes
Addressing the scarcity of labeled data for training classification models in wind turbine planetary gearbox anomaly identification, this study proposes an unsupervised automated detection method. Log Mel-band energy features are extracted from raw vibration signals and fed into an unsupervised anomaly recognition model centered on a U-net autoencoder. A health-state threshold is established based on reconstruction error between model input and output, enabling anomaly identification. The method is validated using factory gearbox test data and operational data from a wind farm in Yangtouya, Shanxi. For factory gearboxes, dual validation is performed using a spectrum amplitude modulation-based signal processing method. Results demonstrate that the proposed method achieves 93.34% recognition accuracy on both factory and wind farm test sets, confirming its capability to automatically and correctly separate abnormal wind turbine gearboxes. The approach eliminates reliance on labeled fault data, offering a scalable solution for full-lifecycle health monitoring, from factory acceptance testing to in-service early anomaly detection, adaptable across different operating conditions and turbine models.
Improved Adaptive Super-Twisting Sliding Mode Control for Permanent Magnet Synchronous Motors
Adaptive super-twisting sliding mode control (ASTSMC) for permanent magnet synchronous motors (PMSM) suffers from prolonged convergence and insufficient disturbance rejection under complex operating conditions. This paper proposes an improved ASTSMC incorporating a fixed-time disturbance observer (FTDO). A power term is introduced into the adaptive super-twisting controller to accelerate convergence far from the origin, while the discontinuous sign function is replaced by a continuous h(s) function to mitigate chattering. The FTDO ensures disturbance estimation converges within a fixed time independent of initial states, overcoming the limitations of traditional and finite-time observers. The estimated disturbance is fed forward to the sliding mode controller for compensation, enhancing robustness. The FTDO design is based on an auxiliary state variable z and its derivative, with error dynamics analyzed via Lyapunov stability. Comparative simulations against conventional disturbance observers and sliding mode controllers validate the proposed strategy. The results demonstrate shorter convergence time, improved dynamic performance, and superior disturbance rejection, making the approach suitable for high-performance servo systems. The method addresses the critical need for robust, fast-response control in electric vehicles, rail transit, and aerospace applications where PMSM drives face significant uncertainties and external disturbances.