• • Two-stage VF2-based graph clustering reduces effective execution sequence variants from 20 to 1, cutting DRAM-related access cycles from 18,595 to 1,502 and raising L2 cache hit rate from 16.09% to 84.54%. This directly mitigates thread warp divergence in SIMT architectures, enabling predictable memory access patterns essential for scaling to thousands of heterogeneous power electronic devices.
• • Static specialization via automatic code generation eliminates dynamic graph parsing, reducing instruction count per control step by 91.8%. This translates to a 2.7x kernel-level speedup, demonstrating that removing runtime interpretation overhead is critical for converting GPU raw compute throughput into actual EMT simulation acceleration.
• • The integrated method achieves a nearly 3x end-to-end speedup over conventional CPU-based simulation for a grid with thousands of renewable energy devices. This performance level is necessary for practical wide-frequency oscillation analysis and stability studies where traditional CPU serial or multi-core parallelism fails to meet efficiency requirements.
• • The approach exhibits excellent scalability, as validated on large-scale scenarios. By discarding idealized homogeneity assumptions and automatically aggregating structurally isomorphic control topologies, it addresses load imbalance that plagues simple grouping methods, making it viable for real-world grids with diverse wind, PV, and storage systems from multiple manufacturers.
Download Full PDF: GPU-Based Parallel Acceleration Method for Electromagnetic Transient Simulation of Large-Scale Renewable Energy Power Grids | SinoTechIntel | SinoGreenTech