Water Quality Prediction and Data Quality Enhancement of the Lhasa River Using Machine Learning
This study develops a multivariate time-series forecasting model for water quality in the Lhasa River, focusing on four key indicators: water temperature, pH, dissolved oxygen, and turbidity. Data preprocessing integrated multiple missing-value imputation strategies and interquartile range (IQR) outlier removal. Boxplots and relative standard deviation (RSD) assessed data distribution and dispersion, while autocorrelation and Pearson correlation analyses revealed periodic patterns and inter-variable relationships. Four representative algorithms—Support Vector Regression (SVR), Extreme Gradient Boosting (XGBoost), CNN-BiLSTM-Attention, and TCN-Transformer—were optimized via Bayesian hyperparameter tuning. Model performance was evaluated using MAE, MSE, RMSE, and R². The study systematically compared the effects of different missing-value handling methods, both independently and combined with IQR outlier removal. Results indicate that CNN-BiLSTM-Attention excels in water temperature prediction, suitable for relatively stable and simple patterns. In contrast, TCN-Transformer demonstrates superior performance for pH, dissolved oxygen, and turbidity, which exhibit strong nonlinearity and long-term dependencies, effectively capturing temporal dependencies and coupling relationships. The findings provide a viable technical route and theoretical reference for river water quality monitoring and intelligent early-warning systems.