APPLICATION OF ARTIFICIAL INTELLIGENCE AND MACHINE LEARNING IN THE PREDICTION OF WATER QUALITY PARAMETERS

Authors

  • Ochiagha Kate Ekwutosi, Aghanwa Charles Ifeanyi, Obiefuna Joy Ngozika, Onyeije Ugomma Chibuzo, Oli Christain Chukwuemeka, Ezeokoye Genevieve Nwanneka Author

Keywords:

Artificial intelligence; machine learning; water quality prediction; water quality index; dissolved oxygen; XGBoost; LSTM; explainable AI

Abstract

Reliable water-quality prediction is increasingly important because laboratory monitoring alone cannot provide the temporal coverage needed for rapid pollution detection, process control, and water-resource planning. This article critically examines how artificial intelligence (AI) and machine learning (ML) are being applied to predict physicochemical and ecological water-quality variables, with attention to dissolved oxygen (DO), pH, electrical conductivity (EC), total dissolved solids (TDS), nutrients, chlorophyll-A, and composite water quality indices (WQIs). We conducted a structured integrative evidence synthesis using 28 DOI-verified peer-reviewed studies published from 2020 to 2026; 25 were empirical modelling studies, and three were recent reviews. The evidence was coded by target variable, data setting, model family, feature strategy, validation approach, and reported performance. The study shows a clear movement from standalone artificial neural networks toward tree-based ensembles, optimized hybrids, explainable AI, and attention/Transformer architectures. Across the empirical set, tree/boosting ensembles and ANN-based models were the most frequently used families, while long short-term memory networks and Transformers were concentrated in time-series and long-horizon forecasting. Results also show that model accuracy depends less on algorithm novelty alone than on data completeness, temporal alignment, feature selection, and validation design. Random Forest and XGBoost perform strongly for WQI classification and reduced-input prediction, while optimized ANN models remain competitive for DO forecasting; self-attentive LSTM and transfer-learning Transformers are particularly useful when temporal dependence or sparse observations dominate. The article proposes an integrated prediction framework combining data-quality control, feature selection, model benchmarking, uncertainty-aware validation, and explainable outputs. The study also emphasizes that benchmark comparisons and data-leakage controls are essential before interpreting high published accuracy as evidence of deployment readiness. It concludes that AI/ML can substantially reduce monitoring cost and improve early warning, but operational deployment requires cross-site validation, transparent uncertainty reporting, and safeguards against data leakage and black-box decision-making.

Downloads

Download data is not yet available.

Downloads

Published

2026-10-05

Issue

Section

Articles