Data preprocessing techniques for statistical photovoltaic forecasting models
Date published
Free to read from
Supervisor/s
Industry supervisor/s
Journal Title
Journal ISSN
Volume Title
Publisher
Department
Course name
Type
ISSN
Format
Citation
Abstract
Accurate photovoltaic power forecasting depends on high-quality data, yet raw meteorological and photovoltaic datasets often contain noise, missing values, and redundant features. Data preprocessing is therefore critical, but existing studies typically apply isolated techniques without a structured framework or systematic evaluation of combined strategies. This paper addresses these gaps by introducing a functional classification of data preprocessing methods and empirically testing sixteen (16) widely used techniques—individually and in combination—on raw photovoltaic and meteorological data. Two scenarios are considered: (i) comparing optimal data preprocessing combinations to a base-case approach using minimal cleaning, and (ii) benchmarking against an established three-step photovoltaic forecasting model. Forecasting performance is assessed using Root Mean Square Error across more than twenty (20) regression algorithms in MATLAB’s Regression Learner App. The experimental findings are further validated through cross–testing with an independent dataset, and an iterative ranking procedure is introduced to examine preprocessing order and cumulative impact. Results demonstrate that integrated data preprocessing strategies significantly outperform both baseline and conventional preprocessing pipelines, highlighting the importance of method synergy. This study provides a reproducible framework for data preprocessing classification and evaluation, offering practical guidance for researchers seeking to optimize photovoltaic forecasting models.
