Chapter Four · failure evidence

What Time Series Forecasting & State-Space got wrong, from 93 dissertations

The records evaluate time series forecasting and state-space estimation techniques across domains such as energy, macroeconomics, finance, and transportation. Across these studies, performance frequently suffered when complex machine learning models failed against simple statistical baselines, data histories were too limited, or models neglected non-stationarity and spatial dependencies. These records come from PhD theses at 32 institutions, 2021 to 2026. Each links to its thesis. They were extracted by language models reading the full text, so treat each as a lead to read, not a verdict.

Complex machine learning and neural architectures underperform simpler statistical and naive baselines

16 theses · 13 institutions

Deep neural networks, transformers, and complex parameterizations frequently achieve worse predictive accuracy than basic autoregressive, moving average, or naive persistence baselines. Researchers often rejected or replaced these heavy models because added architectural complexity failed to justify higher error rates and increased computational overhead.

Tried and failed

time series transformer applied to asset price forecasting. Outcome: worse than baseline. Reason: underperformed simpler autoregressive baseline on predictive accuracy metrics

An Economic Evaluation of NFTs: A Systematic Framework for Prediction and Analysis with Other Comparable Assets · Texas Tech

Tried and failed

ARIMA and recurrent neural networks applied to transient emission time series prediction. Outcome: worse than baseline. Reason: severe non-linearity and limited training sample size caused overfitting and high variance compared to gradient boosting

A deterministic model for wear of piston ring and liner and a machine learning-based model for engine oil emissions · MIT

Tried and failed

hybrid statistical linear model and neural network applied to short-term wind speed forecasting. Outcome: worse than baseline. Reason: non-linear neural layer added complexity without exceeding linear autocorrelation baseline accuracy

Integrating Machine Learning and Weather Analytics for Sizing Variable Generation with Utility-Scale Energy Storage · TXST Digital Repository

Lost to a baseline

ARIMA achieved 1.0 standardized RMSE on Solar time series where Bayesian Synthesis achieved 1.47

Scalable Structure Learning, Inference, and Analysis with Probabilistic Programs · MIT

Lost to a baseline

Naive forecasting method outperformed ARIMA(1,1,0) at k=1 step-ahead forecasting in Lab Study 1 (median MASE = 1.10).

A Graph-Time Signal Processing Approach for Modeling and Monitoring with Applications in Healthcare and Occupational Safety · DSpace at SUNY Buffalo

Lost to a baseline

Transfer function Model VIII (RMSFE 1872.83) was beaten by univariate ARIMA Model II (RMSFE 1259.13) in 10-week out-of-sample forecasting for late 1998.

Forecasting of grain railcar shortage · Iowa State

Lost to a baseline

LSTM model underperformed basic statistical and naive models in long-term annual milk supply forecasting (MASE 4.96 and MAE 4.8M vs MASE 1.00 and MAE 0.97M for Seasonal Naive).

Post-Hoc, Contrastive, Explainable Artificial Intelligence for Time Series and Image Data · Research Repository UCD

Lost to a baseline

Nelson-Siegel-Svensson model achieved higher RMSE (worse forecasting performance) than the simpler Nelson-Siegel (Diebold & Li) model in rolling window AR(1) forecasts.

Estimating and forecasting the yield curve : Sri Lankan government securities market · Institutional Repository University of Moratuwa

Lost to a baseline

For long-term 1-day forecasting, standard SVR (MAE 0.97, MAPE 9.15%) outperformed univariate LSTM (MAE 4.16, MAPE 51.48%) and WEBSEC (MAE 3.14, MAPE 67.99%) in South Plains TX.

Wind Speed Prediction using Classical Time Series and Machine Learning Models: A Comparative Analysis · Texas Tech

Considered and rejected

Considered and rejected: Rejected sequence models (Gated Recurrent Units / GRU) and SDEs for temperature, HVAC consumption, and electricity spot price processes in favor of traditional time series models (SARIMA, Holt-Winters, ARX-GJR-GARCH) for superior interpretability and efficiency.

Three Essays on Energy Markets · JScholarship

Considered and rejected

Considered and rejected: Rejected artificial neural networks (ANNs) and support vector machines (SVMs) for econometric forecasting because literature concluded they either underperformed ARIMA or did not justify added complexity.

Mine to Table: Technology and Policy Strategies for Sustainable Mineral Supply Chains in the Low-Carbon Energy Transition · MIT

Considered and rejected

Considered and rejected: Machine learning time series forecasters (ANNs, SVMs), rejected because statistical models matched or exceeded accuracy with lower computational overhead and less overfitting risk.

Entraining a Robot to its Environment with an Artificial Circadian System · Georgia Tech

Lost to a baseline

Seasonal Naive forecasting beat ProphetNN in mean absolute error at the start and end of the year during low-supply periods (MAE 0.37M vs higher for Prophet models).

Post-Hoc, Contrastive, Explainable Artificial Intelligence for Time Series and Image Data · Research Repository UCD

Lost to a baseline

In small-scale 1-quarter point forecasting for GDP deflator, time-invariant VAR-MN achieved a lower RMSFE (0.9443) than the TV-SV-VAR-MN benchmark.

Bayesian Inference with Applications to Macroeconomics and Financial Market Price Discovery · ResearchWorks

Lost to a baseline

In small-scale 1-quarter point forecasting for Real GDP, standard flat VAR-OLS (0.9544) and VAR-MN (0.7442) outperformed the TV-SV-VAR-MN benchmark.

Bayesian Inference with Applications to Macroeconomics and Financial Market Price Discovery · ResearchWorks

Tried and failed

long short-term memory networks applied to long-horizon financial time series forecasting. Outcome: worse than baseline. Reason: obtained results that were below satisfactory performance levels

An AI Robot-Analyst System for Investment Analysis · Harvard

Lost to a baseline

In Dollar-Mark volatility forecasting for horizon T-out in [601,900], ARFIMA(0,d,0) achieved lower 10-step MSFE (7.216) than RLS (8.012).

Inference methods for locally ordered and common breaks in multiple regressions · OpenBU

Lost to a baseline

In Dollar-Yen volatility forecasting for horizon T-out in [601,900], ARFIMA(0,d,0) achieved lower 1-step (0.632), 5-step (3.868), and 10-step (9.584) MSFE than RLS (0.641, 4.008, 10.220).

Inference methods for locally ordered and common breaks in multiple regressions · OpenBU

Lost to a baseline

In Dollar-Mark volatility forecasting for horizon T-out in [601,900], ARFIMA(1,d,1) achieved lower 5-step MSFE (2.891) than RLS (3.066), and ARFIMA(0,d,0) achieved lower 1-step MSFE (0.632) than RLS (0.641).

Inference methods for locally ordered and common breaks in multiple regressions · OpenBU

Lost to a baseline

The Acemoglu et al. (2016) innovation network model was vastly outperformed in patent activity forecasting by a simple geometric random walk with drift.

Network-dependent dynamics of innovation and production · Oxford

Sequential models suffer from input window mis-specifications and multi-step trajectory divergence

18 theses · 12 institutions

Recurrent and sequential architectures often diverge or degrade when input lookback windows are too short to capture periodic cycles or too long and noisy. Truncating input histories, omitting physical states, and iterating multi-step forecasts over extended horizons lead to severe overestimation and abrupt prediction discontinuities.

Tried and failed

pure simulation-trained hybrid neural network forecasting applied to offshore wind speed prediction. Outcome: worse than baseline. Reason: lacking real-world observational ground-truth data fusion during training led to high prediction error against actual measurements

Data-driven and physics-informed wind prediction in offshore wind farms · University of Nottingham Repository

Tried and failed

temporal averaging of non-linear input variables applied to wind power generation forecasting. Outcome: worse than baseline. Reason: averaging distorts the skewed distribution shape and severely underestimates non-linear cubic power transformations

Multi-decadal wind power forecasting under climate change · Publikationssystem UB Tuebingen

Tried and failed

zero-shot LLM time series forecasting applied to retail demand forecasting. Outcome: did not generalise. Reason: models fail to associate moving holidays with irregular demand surges without explicit proximity prompts

Advancing Time Series Forecasting: Hierarchical Methods, Probabilistic Models, and Domain Knowledge Integration from Power Systems to Retail · Georgia Tech

Tried and failed

short lookback windows in recurrent neural networks applied to time series demand forecasting. Outcome: worse than baseline. Reason: insufficient historical context captured to accurately model longer-term periodic demand patterns

Electricity Load Forecasting in Texas using Neural Networks to Enhance the Power Grid Stability · Texas Tech

Tried and failed

mean demand imputation under unobserved censoring applied to competitive market demand forecasting. Outcome: worse than baseline. Reason: Assuming mean market arrivals when undercut caused systematic demand overestimation and severe performance drops

Market-Based and Policy-Based Conditional Demand Forecaster for Airline Revenue Management · MIT

Tried and failed

Regime-dependent state conditioning in predictive regression applied to text sentiment market forecasting. Outcome: worse than baseline. Reason: Conditioning on recession states statistically significantly increased forecast errors rather than improving predictability

Market-level and firm-level implications of financial distress · Imperial

Considered and rejected

Considered and rejected: Shallow ML models (KNN, Random Forest) were rejected for MPC system identification due to severe overfitting on sequential time-series data.

Multi-objective building system control optimization using machine-learning-based techniques · Harvard

Tried and failed

Gaussian process regression with truncated early-time inputs applied to time-series material property forecasting. Outcome: data insufficient. Reason: Excluding later-time observation points removed critical signal required to accurately predict long-term trajectory

Accelerated assessment of candidate supplementary cementitious materials using statistics and machine learning · Georgia Tech

Lost to a baseline

K-Nearest Neighbors underperformed simpler models (test RMSE of 1.72 °C for summer and 0.84 °C for winter) due to inability to handle time-series trends.

Supply Chain Modeling for Temperature-Sensitive Pharmaceutical Goods · MIT

Considered and rejected

Considered and rejected: Rejected linear regression and classical ML models (e.g., ElasticNet R²=0.086) for digital twin energy forecasting due to poor handling of temporal dependencies.

Towards Inclusivity In Building Performance Simulation A Multimodal Calibration Framework For Disadvantaged Communities · Georgia Tech

Tried and failed

shortest-path search with single-state input applied to temporal duration forecasting. Outcome: worse than baseline. Reason: lacked velocity cues necessary to predict duration effectively

Context-aware motion prediction of human-object interaction · Imperial

Considered and rejected

Considered and rejected: Rejected direct clustering (k-means) on raw latent time-series state vectors because latent Euclidean proximity does not preserve dispatch-relevant forecast trajectories.

Three Essays on Sustainable Operations: Renewable Energy Procurement, Battery Storage, and Equitable Work Scheduling · ResearchWorks

Tried and failed

black-box spatiotemporal forecasting without physical atmospheric state applied to photovoltaic power generation time series. Reason: models produced over-smoothed predictions that failed to capture abrupt, dynamic cloud-induced ramps without atmospheric covariates

Network time series forecasting in photovoltaics power production · EPFL

Tried and failed

increasing input sequence length in recurrent networks applied to solar irradiance forecasting. Outcome: overfit. Reason: historical lag features added noise and increased validation error compared to single current measurements

Exploiting known structure in data-driven models of dynamical systems · UT Austin

Tried and failed

sequence-to-sequence LSTM for trajectory forecasting applied to long-term battery capacity degradation. Outcome: unstable. Reason: exhibited sudden initial prediction discontinuities and severe overestimation of degradation rates

Data-driven early prediction and modeling of lithium-ion battery capacity degradation · Iowa State

Tried and failed

multi-step sequential forecasting over long horizons applied to behavior representation learning. Outcome: unstable. Reason: training failed beyond short horizons whereas aggregate distribution prediction remained stable

Building a foundation model for neuroscience · Georgia Tech

Tried and failed

model predictive control with lstm forecasting applied to solar-powered irrigation scheduling without battery buffer. Reason: solar forecasts had large positive-skewed errors, causing schedule misplacements and significant yield loss without energy storage

Product architectures for solar-powered drip irrigation (SPDI) systems in the Middle East and North Africa · MIT

Considered and rejected

Considered and rejected: Time-series recurrent neural network approaches for dynamic SOH estimation were rejected due to vulnerability to data point loss and the requirement to maintain continuous historical data.

Battery States Monitoring and its Application in Energy Optimization of Hybrid Electric Vehicles · IRIS - POLITO - prod

Data sparsity and short historical records cause severe overfitting and training instability

12 theses · 10 institutions

Time series models frequently fail when trained on small sample sizes, sparse event counts, or short longitudinal time horizons. These constrained data environments produce non-positive definite covariance matrices, prevent model convergence, and cause complex forecasters to lose out to simple heuristic thresholds.

Tried and failed

feature extraction and regression on short time-series applied to predicting skill demand trends. Outcome: worse than baseline. Reason: noise in short time-series made complex models underperform simple heuristic thresholds

Detecting Latent Training Needs Using Large Datasets · EPFL

Tried and failed

recurrent neural network time series forecasting applied to single-metric voter turnout estimation. Outcome: data insufficient. Reason: insufficient historical data points across election cycles to train time series models effectively

Exploring social and economic predictors for U.S. Government elections · Texas Tech

Tried and failed

clustering-based ensemble forecasting applied to intermittent retail demand forecasting. Outcome: did not generalise. Reason: sparse zero-sales data points and intermittent demand in low-sales regions degraded performance

Advanced Data-Driven Methodologies for Enhanced Demand Forecasting in Supply Chain Management · Georgia Tech

Tried and failed

regression and statistical performance trend modeling applied to time-series performance under new competition formats. Outcome: data insufficient. Reason: insufficient historical data points existed following a major regulatory and format change

Evaluating the Impact of Equipment Investments on Olympic Medal Probabilities for Australian Professional Cyclists · MIT

Considered and rejected

Considered and rejected: Rejected deep neural networks as primary ML models in sparse time-series experiments due to insufficient data volume for tuning complex architectures.

Beyond predictions: alignment between prior knowledge and machine learning for human-centric augmented intelligence · Leibniz Universität Hannover Repository

Considered and rejected

Considered and rejected: Rejected using sole quantitative/econometric regression models as the primary evaluation method because the sample size (n=6) and available time series (T=4-13) were insufficient due to data scarcity.

Evaluation of Global Business Services Centers impact on macroeconomic indicators in Central and Eastern Europe countries Globalių verslo paslaugų centrų poveikio Rytų ir Centrinės Europos šalių makroekonominiams rodikliams vertinimas · Mykolas Romeris University / Mykolo Romerio universitetas

Considered and rejected

Considered and rejected: Rejected monthly time scale for time series modeling because monthly data had insufficient SADEs per month and were non-stationary compared to quarterly data.

SERIOUS ADVERSE DRUG EVENTS IN U.S. GERIATRIC POPULATIONS: EMERGENCY DEPARTMENT VISIT TRENDS, CONTRIBUTING FACTORS, AND MONITORING METHODS · JScholarship

Tried and failed

multivariate time series models applied to macro forecasting with small sample size. Outcome: data insufficient. Reason: small sample size caused non-positive definite covariance errors and severe overfitting

An Assessment of EV Adoption and Potential Growth under Evolving Techno-policy Scenarios · Georgia Tech

Tried and failed

LSTM multi-step time series forecasting applied to multi-month ahead streamflow prediction. Outcome: data insufficient. Reason: insufficient monthly training samples led to poor long-horizon performance (NSE <= 0 beyond one month)

Development of generative adversarial networks for spatiotemporal fluid flow, atmospheric and flood predictions · Imperial

Tried and failed

dynamic time series regression on survey data applied to individual longitudinal psychotherapy process tracking. Outcome: data insufficient. Reason: Ceiling effects on post-session self-report rating scales eliminated variance required for time series regression modeling.

Using cognitive-behavioral case formulations to increase relevance and tailor treatment for African American and Hispanic adults experiencing comorbid depression and anxiety · Texas Tech

Tried and failed

ARDL and ECM time-series regression applied to institutional quality economic modeling. Reason: High multicollinearity between predictor variables forced the omission of a key regressor.

The dilemma of natural resource dependency in gulf countries. · Cranfield

Tried and failed

LSTM neural networks for tabular time-series forecasting applied to macroeconomic and election time-series data. Outcome: did not converge. Reason: training frequently failed to converge error terms and accuracy without intensive manual architecture tuning

Exploring social and economic predictors for U.S. Government elections · Texas Tech

Considered and rejected

Considered and rejected: Rejected yearly time scale because it provided too sparse data points for robust time series modeling.

SERIOUS ADVERSE DRUG EVENTS IN U.S. GERIATRIC POPULATIONS: EMERGENCY DEPARTMENT VISIT TRENDS, CONTRIBUTING FACTORS, AND MONITORING METHODS · JScholarship

Considered and rejected

Considered and rejected: Rejected VECM and GARCH models for stock forecasting because weekly portfolio series were stationary by construction and the 150-week in-sample series was too short for GARCH parameter stability.

Three essays on applied economics with high-frequency consumer data · University of Nottingham Repository

Non-stationarity and persistent temporal dynamics degrade long-term forecasting validity

14 theses · 11 institutions

Regressors exhibiting unit roots, structural breaks, and evolving seasonal patterns violate standard stationary time series assumptions. Direct forecasting on non-stationary features or failing to dynamically update models across regime shifts leads to severe out-of-sample error and spurious relationships.

Considered and rejected

Considered and rejected: Rejected forecasting future procedural trends from HES time series data because the time series lacked stationarity and sufficient longitudinal timepoints

Iliac vein stenting in the management of chronic venous outflow obstruction · Imperial

Considered and rejected

Considered and rejected: Rejected OLS, Fixed Effects, and Difference-GMM estimators due to severe weak instrument bias and failure to address dynamic panel endogeneity in persistent growth time series.

Foreign Investment and Institutional Effectiveness in Resource-Dependent Economies · DSpace at SUNY Buffalo

Tried and failed

time-series forecasting on correlation dynamics applied to predicting non-stationary feature relationships. Outcome: did not generalise. Reason: direct forecasting on multivariate correlation time series failed to capture non-stationary dynamics

Robust machine learning methods for high-dimensional datasets with applications in genomics and finance · Imperial

Tried and failed

Augmented Dickey-Fuller stationarity testing applied to macroeconomic time series indicators. Reason: Input variables exhibited unit roots and failed to demonstrate stationarity under unit root testing

An Assessment of EV Adoption and Potential Growth under Evolving Techno-policy Scenarios · Georgia Tech

Tried and failed

polynomial regression for time series forecasting applied to long-term cash flow forecasting. Outcome: did not generalise. Reason: models failed to maintain reasonable accuracy beyond a 12-month horizon

Sustainable Infrastructure Finance: Enhancing Transportation Construction Expenditure Management and Exploring Innovative Financing Mechanisms · Georgia Tech

Tried and failed

standard random forest without linear components applied to persistent macroeconomic time series forecasting. Outcome: did not generalise. Reason: cannot efficiently capture persistent autoregressive dynamics and smooth linear trends in finite samples

Machine Learning Econometrics · Penn

Tried and failed

omitting outlier periods and using standard pre-adjustments applied to time series macroeconomic forecasting. Reason: leads to statistical mis-specification, causing severe overprediction and underestimated uncertainty

Macroeconomic Forecasting: Statistically Adequate, Temporal Principal Components · Virginia Tech

Considered and rejected

Considered and rejected: Rejected applying time-series lag correction methods in Random Forest regression due to irregular lengths of merged township time-series.

Temporal and Spatial Impacts of Extreme Weather Events on Winter Wheat and Milling Oat Yields in Southern Ontario · Carleton University Institutional Repository

Considered and rejected

Considered and rejected: Original time series tracking score evolution over successive semesters was rejected because scores showed no regular longitudinal trends post-certification.

Evaluating Faculty Effectiveness: A Research Study on ACUE Training and Student Perceptions of Instruction · WTAMU Repository

Considered and rejected

Considered and rejected: Standard stationary time series assumptions rejected due to nonstationarity and long-range dependence in regressors

Essays on Time-Varying Models in Time Series and Machine Learning · Leibniz Universität Hannover Repository

Considered and rejected

Considered and rejected: Rejected unit fixed effects in pooled time-series cross-sectional models due to bias introduced when combining slow-moving covariates (GDP, regime type, fragility) with lagged dependent variables.

Offshoring Militarism: U.S. Military Aid and the Limits of American Foreign Policy · ResearchWorks

Considered and rejected

Considered and rejected: Rejected running time-series regressions on raw levels of log aggregate payrolls and manufacturing employment due to non-stationarity and spurious regression risk.

Altitude Sickness: Comparing the Political Economy of Two Periods of Extreme Dollar Strength · Harvard

Tried and failed

LSTM time-series forecasting without dynamic updates applied to photovoltaic power generation forecasting. Outcome: did not generalise. Reason: Inability to capture abrupt transient changes and long-term distributional shifts without continuous dynamic model updates

Applying Digital Twin in a Rotating Mechanical System and a Photovoltaic System · Texas Tech

Tried and failed

periodic time series forecasting for predictive control applied to robot environmental adaptation under phase shifts. Outcome: worse than baseline. Reason: phase-shifted disruptions caused out-of-sync forecasts, reducing performance to reactive baseline levels

Entraining a Robot to its Environment with an Artificial Circadian System · Georgia Tech

Classical linear models fail to capture intermittency, volatility, and non-linear dynamics

12 theses · 10 institutions

Linear autoregressive models like ARIMA degrade over extended forecasting horizons and large datasets because they assume linearity and static seasonality. These methods fail to track conditional heteroskedasticity, sharp spikes, and non-linear atmospheric transitions without dedicated variance or non-linear components.

Tried and failed

ARIMA time series forecasting applied to large-scale wind speed prediction. Outcome: did not generalise. Reason: linear autoregressive models degrade significantly over large datasets and long forecast horizons

Wind Speed Prediction using Classical Time Series and Machine Learning Models: A Comparative Analysis · Texas Tech

Tried and failed

low-order autoregressive and static parametric distribution modeling applied to wind speed time series modeling. Outcome: did not generalise. Reason: underestimated high-order autocorrelation, failing to capture complex temporal dynamics without higher-order ARMA modeling

RELIABILITY CONTRIBUTION OF WIND GENERATION IN ELECTRIC POWER SYSTEMS · HARVEST

Tried and failed

classical time series and statistical forecasting models applied to wind speed prediction. Outcome: did not generalise. Reason: unable to capture intermittent behavior and non-linearity over long-term forecasting horizons or large datasets

Wind Speed Prediction using Classical Time Series and Machine Learning Models: A Comparative Analysis · Texas Tech

Tried and failed

linear time series models without conditional heteroskedasticity applied to intermittent wind power generation forecasting. Outcome: did not generalise. Reason: unable to capture conditional volatility and intermittency without variance modeling components

Forecasting of wind power generation using wind speed and temperature for thambapawani wind farm in Sri Lanka · Institutional Repository University of Moratuwa

Tried and failed

single-model distribution forecasting applied to long-term wind speed forecasting. Outcome: worse than baseline. Reason: produced smaller standard deviation and poorer fit compared to Bayesian Model Averaging

Decommissioning strategy to reduce the cost and risk-driving factors in the offshore wind industry. · Cranfield

Tried and failed

ARIMA and ARIMAX statistical time-series forecasting applied to direct normal irradiance solar forecasting. Outcome: no signal. Reason: linear statistical models failed to capture variance during foggy, highly volatile atmospheric conditions

Machine learning applications for the optimization of renewable energy systems · Iowa State

Tried and failed

static seasonal ARIMA time series modeling applied to time series with evolving seasonality. Outcome: worse than baseline. Reason: fixed seasonal components failed to adapt to time-varying seasonality, resulting in wide prediction intervals

Impact of the 65 mph speed limit on Iowa's rural interstate highways: An integrated Bayesian forecasting and dynamic modeling approach · Iowa State

Tried and failed

seasonal autoregressive integrated moving average imputation applied to wind speed time series data. Outcome: worse than baseline. Reason: Produced significant distributional distortion and high KL divergence compared to non-seasonal ARIMA and LSTM

Data Driven Early Stage Design Support for Offshore Wind Farms · Research Repository UCD

Considered and rejected

Considered and rejected: Rejected SARIMA for long-cycle seasonal forecasting due to computational inefficiency over long periods, using Fourier terms with ARIMA/Ridge instead

A Novel Computational Ethology Framework for Studying Animal Behaviour Under Climate Change (CEFABC2): A Case Study of Little Penguins on Phillip Island · YorkSpace

Considered and rejected

Considered and rejected: Rejected classic linear time series forecasting models (like ARIMA) due to linearity assumptions and inability to handle non-linear multi-target attributes efficiently.

Individual trustworthiness analysis on social media · UT Austin

Considered and rejected

Considered and rejected: Rejected non-ML-based forecasting methods (ARIMA and exponential smoothing) as baselines because they struggle to effectively capture complex patterns and irregularities with non-linear trends

Machine learning-based performance analytics for high-performance computing systems · OpenBU

Lost to a baseline

Benchmark day-hour monthly factor models outperformed July-August day-hour ARIMA models on unstable time series patterns.

Data mining applications for updating missing values of traffic counts. · oURspace

Lost to a baseline

DO-WAVES (dynamic oracle) achieves up to 12% lower photonic power than PROWAVES because ARIMA time-series predictions and linear regression have residual errors.

Energy-efficient architectures for chip-scale networks and memory systems using silicon-photonics technology · OpenBU

State-space and filtering methods degrade due to noise mis-specification and over-smoothing

10 theses · 9 institutions

Linear Kalman filters and dynamic linear models often over-smooth dynamic peaks or amplify sensor noise when process noise parameters are fixed. Denoising procedures can inadvertently strip away target-correlated signals, causing state estimators to perform worse than unintegrated or phase-unaware baselines.

Tried and failed

Kalman filtering with signal quality indices applied to physiological time series denoising. Outcome: worse than baseline. Reason: Denoising removed motion and noise patterns that were informative and correlated with the target event

Self-Aware Machine Learning for Chronic Pathology Monitoring on Wearable Devices · EPFL

Tried and failed

Kalman filtering with constant process noise covariance applied to physiological sensor time series denoising. Reason: Low process noise oversmoothed dynamic peaks while high process noise failed to suppress high-frequency noise

Modeling glucose dynamics during physical activity using a linear model for individuals with Type 1 Diabetes · Harvard

Tried and failed

linear Kalman filter applied to continuous glucose time series tracking. Reason: linear state transitions over-smoothed rapid non-linear spikes and drops

Novel Machine Learning Approaches with Applications in Healthcare and Social Welfare · Georgia Tech

Tried and failed

threshold-based alerting on continuous time-series data applied to continuous physiological vital sign monitoring. Reason: Excessive false-positive alert rates due to high sensor noise and normal physiological fluctuations.

Characterising clinically integrated remote sensing and digital alerting tools for the assessment of health status in community and hospital care · Imperial

Lost to a baseline

Under random time series errors alone, standalone IRS (0.46995 NM/HR 95% error) outperformed the integrated Kalman filter (0.876 NM/HR 95% error)

Integration of global positioning and inertial reference system data inside a flight management computer · Cranfield

Lost to a baseline

Phase-aware forecasting failed to improve upon phase-unaware baselines on ferret and bodytrack due to high-frequency noise and classification misdetections.

Machine learning-based proactive runtime management · UT Austin

Lost to a baseline

Simple Dynamic Linear Model (SDLM) and Exponentially Weighted Moving Average (EWMA) were heavily outperformed on geopolitical time-series by Sample-And-Calibrate (SAC) due to severe under-confidence and lack of sharpness.

Partial Information Framework: Basic Theory and Applications · Penn

Lost to a baseline

Dynamic linear model (SpDynLM) produced higher hold-out prediction error (RMSPE = 0.868) than Graphical Matern (RMSPE = 0.800) in the AR(1) spatial time-series simulation Set 3A

Topics in Modeling of Multivariate Mixed Data Types and Highly Multivariate Spatial Data · JScholarship

Tried and failed

splitting states by task phases applied to macroscopic traffic flow forecasting. Outcome: worse than baseline. Reason: short-duration states were highly susceptible to measurement noise

Modeling and Optimization of Ridesplitting Operations · EPFL

Tried and failed

adding explanatory covariates to dynamic linear models applied to traffic fatality rate forecasting. Outcome: worse than baseline. Reason: explanatory speed variables did not improve predictive performance over the univariate baseline state-space model

Impact of the 65 mph speed limit on Iowa's rural interstate highways: An integrated Bayesian forecasting and dynamic modeling approach · Iowa State

Spatiotemporal and cross-series network dependencies are lost in standard time series formulations

9 theses · 8 institutions

Univariate time series and non-spatial regression models fail to account for dynamic network topologies, spatial non-stationarity, and physical transport delays. Omitting downstream sensor bottlenecks or decoupling temporal graphs from spatial constraints degrades prediction accuracy across transportation and physical grids.

Considered and rejected

Considered and rejected: Rejected non-graph classical time-series methods (ARIMA, Kalman filters, standard linear regression) because they fail to simultaneously capture dynamic spatial topological features and non-spatial edge throughput/bandwidth constraints.

Using machine learning for dynamic resource orchestration & task scheduling, in a radio access network based edge environment · Oxford

Considered and rejected

Considered and rejected: Rejected univariate time series models (e.g., ARIMA) for mobility network prediction due to their inability to model cross-series dependencies

Towards Equitable Access to Essential Services: A Data-Driven Human Mobility Modeling Framework · DSpace at SUNY Buffalo

Tried and failed

traditional machine learning on multispectral time-series applied to crop yield spatial prediction. Outcome: worse than baseline. Reason: shallow models struggled to capture complex spatiotemporal dependencies compared to deep architectures

Application of Precision Agriculture Technologies for Sustainable Crop Management in the Southern High Plains · Texas Tech

Tried and failed

binary classification of precursor states from sensor time-series applied to freeway traffic congestion prediction. Outcome: did not generalise. Reason: sensor coverage omitted the downstream bottleneck location, conflating bottleneck activation with backward shockwave propagation

Identification of Traffic Congestion Precursors using Machine Learning Approaches · Georgia Tech

Tried and failed

non-spatial regression and neural networks applied to spatial production forecasting. Outcome: did not generalise. Reason: Failed to capture spatial non-stationarity in the data.

Oil and gas data analytics: Application of spatial, spatio-temporal, and other machine learning models for a multi-basin unconventional production evaluation and economic analysis · Texas Tech

Tried and failed

time gating in recurrent graph neural networks applied to spatiotemporal traffic forecasting. Outcome: worse than baseline. Reason: time gating fails to capture large spatial correlations across nodes

Machine Learning On Large-Scale Graphs · Penn

Tried and failed

temporal convolution coupled with causal graph learning applied to long-term spatial-temporal traffic forecasting. Outcome: worse than baseline. Reason: None

Bilevel Optimization in the Deep Learning Era: Methods and Applications · Virginia Tech

Considered and rejected

Considered and rejected: Rejected standard continuous-coordinate time-series forecasting (vector regression) because it fails to enforce spatial network continuity and is susceptible to GPS noise.

Deep Generative Models for Trajectory Prediction and Mobility Network Forecasting · YorkSpace

Considered and rejected

Considered and rejected: Rejected standard matrix/tensor decomposition alone for mobility forecasting due to its inability to capture dynamic temporal evolution without temporal regularization/prediction modules

Towards Equitable Access to Essential Services: A Data-Driven Human Mobility Modeling Framework · DSpace at SUNY Buffalo

Considered and rejected

Considered and rejected: Rejected simulating range dependence using discrete mooring time series alone, which misses wave-induced horizontal spatial variability along the propagation path.

CHARACTERIZATION OF INTERNAL WAVES AND TIDES OVER A CALDERA IN THE NEW ENGLAND SEAMOUNT CHAIN AND THEIR IMPACT ON SHORT-TERM ACOUSTIC PATH AVAILABILITY · Calhoun

Left open by the authors

Problems the authors named and did not get to.

Left open

Evaluate wind time-series imputation models (e.g., LSTM, ARIMA) on data gaps exceeding 10% and during prolonged outages. Blocker: None

Data Driven Early Stage Design Support for Offshore Wind Farms · Research Repository UCD

Left open

Develop hybrid or ensemble forecasting models combining multiple RNN architectures with statistical models for multi-step wind power forecasting. Blocker: None

Deep Learning-Based Medium to Long-Term Multi-Step Ahead Wind Power Generation Forecasting · TXST Digital Repository

Left open

Develop and benchmark wind power forecasting models using the compiled open-source wind power dataset collection. Blocker: None

Multi-decadal wind power forecasting under climate change · Publikationssystem UB Tuebingen

Left open

Implement transformer architectures and model distillation techniques for multi-step ahead wind power generation forecasting. Blocker: None

Deep Learning-Based Medium to Long-Term Multi-Step Ahead Wind Power Generation Forecasting · TXST Digital Repository

Left open

Extend the energy management system framework to support wind power forecasting and integration using meteorological data fusion and regression models. Blocker: None

An improved energy management system framework for solar energy integration. · Cranfield

Left open

Model long-range dependency in wind power time series residuals using fractional Lévy stable motion (fLsm) with parameters alpha and H. Blocker: None

Graphon Mean Field Games with Finite States and Forecasting Models for the Energy Market · IRIS - UNITN - prod

Left open

Develop a Bayesian deep learning model for wind power forecasting to capture uncertainty. Blocker: None

Towards intelligent operation of future power system: bayesian deep learning based uncertainty modelling technique · Imperial

Left open

Extend the univariate wind power forecasting models to multivariate architectures by incorporating meteorological and environmental exogenous variables. Blocker: None

Deep Learning-Based Medium to Long-Term Multi-Step Ahead Wind Power Generation Forecasting · TXST Digital Repository

Left open

Benchmark the adaptive linear regression drift-handling strategy against online ARIMA, SVM, Kalman filters, and RNNs on streaming time-series data. Blocker: None

Modelado Predictivo en Flujo de Datos de Procesos con Deriva de Concepto y su Aplicación al Turismo en Canarias · accedaCRIS

Left open

Compare Generalized Network Autoregressive (GNAR) forecasting performance against Bayesian Vector Autoregression (BVAR) models on multivariate time series benchmarks. Blocker: None

New statistical and machine learning methods for time series forecasting · Imperial

Checking a claim in this area?

We can run the same search on any method or claim. If nothing turns up, we will say so, and that proves nothing on its own.