Chapter Four · failure evidence

What Ensemble Learning got wrong, from 47 dissertations

The records detail empirical challenges encountered when designing and applying ensemble learning methods across various domains. Many attempts failed because complex ensembling strategies fell short of simpler base models, suffered from overfitting, or introduced computational bottlenecks. These records come from PhD theses at 25 institutions, 2021 to 2026. Each links to its thesis. They were extracted by language models reading the full text, so treat each as a lead to read, not a verdict.

Complex ensembling and stacking underperform simpler baselines or individual base models

15 theses · 13 institutions

Stacked generalization and meta-learning models frequently lost to simpler linear baselines, standalone base models, or basic voting schemes. In multiple cases, aggregating model predictions produced higher prediction error and variance than the constituent learners on their own.

Tried and failed

random forest regression applied to predicting physical parameters from relaxometry data. Outcome: worse than baseline. Reason: non-linear ensemble failed to outperform a simpler quadratic general linear model baseline

Quantitative Microstructural Imaging for Clinical Use · EPFL

Tried and failed

test-time representation optimization with MLP feature reconstruction applied to ensemble feature aggregation for downstream regression. Outcome: worse than baseline. Reason: regularization and architectural tuning failed to outperform simple ensemble averaging on downstream linear regression

BEYOND VANILLA FINETUNING: APPROACHES TO MAXIMIZE THE BENEFITS OF PRETRAINING IN COMPUTER VISION · Cornell

Lost to a baseline

Ensemble methods (DRF, Stacked Ensemble, GBM) were beaten on out-of-sample evaluation data by the simpler Mallows's Cp linear regression model due to severe overfitting.

A Sequential Modeling Approach to Explain Complex Processes and Systems · Virginia Tech

Lost to a baseline

For Clinical+ perinatal risk prediction, classic Unweighted Logistic Regression achieved a higher test AUC (0.64, 95% CI: 0.59-0.69) than the tree-based Unweighted GBM (0.62, 95% CI: 0.57-0.67) and the Unweighted GBM-Based Ensemble (0.63, 95% CI: 0.59-0.69)

Advancing Evidence-Based Maternity Care: Empirical Studies of Technologies, Policies, and Clinical Practices in the U.S. and Abroad · Harvard

Tried and failed

stacked generalization ensembles with meta-learners applied to time-series crop yield prediction. Outcome: worse than baseline. Reason: violation of the IID assumption in temporal data despite using blocked sequential cross-validation

Optimized ensemble learning and its application in agriculture · Iowa State

Tried and failed

stacked generalization ensemble with neural networks applied to high-latitude ionospheric electron density modeling. Outcome: worse than baseline. Reason: Ensemble generalization degraded significantly on out-of-distribution test data relative to individual empirical physical models.

Topside Ionospheric Modeling using Machine Learning · Georgia Tech

Lost to a baseline

Stacking Mean Regression (SMR) produced higher error and variance than simpler Voting Mean Regression across all ensembles

Key Concepts and Methods in Point and Probabilistic Forecasting of Electricity Prices · IRIS - POLITO - prod

Lost to a baseline

Average weighted ensemble achieved lower RRMSE (8.87%) than optimized weighted ensemble (9.22%) in 2019 because LASSO performed poorly in training but well in the test year.

Decision support tools for farm management by integrating domain knowledge with machine learning and optimization techniques · Iowa State

Lost to a baseline

The first ensemble model (E1) was outperformed by both the standalone RNN and LSTM models on R2 (E1 achieved R2 0.8765 vs RNN 0.8977 and LSTM 0.9190) and error metrics (MSE 0.8532 vs RNN 0.8036, LSTM 0.7148; MAE 0.6705 vs RNN 0.6310, LSTM 0.4949)

Integrating Transactive Energy and Machine Learning For Re-Energizing Wastewater Treatment Plants · TXST Digital Repository

Lost to a baseline

Stacked ensemble model underperformed individual base models on broadleaf-only plots (SEM RMSE = 112.55 Mg ha-1 vs. GLM RMSE = 105.72 Mg ha-1, GBM RMSE = 111.41 Mg ha-1)

Advanced Techniques For Prediction of Forest Above Ground Biomass Using Satellite Remote Sensing Data · IRIS - UNITN - prod

Lost to a baseline

Averaging uncalibrated model scores in a machine-learning-only ensemble performed significantly worse than individual constituent models (Models 1, 2a, and 4).

Strong gravitational lenses in the era of wide-field surveys · Oxford

Lost to a baseline

NN ensembles retrained in an active learning loop with ensemble uncertainty-guided eABF-GaMD underperformed single NNs with GMM-based uncertainty in test set prediction error.

Enhancing Robustness of Neural Network Interatomic Potentials through Sampling Methods and Uncertainty Quantification · MIT

Lost to a baseline

Stacking ensemble model (96% accuracy) lost to standalone SVM classifier (100% accuracy) in at-risk student prediction

Gradgroom : an integrated framework for an educational early warning system with mentor matching · Institutional Repository University of Moratuwa

Considered and rejected

Considered and rejected: Rejected neural network stacking/merging of ensemble predictions because it required fixing the number of paths beforehand and performed worse or comparable to parameter-free inverse variance merging

Making Computer Vision Models Robust and Adaptive · EPFL

Considered and rejected

Considered and rejected: Ensemble modeling (SuperLearner combining RF, GLMNet, BART, XGBoost) was rejected in favor of random forest alone due to equivalent performance.

HARNESSING ANTIBODY KINETICS TO IMPROVE EPIDEMIOLOGIC INFERENCE: CASE STUDIES IN CHOLERA AND SARS-COV-2 · JScholarship

Advanced weighting and aggregation strategies fail to improve over simple averaging

8 theses · 7 institutions

Techniques designed to dynamically weigh or aggregate predictions often failed to deliver benefits over standard mean ensemble averaging. In some settings, aggregation weights collapsed to a single model, initial strong learners dominated subsequent ones, or parametric layers degraded output quality.

Tried and failed

k-means clustering for sub-ensemble selection applied to spatiotemporal trajectory forecasting. Outcome: worse than baseline. Reason: clustering selected sub-ensembles yielded little to no forecast error reduction over full ensemble

Understanding tropical cyclones using machine learning with satellite imagery · Imperial

Tried and failed

boosting with computed model weight aggregation applied to weakly supervised ensemble learning. Outcome: worse than baseline. Reason: initial strong learners dominated the ensemble predictions, preventing effective contribution from subsequent weak models

On the Resource Efficiency of Language Models · Georgia Tech

Tried and failed

ensemble averaging across multiple levels of theory applied to activation free energy prediction. Outcome: worse than baseline. Reason: averaging solvation energies across parameterisation levels did not improve accuracy over single-level calculations

Solvent design assisted by mechanistic insights: methods and application to peptide synthesis · Imperial

Tried and failed

weighted ensembling of tuned gradient boosting models applied to tabular regression target prediction. Outcome: worse than baseline. Reason: the ensemble weights collapsed to the single best performing model configuration without adding diversity

The Beautiful (Computer) Game: How Data Science Will Revolutionize the World's Most Popular Sport · UT Austin

Lost to a baseline

Stacked LASSO and Stacked LightGBM achieved lower RMSE (632 kg/ha and 756 kg/ha, respectively) than the Average Ensemble (761 kg/ha) for state-level corn yield forecasts.

Optimized ensemble learning and its application in agriculture · Iowa State

Considered and rejected

Considered and rejected: Rejected model fusion / ensemble averaging across classifiers because disparate model performance (U-Net/LightGBM vs XGBoost) reduced overall quality and destroyed interpretability.

Interpretacja danych geoprzestrzennych przy użyciu wyjaśnialnych metod uczenia maszynowego Interpretation of geospatial data using explainable machine learning methods · AMUR - Repozytorium Uniwersytetu im. Adama Mickiewicza w Poz

Tried and failed

learnable aggregation of expert prediction logits applied to multimodal mixture of experts ensemble. Outcome: worse than baseline. Reason: parametric aggregation layers underperformed simple mean averaging

Multimodal Land Cover Mapping from Remote Sensing Imagery, Species Observations, and Language · EPFL

Tried and failed

averaging kernel matrices before computing attribution applied to ensemble data attribution across models. Outcome: worse than baseline. Reason: averaging intermediate kernel representations degrades attribution quality compared to ensembling final attribution predictions

Machine Learning through the Lens of Data · MIT

Tree ensembles overfit or fail to generalize compared to regularized or linear models

3 theses · 3 institutions

Tree-based ensemble architectures overfit feature spaces and internal dataset characteristics relative to kernel methods or regularized linear models. In cross-dataset and physical regression tasks, this overfitting degraded out-of-sample generalization.

Tried and failed

merging decision trees trained on heterogeneous feature sets applied to ensemble regression modeling. Outcome: did not generalise. Reason: combining trees grown independently on relative versus absolute feature scales caused severe prediction errors

Processing and Analysis of the Seismocardiogram to Enable Estimations of Blood Volume Decompensation Status · Georgia Tech

Tried and failed

Tree-based ensemble regression models applied to crystal dielectric constant prediction. Outcome: overfit. Reason: Tree ensembles overfit the feature space compared to kernel methods like SVR and KRR

Materials Design of Complex Dielectric Crystals · Imperial

Tried and failed

non-linear ensemble models for speech emotion recognition applied to cross-dataset acoustic emotion prediction. Outcome: did not generalise. Reason: complex models overfit internal dataset characteristics compared to simpler regularized linear models

Creating Links: Building an Educational Platform to Ask Relevant Questions in Education · MIT

Linear models and simple baselines fail to capture complex non-linear patterns

6 theses · 5 institutions

Simpler linear or naive models repeatedly underperformed non-linear ensemble methods like Random Forests and Gradient Boosted Trees. They lacked the capacity to capture non-linear feature interactions and suffered when handling collinear predictors.

Tried and failed

regularized logistic regression for risk prediction applied to tabular pregnancy outcome data. Outcome: worse than baseline. Reason: linear models could not capture complex non-linear feature interactions as effectively as tree ensembles

Development and Validation of an Interpretable Risk Prediction Model for Stillbirth using Machine Learning · Harvard

Lost to a baseline

Linear regression model yielded significantly higher prediction error compared to Random Forest and AdaBoost ensemble models for conductivity prediction.

Real-time monitoring of neuronal cells through electrohydrodynamic patterning of flexible graphene microelectrodes · Iowa State

Lost to a baseline

Logistic regression (CV AUC 0.6673, test AUC 0.6613) and Naïve Bayes (CV AUC 0.4711, test AUC 0.4760) lost to non-linear ensemble models Gradient Boosted Trees (test AUC 0.6782) and Random Forest (test AUC 0.6727) on predicting Zone A reenlistment

PRICE ELASTICITY OF MONETARY INCENTIVES ON DRIVING DESIRED BEHAVIOR · Calhoun

Considered and rejected

Considered and rejected: Rejected linear regression from the final scATAC-Express ensemble due to poor predictive capacity and vulnerability to collinearity

Integrating genomic and multiomic data for computational analysis of gene regulation in circulating immune cells · Georgia Tech

Considered and rejected

Considered and rejected: Rejected Logistic Regression in the final ensemble because it added no AUC performance gain over tree-based models and required separate data preparation pipelines.

A DATA-DRIVEN APPROACH TO PREDICTING AUSTRALIAN BUSHFIRES · Calhoun

Considered and rejected

Considered and rejected: Rejected linear regression, Kernel Ridge, and LASSO due to inferior test MAE compared to ensemble tree methods (GBR, ETR, RFR).

Material discovery and modelling for solid-state hydrogen storage and fuel cell applications · University of Nottingham Repository

Ensemble inference incurs excessive computational overhead and latency

3 theses · 3 institutions

Deploying ensemble models often introduced severe computational and latency bottlenecks that impeded low-latency streaming and operational settings. The scaling of forward passes and model count made full ensembling impractical under tight real-time constraints.

Tried and failed

Ensemble learning algorithms applied to Low-latency streaming classification. Outcome: too slow. Reason: Inference time scaled unfavorably with ensemble size and sample count.

Driver distraction detection using experimental methods and machine learning algorithms. · Cranfield

Considered and rejected

Considered and rejected: Rejected 4DVAR and ensemble-based data assimilation (EnKF/3DEnVar) due to prohibitive operational computational costs and latency constraints.

The assimilation of surface observations in the European Alpine region · IRIS - UNITN - prod

Considered and rejected

Considered and rejected: Rejected Bayesian neural networks (e.g. MC dropout) for mapping uncertainty due to high computational overhead of multiple forward passes compared to small ensembles.

ACTIVE LEARNING OF VISION-BASED REPRESENTATIONS FOR ROBOTICS · Penn

Ensembling struggles with data quality issues and distributional assumptions

3 theses · 3 institutions

Ensembles degraded or produced negative performance metrics when subjected to misspecified parametric distribution assumptions or widespread missing diagnostic data across sites. Neural network ensembles also suffered from localized prediction failures in specific input ranges.

Tried and failed

regularized regression and tree ensembles applied to clinical outcome prediction. Outcome: worse than baseline. Reason: high rates of missing diagnostic predictor data across sites

APPLYING MATHEMATICAL MODELS TO IMPROVE CLINICAL EVALUATION AND PREDICTION · Penn

Tried and failed

deep neural network ensembles with input augmentation applied to tabular environmental sensor regression. Outcome: worse than baseline. Reason: localised predictive failure in specific input feature ranges, underperforming simple mean baseline

Trustworthy Soft Sensing in Water Supply Systems using Deep Learning · Virginia Tech

Tried and failed

ensemble model output statistics with parametric distributions applied to ensemble weather forecast postprocessing. Outcome: worse than baseline. Reason: distributional assumptions yielded poor fit, leading to degraded error metrics and negative R-squared values

Leveraging Artificial Intelligence and Numerical Weather Prediction to Build Custom Smart Home Software · Texas Tech

Left open by the authors

Problems the authors named and did not get to.

Left open

Develop ensemble prediction models combining machine learning and deep learning algorithms for ordinal disease incidence and hospitalisation targets. Blocker: None

Epidemiology and forecasting of influenza and COVID-19 · Imperial

Left open

Ensemble ESLSTM-MTF models initialized with diverse hyperparameters and custom loss functions that penalize extreme yield errors. Blocker: None

Enhancing Winter Wheat Crop Yield Predictions: A Data-Driven, Incremental and Integrative Approach with Machine Learning · Research Repository UCD

Left open

Implement ensemble methods combining fault detection and identification models trained on different feature subsets to improve prediction performance. Blocker: None

Fault Detection and Identification of Large-scale Dynamical Systems · MIT

Left open

Implement an evaluation pipeline comparing the predictive performance and resource usage of ensemble data monitors against periodic retraining baselines. Blocker: None

A Toolkit for Synthetic Data Generation and Drift Detection in Regression Scenarios · Carleton University Institutional Repository

Left open

Develop training objectives for diverse ensembles to allow sub-models to use efficient architectures at reduced training cost. Blocker: No specific mathematical formulations or target metrics are provided beyond a broad research direction

Towards Efficient and Robust Deep Neural Network Models · DukeSpace

Left open

Develop hybrid or ensemble forecasting models combining multiple RNN architectures with statistical models for multi-step wind power forecasting. Blocker: None

Deep Learning-Based Medium to Long-Term Multi-Step Ahead Wind Power Generation Forecasting · TXST Digital Repository

Left open

Integrate random forest, neural networks, CART, LASSO, ridge, and PCA into the SuperLearner-hdPS ensemble framework for small sample sizes. Blocker: None

Real-World Bleeding with Ibrutinib in B-Cell Malignancies · Penn

Left open

Implement and integrate transformer-based architectures or ensemble techniques into the solar power prediction pipeline. Blocker: None

Sustainable Solar-Powered EV Charging System Design Using Machine Learning, DC Fast Charging, and an Intelligent DMPPT Optimization Technique · Texas Tech

Left open

Analyze and identify root causes for the ensemble transfer model's performance drop on the PU6 and PU8 datasets. Blocker: None

Semantic Segmentation of Satellite Imagery using Positive and Unlabeled Learning. · Scholars' Bank

Left open

Fit and prune deeper regression trees on the ensemble HFT probability scores to extract specific high-frequency trading strategies. Blocker: Requires proprietary order book data with labeled or identifiable HFT orders.

Three Essays on Financial Economics · IRIS - UNITN - prod

Checking a claim in this area?

We can run the same search on any method or claim. If nothing turns up, we will say so, and that proves nothing on its own.