Chapter Four · failure evidence
What Deep Learning & Pretraining got wrong, from 23 dissertations
The records evaluate deep learning and pretrained language models across biomedical, tabular, language, and forecasting applications. Although deep neural methods offer powerful capabilities, they frequently suffer from overfitting on small datasets, fail against simpler classical baselines, or face domain adaptation and deployment limitations. These records come from PhD theses at 15 institutions, 2021 to 2026. Each links to its thesis. They were extracted by language models reading the full text, so treat each as a lead to read, not a verdict.
Deep learning models are rejected due to data scarcity and high overfitting risk in small sample regimes
Deep learning models were rejected across behavioral, biomedical, tabular, chemometric, and electroencephalogram studies because small sample sizes created a severe risk of overfitting. Authors noted that deep neural networks required excessive amounts of data and computational overhead while offering poor interpretability compared to classical methods.
Considered and rejected
Considered and rejected: Rejected using complex deep learning techniques without multi-task framing on small-sample behavioral datasets due to severe overfitting and poor generalizability.
Towards Human-Centered Behavioral Sensing for Student Support · ResearchWorks
Considered and rejected
Considered and rejected: Rejected Deep Learning (DL) models due to high risk of overfitting on small biomedical sample sizes, computational overhead, and lack of interpretability.
Considered and rejected
Considered and rejected: Rejected OpenNN and deep learning frameworks (Keras/TensorFlow) due to narrow focus on neural networks and excessive data requirements for small tabular data
Fatigue Monitoring System · SUPSI - ARIS
Considered and rejected
Considered and rejected: Rejected deep learning approaches due to limited training data and high susceptibility to overfitting in small EEG sample regimes.
Instrumentation for daily-life brain-computer interfaces · IRIS - POLITO - prod
Considered and rejected
Considered and rejected: Deep learning / ANN models were deprioritized over PLS/SVM due to high chemometric collinearity and small sample size ('small n, large p' regime) risking overfitting without massive datasets.
Sensor Approaches for the Non-Destructive Assessment of Food Safety and Authenticity · Cranfield
Deep architectures underperform or are rejected in favor of simpler classical baselines
Deep learning approaches were beaten by or rejected in favor of simpler models such as Random Forest, logistic regression, and optical flow interpolation. These classical baselines provided lower whole-image error, better false positive reduction, competitive tabular performance, and faster polynomial-time explainability.
Tried and failed
deep learning on raw signals and scalograms applied to arterial Doppler waveform classification. Outcome: worse than baseline. Reason: deep models underperformed handcrafted feature-based logistic regression classifier
Testing for peripheral arterial disease in diabetes · Imperial
Considered and rejected
Considered and rejected: Deep learning architectures rejected in favor of Random Forest due to competitive tabular performance and fast polynomial-time SHAP computation via TreeExplainer.
Considered and rejected
Considered and rejected: Rejected using deep learning architecture over Random Forest classifier because Random Forest was simpler and reduced false positives by 50% vs 35% for DNN.
Molecular simulation on two-dimensional nanoporous material for energy-efficient separation · EPFL
Lost to a baseline
Linear and optical flow interpolation yielded lower whole-image MSE than deep learning methods (R2UNet/proposed) because the latter deprioritized background non-ROI pixels.
Controllable synthetic algorithms and evaluations for annotated clinical image synthesis · Imperial
Domain adaptation and specialized pretrained representations degrade downstream performance
Applying domain fine-tuned language models to time series forecasting led to performance degradation across all horizons and metrics relative to standard deep learning models. Similarly, domain-adapted language representations underperformed general BERT and GPT-2 on summarization and degraded sentiment classification recall when trained on raw chatbot text.
Tried and failed
domain fine-tuned zero-shot language model applied to time series forecasting. Outcome: worse than baseline. Reason: exhibited performance degradation across all horizons and metrics compared to standard deep learning models
Deep Learning Methods for Built Environment Operational Management · Virginia Tech
Tried and failed
Domain-specific pretrained language model embeddings applied to Extractive text summarization. Outcome: worse than baseline. Reason: Domain-adapted BERT underperformed general domain BERT and GPT-2 on extractive summarization metrics
Tried and failed
pretrained language model embeddings on raw text applied to chatbot text sentiment classification. Outcome: worse than baseline. Reason: raw chatbot-generated text representations degraded classification recall compared to customer text alone
Three essays on strategic digital initiatives in modern organizations · Georgia Tech
Fine-tuning and scaling pretrained transformers causes overfitting or underperforms smaller models
Fine-tuning pretrained transformers for continuous regression suffered from fluctuating validation loss in later epochs and failed to generalize to unseen data. Furthermore, larger pretrained transformer embeddings underperformed smaller DistilBERT representations across all evaluated downstream text classification tasks.
Tried and failed
fine-tuning pretrained transformer for continuous regression applied to short text evaluation scores. Outcome: overfit. Reason: validation loss fluctuated in later epochs, failing to generalize to unseen data
Tried and failed
larger pretrained transformer embeddings applied to fine-grained text classification. Outcome: worse than baseline. Reason: DistilBERT outperformed full BERT representations across all tested downstream classifiers.
Spatiotemporal Event Forecasting and Analysis with Ubiquitous Urban Sensors · Virginia Tech
Pretrained models and large deep architectures are rejected due to temporal leakage and latency limits
Off-the-shelf contemporary large language models were rejected for historical financial forecasting because severe temporal pretraining leakage invalidated their use. In addition, full-scale GPU and cloud deep learning baselines were rejected because they flattened features and failed to optimize for low edge latency.
Considered and rejected
Considered and rejected: Rejected off-the-shelf contemporary LLMs for historical financial forecasting due to severe temporal pretraining leakage.
Considered and rejected
Considered and rejected: Rejected comparing against full-scale GPU/cloud deep learning baselines (Transformers/CNN-LSTM) because they treat input features as flat vectors and do not optimize for low edge latency
LIDIT: Low-Latency Intrusion Detection in IoMT Devices using TinyML · Queens University Institutional Repository
Left open by the authors
Problems the authors named and did not get to.
Left open
Explore and evaluate improved deep learning architectures for classifying Electronic Theses and Dissertations. Blocker: None
Classifying ETDs · Virginia Tech
Left open
Tune hyperparameters and architectures for deep learning emergency department arrival forecasting models across 1h to 48h horizons. Blocker: Requires the emergency department arrival dataset used in the thesis, which is likely private hospital data.
Left open
Evaluate pretrained transformer models with subword tokenization and transfer learning on scraped job postings to predict bankruptcy and corporate growth. Blocker: Scraped job listings dataset may not be publicly archived with the thesis
Assessing Corporate Growth and Bankruptcy Risk Using Public Data Proxies · Harvard
Left open
Apply transfer learning and transformer architectures to improve generalization of deep learning models across diverse seismic datasets. Blocker: The objective is a broad research direction lacking specific datasets, target tasks, or evaluation metrics.
Improving accuracy and efficiency of seismic data analysis using deep learning · UT Austin
Left open
Develop deep learning methods to reduce the computational burden of real-time SDRE feedback stabilization for nonlinear PDEs. Blocker: The proposal lacks any concrete model architecture, training formulation, or target performance metrics.
State-dependent Riccati equation feedback stabilization for nonlinear PDEs · Imperial
Left open
Evaluate deep learning architectures against shallow models for GHG and odour emissions from biological nutrient removal wastewater data. Blocker: Lacks access to the specific cold-region wastewater treatment plant field and laboratory monitoring dataset.
Left open
Develop tabular deep learning architectures or evaluate feature selection across wider combinations of cardinality and coefficient of variation. Blocker: None
Data Value Analytics for Feature Selection and Dataset Retrieval · Research Repository UCD
Checking a claim in this area?
We can run the same search on any method or claim. If nothing turns up, we will say so, and that proves nothing on its own.