Chapter Four · failure evidence
What Probabilistic Latent Modeling got wrong, from 67 dissertations
The records document recurring failures and trade-offs encountered when applying probabilistic latent variable models across text, spatial, tabular, and dynamical data. Researchers frequently find that standard latent modeling approaches struggle with data sparsity, distribution misspecification, optimization instability, and competition from simpler deterministic baselines. These records come from PhD theses at 24 institutions, 2021 to 2026. Each links to its thesis. They were extracted by language models reading the full text, so treat each as a lead to read, not a verdict.
Latent Dirichlet Allocation fails on short or sparse text documents
Short texts such as tweets, survey responses, comments, and search queries lack sufficient word co-occurrence and document length for Latent Dirichlet Allocation to infer coherent topic representations. As a result, pairwise similarities diverge from human judgment, topic coherence degrades, and the method is consistently rejected in favor of mixture models or dense neural embeddings.
Tried and failed
Latent Dirichlet Allocation topic modelling applied to short-text survey responses. Outcome: worse than baseline. Reason: standard LDA struggles with sparsity and lack of word co-occurrence in short texts
Examining structural, systemic, and enabling approaches to sustainable food system transformations · Imperial
Tried and failed
Latent Dirichlet Allocation topic modeling applied to short social media texts. Outcome: worse than baseline. Reason: sparse word co-occurrences in short documents led to poor topic coherence compared to NMF and neural embeddings
Assessing mental wellbeing in urban areas using social media data: understanding when and where urbanites stress and de-stress · Georgia Tech
Considered and rejected
Considered and rejected: Rejected Latent Dirichlet Allocation (LDA) text topic features due to poor perplexity/interpretability on 89 profile tweet aggregations.
Personality-Driven Social Media Curation: How Personality Traits Affect Following Decisions on Twitter · Scholars' Bank
Considered and rejected
Considered and rejected: Rejected Latent Dirichlet Allocation (LDA) for short sanction text messages, choosing biterm topic modeling instead.
Three Papers on Peer Sanctioning, its Evaluation, and its Justification · DukeSpace
Tried and failed
Latent Dirichlet Allocation for semantic similarity applied to short text categorization and clustering. Outcome: no signal. Reason: LDA pairwise similarity matrices differed substantially from human judgment correlation matrices
Understanding Cyber Attacks Using Text Mining · Texas Tech
Tried and failed
Latent Dirichlet Allocation topic modeling applied to short user comment text corpus. Outcome: no signal. Reason: failed to reveal distinct latent themes or nuanced categorization across user comments
Public Use of Open Access Research: Evidence from the National Academies and Harvard DASH Repository · Georgia Tech
Tried and failed
Latent Dirichlet Allocation topic modeling applied to short conversational text queries. Outcome: no signal. Reason: text length was too brief and sparse to extract coherent topics
Evaluating the explainability of AI-driven clinical decision support tools · Imperial
Considered and rejected
Considered and rejected: Rejected Latent Dirichlet Allocation (LDA) in favor of the Gibbs Sampling Dirichlet Multinomial Mixture (GSDMM) model for topic modeling due to GSDMM's superior handling of short, dense tweet text
Decision support for sustainable energy systems, energy economics, urban mobility, and emerging methods in information systems research · Leibniz Universität Hannover Repository
Considered and rejected
Considered and rejected: Rejected Latent Dirichlet Allocation (LDA) for short survey priority responses due to poor performance on documents with fewer than 50 words.
Examining structural, systemic, and enabling approaches to sustainable food system transformations · Imperial
Considered and rejected
Considered and rejected: Rejected standard Latent Dirichlet Allocation (LDA) for clustering short extracted action spans due to severe sparsity and inability of synonyms to co-occur in 1-10 word spans.
Considered and rejected
Considered and rejected: Rejected Latent Dirichlet Allocation (LDA) for query topic modeling because short search queries (avg 5.2 words) yield sparse vectors and ignore word sequence order.
EFFECTS OF INFORMATION ON DIGITAL PLATFORMS AND ONLINE ECOSYSTEMS · Penn
Linear factor models and principal component analysis fail to capture genuine latent factor structures
Linear projections and principal component methods frequently capture sample noise instead of true weak factors, resulting in poor out-of-sample generalization and failure to converge to true components. Furthermore, deterministic principal components fail to account for measurement error, time-series structure, or parameter uncertainty when modeling underlying latent factors.
Tried and failed
non-negative and empirical Bayes matrix factorization applied to variant-phenotype cluster decomposition. Outcome: worse than baseline. Reason: failed to improve biological relevance of clusters compared to Latent Dirichlet Allocation
Characterizing Regulatory Elements and Non-Coding Variants in the Human Genome · Harvard
Tried and failed
projecting latent factor decompositions to recover unprojected components applied to stochastic discount factor decomposition. Outcome: did not converge. Reason: projected weak-factor and idiosyncratic components failed to converge to their true unprojected counterparts
Asset pricing with unsystematic risk · Imperial
Tried and failed
retaining high-order principal components for weak signals applied to linear factor asset pricing models. Outcome: did not generalise. Reason: high-order components captured primarily sample-specific noise rather than true weak latent factors, causing severe out-of-sample degradation
Asset pricing with unsystematic risk · Imperial
Lost to a baseline
Linear PCA dimensionality reduction on agent-based stochastic data found a latent space size of n - 1 (25x larger than non-linear DR), resulting in poor reducibility compared to non-linear methods.
Reduced Order Non-INtrusive (RONIN) Modeling for Strategic Defense Planning · Georgia Tech
Considered and rejected
Considered and rejected: Principal component analysis was rejected in favor of common-factor analysis (unweighted least squares) because error variance was unknown and generic latent constructs were sought
Considered and rejected
Considered and rejected: Rejected non-parametric Principal Component Analysis (PCA) for latent factor extraction due to lack of time-series structure and inability to propagate parameter uncertainty
Parameter Uncertainty, Cashflow Betas, and Earnings Announcement Premia · Scholars' Bank
Considered and rejected
Considered and rejected: Rejected principal component analysis (PCA) because it cannot model underlying latent factors via covariance
An Exploration of Fear of Sleep and Experiential Avoidance in the Context of PTSD and Insomnia Symptoms · Scholars' Bank
Considered and rejected
Considered and rejected: Rejected factor analysis (FA) in favor of PCA for factor score generation to describe empirical dataset variance without imposing latent variable normality assumptions.
Whole-water systems modelling for sustainable catchment management · Imperial
Considered and rejected
Considered and rejected: Deterministic PCA with arbitrary truncation threshold (rejected because it incorporates undesired noise and lacks probabilistic modeling of model reduction error)
Considered and rejected
Considered and rejected: Rejected Principal Component Analysis (PCA) in favor of Exploratory Factor Analysis (EFA) because PCA does not model latent constructs with conceptual meaning and assumes zero measurement error.
Evaluating the Efficacy of Talent Identification and Development in the National Hockey League Entry Draft · YorkSpace
Variational autoencoders underperform simpler deterministic and linear baselines
Probabilistic latent representations from variational autoencoders frequently underperform simpler alternatives like deterministic autoencoders, principal component analysis, and nearest neighbor baselines in reconstruction and prediction tasks. Adding variational inference overhead often increases prediction uncertainty, degrades cluster separation, and fails to provide benefits over standard point-estimate baselines.
Tried and failed
sequence-to-sequence variational autoencoder applied to sketch to latent space mapping. Outcome: worse than baseline. Reason: failed to show significant improvement over simpler vector KNN regression despite higher training complexity
Tried and failed
surrogate modeling on variational autoencoder latent space applied to property prediction from spatial representations. Outcome: worse than baseline. Reason: nonlinear latent embedding degraded forward predictions and increased uncertainty compared to linear dimensionality reduction
Neural Inverse Microstructure Design with Bayesian Scale-Bridging · Georgia Tech
Tried and failed
variational probabilistic embedding applied to human comparison choice modeling. Outcome: worse than baseline. Reason: the underlying choice likelihood model is identical to point-estimate maximum likelihood estimation
Lost to a baseline
In pure embedding quality (Recall@1), probabilistic models slightly lagged behind standard supervised Cross-Entropy baseline on ViT-Medium
Uncertainties of Latent Representations in Computer Vision · Publikationssystem UB Tuebingen
Lost to a baseline
VAE + Gaussian process regression on latent variables lost to a simple null baseline (predicting test QY using the most similar sequence in the training set), which achieved matching correlation.
INVESTIGATING THE STRUCTURAL REQUISITES OF PHOTOPHYSICAL PROPERTIES IN FLUORESCENT BIOMOLECULES · JScholarship
Lost to a baseline
VAE reconstruction MSE underperformed PCA and NMF when using ≥ 5 latent parameters
Characterising & classifying the local population of ultracool dwarfs with Gaia DR2 and EDR3 · Imperial
Lost to a baseline
AE and VAE Latent Noise Segmentation achieved lower ARI (0.926 and 0.918) than simple noisy reconstruction clustering (0.982 and 0.977) on the Gradient Occlusion dataset.
Differences in Visual Perception in Humans and Deep Neural Networks · EPFL
Lost to a baseline
Large VAE with latent size 128 (44.3 AUPRC) was beaten by PCA difference vector baseline (67.3 AUPRC) on the Floods dataset with k=3 history.
Considered and rejected
Considered and rejected: Rejected stochastic latent sampling in VAEs, choosing deterministic autoencoders with explicit decoder regularization and ex-post density estimation.
Reining in the Deep Generative Models · Publikationssystem UB Tuebingen
Considered and rejected
Considered and rejected: Rejected standard VAE in favor of deterministic autoencoder because generative latent distribution modeling was unnecessary for embedding compression.
Neural Compression for Scalable Question-Answer Retrieval · DalSpace
Variational latent variable models suffer from training instability, sampling issues, and posterior collapse
Enforcing sparsity, autoregressive covariance, or stochastic sampling in variational latent spaces often leads to unstable training and fails to resolve posterior collapse. In addition, random latent sampling in sparse data regimes causes low generation success rates, while inference via Monte Carlo sampling or latent inversion introduces high computational complexity and convergence failures.
Tried and failed
L1 or L2 regularization on posterior parameters applied to variational autoencoder latent representations. Outcome: worse than baseline. Reason: Failed to increase representation sparsity compared to vanilla variational autoencoder baseline.
Injecting Inductive Biases into Distributed Representations of Text · Cambridge
Tried and failed
variational autoencoder for sparse latent representation applied to text sentence embeddings. Outcome: unstable. Reason: struggled to achieve steady and consistent sparsity across hyperparameter configurations
Injecting Inductive Biases into Distributed Representations of Text · Cambridge
Tried and failed
autoregressive latent covariance structure applied to variational autoencoders. Reason: did not prevent posterior collapse when it occurred in standard VAE baseline
Tried and failed
i.i.d. latent sampling in variational autoencoders applied to set and graph generation. Reason: It optimizes an excessively loose lower bound on the true evidence lower bound.
Equivariant Neural Architectures for Representing and Generating Graphs · EPFL
Lost to a baseline
Separately trained predictor and VAE required significantly more optimization steps and exhibited higher variance during latent space search compared to simultaneously trained VAE with predictor
Considered and rejected
Considered and rejected: Rejected Monte Carlo sampling from the marginal word distribution during CVAE inference in favor of feeding expected latent alignments to avoid high computational complexity
LOOKING INTO ACTORS, OBJECTS AND THEIR INTERACTIONS FOR VIDEO UNDERSTANDING · JScholarship
Tried and failed
latent space sampling in variational autoencoder applied to random 3D mesh generation. Outcome: data insufficient. Reason: training data sparsity led to low valid generation success rate when sampling the latent space randomly
Tried and failed
latent space inversion via optimization and encoder applied to deep progressive generative adversarial networks. Outcome: did not converge. Reason: optimization and encoder-based methods fail to accurately invert deep multi-scale generative models
Considered and rejected
Considered and rejected: Rejected Variational Autoencoders (VAEs) for sound-masking noise generation because GANs avoid deterministic bias and perform better with discrete latent variables.
Robust Defenses Against Adversarial Machine Learning in Internet of Things Security · Carleton University Institutional Repository
Standard Gaussian latent priors cause distributional mismatch and over-smoothing
Assuming continuous Gaussian priors creates dense and overlapping latent spaces that hinder distinct cluster separation and degrade performance on non-Gaussian or tabular systems. In dynamical and fluid modeling, the enforcing of smooth Gaussian latent spaces leads to underpredicted kinetic energy and loss of high-frequency spectral content.
Lost to a baseline
Vanilla autoencoders achieved slightly higher silhouette scores for cluster separation compared to VAE due to VAE's continuous, overlapping latent space.
Machine Understanding of Architectural Space: From Analytical to Generative Applications · EPFL
Considered and rejected
Considered and rejected: Rejected enforcing a Gaussian prior via adversarial discriminators on the style latent bottleneck due to training instability and degraded RMSE
Using wearable sensors and machine learning to enrich lower limb assistive devices · ResearchWorks
Tried and failed
conditional variational autoencoder with factorised latent space applied to 3D multi-object scene completion. Reason: failed to learn a structured prior, yielding incomplete reconstructions when sampling from Gaussian noise
Scene understanding for 3D multi-object scenes: labelling, reasoning and decomposing · Imperial
Tried and failed
variational autoencoders with Gaussian priors applied to systems with strongly non-Gaussian statistics. Outcome: did not generalise. Reason: standard Gaussian latent priors poorly represent strongly non-Gaussian statistics
Physics-Driven Machine Learning for Applications in Geophysical Fluid Dynamics · MIT
Tried and failed
variational autoencoder for dimensionality reduction before clustering applied to high-dimensional mass spectrometry imaging data. Reason: Gaussian prior forced an overly dense latent space, preventing separation into distinct clusters
Mass spectral imaging of clinical samples using deep learning · Imperial
Tried and failed
variational autoencoder transformer reduced order model applied to turbulent fluid flow prediction. Reason: Smooth latent representations underpredicted kinetic energy and lost high-frequency spectral content
A Dynamic Model of Unpiloted Aerial Systems for Complex Ship Airwake Environments · Georgia Tech
Lost to a baseline
Variational Autoencoders (VAEs) lose to standard Autoencoders (AEs) when training SNR and testing SNR perfectly match due to gaps in the latent space
Model and data driven approaches to wireless image transmission · Imperial
Considered and rejected
Considered and rejected: Rejected Variational Autoencoders (VAEs) in FASTER-CE to avoid assuming Gaussian distributions over tabular latent spaces
Multi-objective approaches towards trustworthy machine learning · UT Austin
Latent Dirichlet Allocation yields uninterpretable, unstable, or poorly aligned topic structures
Unsupervised topic models often generate topics filled with generic, loosely related words that fail to map onto predefined domain categories or theoretical constructs. Models also exhibit instability over longitudinal corpora, where topic tracking is sensitive to initial text sequence ordering and fails to produce consistent results across time.
Tried and failed
Latent Dirichlet Allocation topic modeling applied to organizational text corpora. Outcome: no signal. Reason: generated uninterpretable topics with generic words failing to map to distinct theoretical constructs
Tried and failed
unsupervised Latent Dirichlet Allocation applied to text classification against predefined taxonomy. Reason: unsupervised topic clusters did not correspond one-to-one with predefined target taxonomy categories
Analytics-Enabled Quality and Safety Management Methods for High-Stakes Manufacturing Applications · MIT
Considered and rejected
Considered and rejected: Rejected using Latent Dirichlet Allocation (LDA) for document topic modeling due to non-accurate and inconsistent inferring results.
Using Machine Learning And Natural Language Processing To Improve Scientific Processes · Penn
Considered and rejected
Considered and rejected: Rejected Latent Dirichlet Allocation (LDA) for topic modeling due to reliance on bag-of-words, requirement of prior topic number determination, and ignoring word ordering/semantics in favor of Top2Vec.
Three Essays on the Use of Health Information Technology to Improve Patient Safety: Cases of Technology Design and Participation in Professional-Only Healthcare Forums · DSpace at SUNY Buffalo
Considered and rejected
Considered and rejected: Rejected Latent Dirichlet Allocation (LDA) for issue report topic extraction and review aspect extraction because it produces loosely-related words when vocabulary is large
Machine Learning Driven Guidance for Software Maintenance: Enhancing Code Management and Features · Queens University Institutional Repository
Considered and rejected
Considered and rejected: Rejected Latent Dirichlet Allocation (LDA) in favor of Non-negative Matrix Factorization (NMF) due to NMF's higher topic coherence and model stability.
Essays on central bank communication · UT Austin
Tried and failed
undirected latent Dirichlet allocation applied to longitudinal corporate annual report text. Outcome: unstable. Reason: critically dependent on initial text sequence and could not reliably track recurring topics across decades
Gaussian process latent models and Bayesian optimization surrogates fail in dynamic or complex settings
Approximations to deep Gaussian process posteriors produce poor uncertainty quantification and inaccurate variance estimates for acquisition functions in Bayesian optimization. Standard Gaussian process regression also fails to track time-varying environmental shifts and incurs severe hyperparameter complexity in high-dimensional spaces.
Tried and failed
Gaussian approximation of deep Gaussian process posteriors applied to Bayesian optimization uncertainty estimation. Reason: yielded poor uncertainty quantification and inaccurate variance estimates for acquisition functions
Physics-informed Machine Learning for Digital Twins of Metal Additive Manufacturing · Virginia Tech
Tried and failed
Standard Bayesian optimization without contextual modeling applied to optimization under dynamic periodic environmental drift. Outcome: did not converge. Reason: Standard Gaussian process regression could not track or account for dynamic time-varying environmental context shifts
Machine Learning Applications for Improving Accelerator Operations · Cornell
Lost to a baseline
Gaussian Process surrogate underperformed GradientBoosting during Bayesian optimization, dipping below p=0.05 around iteration 50.
Data-Driven Design of Recycling-Friendly Aluminium Alloys · MIT
Lost to a baseline
One latent GP framework performed worse (median RMS = 32.9 m s-1) than the baseline FF' model (RMS = 28.5 m s-1) on synchronous synthetic test data.
The epoch of giant planet migration : searching for young planets within the stellar noise · UT Austin
Considered and rejected
Considered and rejected: Rejected Gaussian process latent variable models (Bayesian approach) for mocap synthesis due to hyperparameter complexity, kernel selection issues, and loss of nuances in high-dimensional space
Predicting lower limb kinematics and kinetics from internal measurement units using deep learning · Imperial
Left open by the authors
Problems the authors named and did not get to.
Left open
Perform biological interpretation and functional annotation of latent topics and theme k-mers identified from shotgun metagenomic topic models. Blocker: None
Microbiome and Metagenomics: Statistical Methods, Computation and Applications · Penn
Left open
Investigate the causal implications and identification strategies of latent topics extracted from text data. Blocker: Lacks specific causal hypotheses, targets, identification strategies, or experimental frameworks
LEARNING AND INFERENCE IN DIGITAL MARKETS: METHODS AND APPLICATIONS · Cornell
Left open
Benchmark co-segregation-based Bayesian linear regression feature selection against standard methods like PCA and random forests for microbial phenotypic GWAS. Blocker: None
Probing Genomic, Metabolic, And Phenotypic Evolution In Microbes Using Comparative And Experimental Evolution Data · Georgia Tech
Left open
Develop a probabilistic model of microbial biomass growth to scale column-level dynamics to ecosystem-level microbial community behavior. Blocker: Lack of mathematical formulation, target datasets, and specific probabilistic framework defined in the thesis
Investigation of microbial community dynamics in soil during variable hydrological forcing · EPFL
Left open
Extract literature corpora from academic databases beyond Scopus to analyze topic trends via Latent Dirichlet Allocation. Blocker: None
Error component models for traffic crash severity analysis · Texas Tech
Left open
Incorporate monotonicity, fairness, or probabilistic/Bayesian priors into constrained low-rank approximation frameworks. Blocker: Lacks specific mathematical formulations, optimization objectives, and evaluation datasets or metrics.
Algorithms for Data Fusion, Representation Learning, and Scalable Clustering based on Constrained Low-Rank Approximation · Georgia Tech
Left open
Initialize variational autoencoders or graph-based multi-scale clustering algorithms with latent factors learned from Integrative Hierarchical Poisson Factorisation. Blocker: None
Left open
Extract high-level semantic features from visual stimuli to model and explain latent aesthetic taste criteria among collaborative filtering user cohorts. Blocker: None
Left open
Develop unsupervised methods to select subsets of latent dimensions in variational autoencoders for anomaly detection. Blocker: None
Anomaly Detection via Latent Variables Learned by Variational Autoencoders · Queens University Institutional Repository
Left open
Model agent intent using a continuous latent space with variational autoencoders for approximate inference in multiagent reinforcement learning. Blocker: None
Robust and Scalable Multiagent Reinforcement Learning in Adversarial Scenarios · MIT
Checking a claim in this area?
We can run the same search on any method or claim. If nothing turns up, we will say so, and that proves nothing on its own.