Chapter Four · failure evidence
What Bootstrap Resampling & Permutation got wrong, from 82 dissertations
The records document numerous empirical and theoretical failures when applying bootstrap resampling and permutation testing across diverse statistical models. These difficulties include severe computational bottlenecks, breakdown under dependent or non-exchangeable structures, asymptotic inconsistency at parameter boundaries, and poor confidence interval coverage. These records come from PhD theses at 26 institutions, 2021 to 2026. Each links to its thesis. They were extracted by language models reading the full text, so treat each as a lead to read, not a verdict.
Permutation testing and bootstrap resampling incur prohibitive computational overhead
Researchers frequently rejected permutation testing and bootstrapping because repeatedly retraining models, generating null distributions, or evaluating combinatorial permutations created unsustainable runtimes. Iterating these resampling procedures proved especially intractable for high-dimensional spike matrices, large neuroimaging cohorts, genomic sequences with high missingness, and complex ensemble classifiers.
Considered and rejected
Considered and rejected: Bootstrapping for regression standard errors rejected due to computational infeasibility with minimal difference
A Computational Approach to Recontextualization in Human Reading Behavior · Harvard
Tried and failed
Permutation-based pointwise contrast testing applied to functional data analysis contrasts. Outcome: too slow. Reason: Excessive computation time and conservative empirical test sizes
Functional Data Models for Raman Spectral Data and Degradation Analysis · Virginia Tech
Tried and failed
phylogenetic bootstrap inference on high-missingness sequence alignments applied to large reduced-representation genomic datasets. Outcome: too slow. Reason: bootstrapping datasets with 60% or 70% missing data was computationally intractable
Considered and rejected
Considered and rejected: Permutation-based significance testing for every regional learner due to extreme computational overhead (two to three orders of magnitude slower than analytic approximation).
Advancing Statistical Inference For Population Studies In Neuroimaging Using Machine Learning · Penn
Considered and rejected
Considered and rejected: Rejected Threshold-Free Cluster Enhancement (TFCE) permutation testing for the 13,680-subject UK Biobank cohort due to excessive computational burden.
Advances in statistical methods for large-scale binary-valued neuroimaging data · Oxford
Considered and rejected
Considered and rejected: Rejected standard permutation testing across sample indices for large sample sizes in KMERF due to high computational overhead of retraining forests per permutation, choosing a chi-square test approximation instead.
Random Forest for Hypothesis Testing: Development and Application to Cancer Detection · JScholarship
Considered and rejected
Considered and rejected: Rejected non-parametric local moderated t-statistic in cluster-based permutation inference due to intractable computational burden (10 hours for 1000 permutations on 40 cores).
Bayesian Methods for Analyzing Functional Magnetic Resonance Imaging Data · JScholarship
Considered and rejected
Considered and rejected: Rejected using all possible permutations in DBFinder due to conservative FDR penalization in small sample sizes and computational cost, opting instead for balanced permutations.
Statistical Methods for Mobile Health and Genomics Data · Harvard
Considered and rejected
Considered and rejected: Rejected permutation-based surrogate data testing for high-dimensional spike-field matrices due to extreme computational expense with increasing recording dimensionality, replacing it with an analytical Marchenko-Pastur bound.
Brain as a Complex System, harnessing systems neuroscience tools & notions for an empirical approach · Publikationssystem UB Tuebingen
Considered and rejected
Considered and rejected: Rejected random shuffling permutation for feature sensitivity due to excessive computational overhead from multiple shuffle iterations.
Fairness specification and repair for machine learning pipeline · Iowa State
Considered and rejected
Considered and rejected: Rejected standard permutation testing of retraining entire classifiers for thousands of iterations in MIGHT due to extreme computational inefficiency, adopting two-forest tree swapping instead.
Random Forest for Hypothesis Testing: Development and Application to Cancer Detection · JScholarship
Considered and rejected
Considered and rejected: Rejected running JLIM null distributions separately for each individual cell in scJLIM due to prohibitive computational expense; replaced with a single precomputed null distribution of N=100,000 permutations shared across cells
Computational methods for dissecting multicellular mechanisms of complex diseases · MIT
Resampling fails when data violate exchangeability or possess complex dependency structures
Standard and discrete bootstrap methods failed or were abandoned when data contained temporal autocorrelation, spatial clustering, network dependence, or correlated predictors. Unstructured permutations broke joint feature distributions and temporal trajectories, leading to severe undercoverage, inflated Type I errors, and invalid null distributions.
Tried and failed
naive two-sample bootstrap for Wasserstein distance applied to two-sample hypothesis testing. Reason: inconsistent for null limit distribution due to extra covariance drift, requiring pooled distribution resampling
Robust and Geometry-Aware Machine Learning Using Optimal Transport · Cornell
Considered and rejected
Considered and rejected: Rejected standard permutation testing due to non-exchangeability of correlated error structures and severe computational burden, opting instead for residual projection and Pearson type III moment-matching approximations
Generalization of kernel machine methods for association testing of multi-omics data · ResearchWorks
Tried and failed
standard and cluster bootstrap for prediction intervals applied to sequential game outcome probability forecasting. Outcome: did not generalise. Reason: complex temporal and sequential dependencies cause severe undercoverage of nominal prediction sets
MOVING BLACK-BOXING TOWARDS STATISTICS: CASE STUDIES FROM AMERICAN FOOTBALL · Penn
Tried and failed
moving block bootstrap with self-normalization statistic applied to GARCH time series mean hypothesis testing. Outcome: worse than baseline. Reason: failed to improve coverage accuracy and degraded performance under several GARCH(1,1) conditions
Tried and failed
unit-level bootstrap variance estimation applied to bipartite network causal inference. Reason: outcome-level resampling fails to account for network dependence across intervention units, severely underestimating standard errors
Considered and rejected
Considered and rejected: Standard/basic bootstrap rejected for time-series panel data in favor of block bootstrap because standard bootstrap fails to capture time-dependence structure.
Educational Disparities in Chronic Pain and Life Expectancy: Gaps and Pathways · DSpace at SUNY Buffalo
Considered and rejected
Considered and rejected: Discrete bootstrapping across independent observations, rejected due to poor support of the continuous exposure and failure to respect spatiotemporal correlation
Advancing Data Science Methods for Environmental Health Policy Design and Evaluation · Harvard
Considered and rejected
Considered and rejected: Rejected full independent sample bootstrapping in favor of block bootstrapping treating each country as an independent block to avoid invalid independence assumptions across time
The Impact of Culture on Non-Life Insurance Consumption · Penn
Considered and rejected
Considered and rejected: Rejected standard SHAP and permutation explainability methods directly on raw temporal series because permuting individual time steps breaks temporal autocorrelations and trajectory dependencies.
Survival Analysis Using Machine Learning for Longitudinal, Multimodal, and High-dimensional Data for Applications in Cardiology · JScholarship
Tried and failed
permutation test on neural network input gradients applied to feature significance testing with correlated predictors. Reason: permuting individual features broke joint predictor distributions, causing inflated Type-I error rates under feature correlation
Tried and failed
unstructured permutation testing of temporal labels applied to temporally imbalanced sequence data. Reason: temporal density imbalances across sample periods generated artificially inflated statistical significance
Studies in Bacterial Genome Dynamics · Harvard
Tried and failed
Empirical copula permutation tests for correlation homogeneity applied to Testing heterogeneous correlation structure. Reason: Failed to control Type I error under high overall sample correlation.
Testing for associations in a heterogeneous population · Texas Tech
Non-smooth derivatives and boundary constraints cause asymptotic bootstrap inconsistency
Bootstrapping failed to provide consistent estimation when parameters lay on simplex boundaries, near zero, or under directional Hadamard derivatives in Wasserstein distances. Resampling also failed in overparametrized linear regressions, post-LASSO estimators, and constrained network settings, producing bounded parameter coverage collapse and severe variance underestimation.
Tried and failed
pair and residual bootstrap applied to high-dimensional linear regression. Reason: consistently underestimates true parameter variance in overparametrized regimes even with optimal regularization
Theoretical characterization of uncertainty in high-dimensional machine learning · EPFL
Tried and failed
standard bootstrap confidence intervals applied to constrained maximum likelihood estimation. Reason: finite samples frequently contained zero false negatives causing boundary issues
Essays on Data Science: Computational Measurement for Learning and Teaching · Harvard
Tried and failed
m-out-of-n subsampling and bootstrap inference applied to optimal transport inference. Reason: exhibited poor size control, over-rejecting significance levels and requiring impractically large sample sizes to converge
Essays on the US Mortgage Market · Harvard
Tried and failed
naive empirical bootstrap applied to sliced Wasserstein distance hypothesis testing. Reason: inconsistent when the set of optimal potentials is non-unique under the alternative hypothesis
Statistical Inference for Regularized Optimal Transport · Cornell
Tried and failed
standard nonparametric bootstrap applied to empirical max-sliced Wasserstein distance. Reason: non-linear directional Hadamard derivatives cause bootstrap inconsistency
Statistical Inference for Regularized Optimal Transport · Cornell
Tried and failed
random rotation bootstrap on spherical point patterns applied to cross summary statistics null envelopes. Reason: replicates collapse to zero variance at maximum spherical distance pi due to geometric constraints
Tried and failed
standard nonparametric bootstrap applied to max-sliced Wasserstein and Gromov-Wasserstein distances. Reason: nonlinearity of the directional derivative makes the bootstrap statistically inconsistent
Computational and Statistical Properties of Optimal Transport-Based Distances · Cornell
Considered and rejected
Considered and rejected: Rejected using standard multinomial bootstrap for inference on simplex-boundary parameters due to asymptotic failure.
Measurement Error in Microbiome Sequencing Experiments: Statistical and Scientific Considerations · ResearchWorks
Considered and rejected
Considered and rejected: Rejected re-fitting the propensity score using post-LASSO within each bootstrap resample because bootstrap is inconsistent for LASSO estimators.
Propensity Score Methods For Causal Subgroup Analysis · DukeSpace
Considered and rejected
Considered and rejected: Rejected bias correction in the parametric bootstrap for selection coefficient confidence intervals due to optimization instability when s is near zero.
Statistical Inference Using Identity-by-Descent Segments: Perspectives on Recent Positive Selection · ResearchWorks
Tried and failed
Wald and bootstrap interval estimators applied to constrained network parameters under RDS. Reason: Failed to achieve valid nominal coverage, producing under 50% coverage for bounded partnership probability parameters
Bootstrap confidence intervals exhibit severe undercoverage and distorted widths
Percentile, Studentized, and Wald bootstrap intervals frequently failed to achieve their nominal coverage rates under heavy censoring, nearest neighbor matching, neural network predictions, and large sensitivity parameters. In several settings bootstrap intervals produced biased bounds, over-corrected bias, or generated excessively wide confidence intervals compared to asymptotic baselines.
Lost to a baseline
Calibration-bootstrap method showed worse coverage accuracy (over-coverage on lower bounds, under-coverage on upper bounds) compared to predictive-distribution methods under heavy censoring
Prediction interval methods for reliability data · Iowa State
Lost to a baseline
Under large variance differences (σ2^2 = 2), Student's t-bootstrap CI width was substantially wider (AW 248.07) than asymptotic CI (AW 98.92) at sample sizes n1=n2=10.
Statistical methods and analysis with real-life data · oURspace
Considered and rejected
Considered and rejected: Bootstrapping with replacement for confidence intervals, rejected because it produces confidence intervals systematically biased below upper bounds
Boundary line methodology for yield gap analysis of farm systems · University of Nottingham Repository
Considered and rejected
Considered and rejected: Rejected tree bootstrap variance estimation for RDS-II as overly conservative with studentized intervals failing to remain within [0, 1].
Tried and failed
nonparametric bootstrap rank confidence intervals applied to sparse multi-environment trial rankings. Outcome: did not generalise. Reason: Uniform interaction distributions degraded coverage under severe subsampling compared to bell-shaped distributions
Accounting for rank uncertainty in decision making for plant breeding · Iowa State
Tried and failed
wild bootstrap for nearest neighbor matching estimators applied to confidence interval estimation in high overlap. Reason: produced severe undercoverage and underestimated confidence interval width in nearest-neighbor matching
Tried and failed
bootstrap neural network ensemble for uncertainty quantification applied to trajectory prediction and control prediction intervals. Outcome: did not generalise. Reason: severe under-coverage resulting in extreme coverage width-based penalties
Scientific Machine Learning for Engineering Systems: Optimal Control, Prognostics, and Uncertainty Quantification · Virginia Tech
Tried and failed
bootstrap bias correction applied to small area MSE estimation. Reason: standard bootstrap adjustment produced an over-correction of the estimation bias
Small area estimation and graphical model for complex surveys · Iowa State
Lost to a baseline
FQE (Bootstrap) produced tighter confidence intervals than BONDIC on Inverted-Pendulum, though it failed the nominal coverage rate across trials
Efficient and safe off-policy evaluation : from point estimation to interval estimation · UT Austin
Lost to a baseline
Percentile bootstrap CI coverage for A1,2 in high overlap (82.8%) was lower than chi-square-based intervals (90.3%).
New methods in home-range overlap and clustering · Iowa State
Lost to a baseline
Percentile and Wald bootstrap CI coverage drops below nominal 95% (to ~88-90%) under large sensitivity parameter values (α = 0.6) compared to smaller α values.
Permutation and bootstrap testing suffer from reduced statistical power and distorted test statistics
Permutation and bootstrap hypothesis tests were repeatedly outperformed by parametric tests and simpler baselines that achieved higher statistical power. Furthermore, permutation testing with small samples or shrinkage exhibited skewed null distributions, elevated false discovery rates on weak effects, or inverted feature importance rankings.
Lost to a baseline
The multivariate JA test with wild bootstrap exhibits lower empirical power under heteroskedasticity than under Gaussian or t(5) error distributions.
Three Essays in Econometrics · Carleton University Institutional Repository
Lost to a baseline
Pivotal bootstrap testing has lower statistical power compared to oracle Gaussian testing under identical sample sizes.
TOPICS IN MODERN REGRESSION MODELING · Cornell
Lost to a baseline
Linear regression selective MSE: bootstrap methods (Doubt-Var/Int) were beaten by standard PlugIn-REG and cross-fitting baselines.
Topics in Selective Prediction · IRIS - SNS - prod
Tried and failed
exact permutation testing for differential region detection applied to small sample genomic data. Outcome: no signal. Reason: extreme group swap permutations skewed the null false discovery rate distribution in small samples
Statistical Methods for Mobile Health and Genomics Data · Harvard
Tried and failed
variance stabilization with permutation testing applied to differential ChIP-seq signal detection. Reason: elevated false discovery rate when detecting weak induced effect sizes
Quantitative analysis of ChIP-seq signals and transcriptomes · Georgia Tech
Tried and failed
Ledoit-Wolf shrinkage in max-T FWER framework applied to resampling-based multiple hypothesis testing. Outcome: worse than baseline. Reason: Proved overly conservative and underestimated p-values compared to standard permutation testing.
Statistical Models for Alternative Splicing with Applications to Heterogeneous Disease · Penn
Tried and failed
fixed summary statistics for matrix permutation estimation applied to permuted monotone matrix recovery. Outcome: worse than baseline. Reason: failed to adapt to heterogeneous signal strengths across samples
Problems In High-Dimensional Statistics And Applications In Genomics, Metabolomics And Microbiomics · Penn
Lost to a baseline
Sim-5 (classification permutation test comparison): Classification permutation test using random forests (package cpt) achieved substantially lower power to detect mean shifts across δ than GFKS-1, GFKS-2, adjusted KS, and adjusted t-tests.
Contributions To Multivariate Matching In Observational Studies · Penn
Lost to a baseline
In small sample and unbalanced strata settings (n10 < 10), the Louis estimator with fixed weight (q0=0.7) exhibited slightly higher power than the bootstrapped optimal-weight estimator.
Tried and failed
backward permutation feature importance applied to ensemble postprocessing neural networks. Reason: produced inverted, unphysical predictor rankings conflicting with forward permutation tests
Leveraging Artificial Intelligence and Numerical Weather Prediction to Build Custom Smart Home Software · Texas Tech
Lost to a baseline
Under Gaussian distributions, all rank-based distance covariance tests (Wilcoxon scores and normal scores) and permutation tests were uniformly outperformed by the parametric Gaussian Likelihood Ratio Test.
Distribution-Free Consistent Tests of Independence via Marginal and Multivariate Ranks · ResearchWorks
Resampling causes numerical divergence and singular design matrices
Generating resamples frequently produced empty fixed-effect cells, singular design matrices, zero-variance cluster covariates, or extreme boundary values that destabilized downstream estimation. These resamples resulted in iterative optimization divergence in mixed models, blown-up standard errors in regression calibration, and negative variance estimates in double bootstrap routines.
Tried and failed
bootstrap variance estimation after regression calibration applied to survival models with measurement error. Outcome: unstable. Reason: a significant fraction of bootstrap resamples produced extreme coefficient estimates, blowing up standard error
Tried and failed
Parametric bootstrap simulation of dispersion parameters applied to Laplace-approximated mixed-effects models. Outcome: did not converge. Reason: High variance in simulated parameters caused numerical divergence during iterative optimization
Methods for spatial hierarchical generalized linear mixed models · Iowa State
Tried and failed
parametric double bootstrap mean squared error estimation applied to small area estimation under non-linear models. Outcome: unstable. Reason: produced negative variance estimates preventing confidence interval construction
Tried and failed
stratified cluster bootstrap sampling applied to complex survey-weighted longitudinal data. Outcome: unstable. Reason: produced unstable, biased, and asymmetric sampling distributions
Considered and rejected
Considered and rejected: Residual bootstrap for variance estimation in bounded Tobit longitudinal data was rejected in favor of random weighting due to repeated sampling of extreme boundary values.
Statistical Models for Identification of Treatment-sensitive Subgroups Based on Longitudinal Outcomes in Clinical Trials · Queens University Institutional Repository
Considered and rejected
Considered and rejected: Decided against non-parametric bootstrapping for calculating mediation indirect effects due to convergence failures and inflated error estimates
Considered and rejected
Considered and rejected: Rejected traditional resampling bootstrap because cells with many fixed effects became empty, choosing Dirichlet random weighting instead
Essays in Labor Economics · JScholarship
Considered and rejected
Considered and rejected: Decided against pair bootstrapping because resampling clusters causes covariates to lack variance/drop levels and results in varying sample sizes with small cluster counts.
Cluster wild bootstrapping to handle dependent effect sizes in meta-analysis with small number of studies · UT Austin
Tried and failed
adaptive ensemble kernels in bootstrap testing applied to nonlinear hypothesis testing. Outcome: unstable. Reason: adaptive estimation caused unstable test statistics and underestimated the alternative kernel compared to fixed alternatives
Considered and rejected
Considered and rejected: Nonparametric pairs (x-y) bootstrap was rejected because it frequently yields singular design matrices during resampling in simulations.
Small sample sizes and sparse records undermine resampling validity
Resampling on restricted sample sizes failed due to poor support of continuous exposures, unobserved rare species, and sample bias that missed true distribution peaks. Short or sparse record lengths prevented effective noise reduction, yielded uninformative inclusion probabilities, and led researchers to reject bootstrapping in favor of analytical standard errors.
Tried and failed
weighted likelihood bootstrap applied to normal mean posterior estimation. Outcome: did not generalise. Reason: produced non-normal posterior samples with variance independent of the observation noise scale
Objective Bootstrap Posterior distributions · UT Austin
Tried and failed
Discrete bootstrap resampling applied to continuous treatment effect estimation. Outcome: data insufficient. Reason: Poor support of continuous exposure in small samples
Causal Inference Methods To Evaluate Health Policies With Spillover · Penn
Tried and failed
bootstrap inclusion probabilities for sparse regression selection applied to governing equation discovery from time series. Outcome: no signal. Reason: inclusion probabilities failed to clearly separate true model terms from false positives
Parameter estimation and inference for nonlinear dynamical systems · Cornell
Considered and rejected
Considered and rejected: Bootstrap confidence interval methods (Moving Block / Semiparametric); rejected due to insufficient contiguous record lengths in sparse SWOT Cal/Val data.
Spatiotemporal tidal prediction and analysis through physics-informed machine learning · Oxford
Tried and failed
bootstrap resampling on small sample sizes applied to side-channel leakage detection. Outcome: data insufficient. Reason: sample bias in restricted trace subsets failed to capture true statistical distribution peaks
Towards Comprehensive Side-channel Resistant Embedded Systems · Virginia Tech
Tried and failed
bootstrap sub-sampling to regularize covariance response matrix applied to linear response operator estimation. Outcome: no signal. Reason: sub-sampling and masking cross-correlation terms failed to reduce sampling noise sufficiently in short records
Considered and rejected
Considered and rejected: Non-parametric biodiversity estimators (Chao1, Chao2, Jackknife, bootstrapping) rejected for biogeographic datasets where surveyed plot area is small (<2%) and rare species are largely unobserved.
Feature selection by Information Imbalance optimization: Clinics, molecular modeling and ecology · IRIS - SISSA - prod
Considered and rejected
Considered and rejected: Rejected using bootstrapping for confidence interval calculation on manually measured bed thicknesses due to small sample sizes (N <= 8), using standard error instead.
The regional sedimentary record of Arabia Terra, Mars. · JScholarship
Left open by the authors
Problems the authors named and did not get to.
Left open
Extend iterated 2SLS outlier test statistics to non-normal reference distributions using Student-t and empirical bootstrap distributions. Blocker: None
Left open
Derive the asymptotic distribution theory for sign-preserved spatially varying coefficient estimators to eliminate the need for bootstrap inference. Blocker: Requires advanced mathematical statistics theory rather than software engineering.
Generalized, quantile and constrained nonparametric regression for spatial data · Iowa State
Left open
Compute bootstrap confidence intervals for space technology venture capital return metrics instead of assuming normal distribution. Blocker: Requires access to proprietary CB Insights and BryceTech venture capital datasets used in the thesis.
Houston, We Have Profits: Analyzing Venture Capital Investment in the Space Technology Industry · Harvard
Left open
Evaluate selection model extensions incorporating robust variance estimation and bootstrapping for dependent effect sizes in meta-analysis via simulation. Blocker: None
Left open
Prove that the exponentially tilted Bayesian bootstrap posterior is the noninformative limit of the Kitamura-Otsu formulation as base measure approaches zero. Blocker: None
Weighting and moment conditions in Bayesian inference · Cambridge
Left open
Implement a full bootstrap routine over the multi-step structural estimation procedure to compute standard errors accounting for earlier-stage estimation error. Blocker: Requires access to the thesis's specific empirical bidding datasets and estimation codebase.
Left open
Extend the bootstrap aggregation (bagging) acquisition function procedure to fully Bayesian MCMC inference for Deep Gaussian Processes. Blocker: None
Physics-informed Machine Learning for Digital Twins of Metal Additive Manufacturing · Virginia Tech
Left open
Perform conformal bootstrap using mixed phi, s, t correlators at higher derivative truncation Lambda to determine the O(3) model fate in 2<d<3. Blocker: None
Left open
Implement Chung 2019 transposition permutations to reduce the permutation test cost for energy distance goodness-of-fit tests on manifolds. Blocker: None
Goodness of fit tests for object data · Texas Tech
Left open
Develop fast analytical approximations for small-sample (N < 20) distance correlation hypothesis testing to replace permutation tests. Blocker: None
Multiscale Statistical Hypothesis Testing for k-Sample Graph Inference · JScholarship
Checking a claim in this area?
We can run the same search on any method or claim. If nothing turns up, we will say so, and that proves nothing on its own.