Chapter Four · failure evidence
What Difference-in-Differences Design got wrong, from 64 dissertations
Difference-in-differences studies frequently encounter methodological breakdowns arising from pre-treatment parallel trend violations and two-way fixed effects estimation biases under staggered adoption. Researchers also report vulnerabilities to unmeasured time-varying confounders, spatial spillover interferences, and sensitivity to outcome transformations and aggregation levels. These records come from PhD theses at 19 institutions, 2021 to 2026. Each links to its thesis. They were extracted by language models reading the full text, so treat each as a lead to read, not a verdict.
Pre-treatment parallel trends violations and failed falsification tests invalidate causal identification
Empirical evaluations frequently failed because treated and control cohorts exhibited divergent pre-intervention trajectories or failed placebo tests before the intervention occurred. These pre-existing differences prevented researchers from establishing valid counterfactual comparisons and forced the rejection of standard difference-in-differences designs.
Tried and failed
difference-in-differences policy evaluation applied to subgroup employment after policy intervention. Reason: parallel trends assumption was violated in pre-treatment periods, preventing causal identification
The Impact of Medicaid Policy Change on People with Disabilities · Harvard
Tried and failed
doubly robust difference-in-differences estimator applied to regional income inequality aggregate indices. Outcome: no signal. Reason: violation of parallel-trend assumptions and no statistically significant effect on aggregate inequality metrics
Essays on income inequality in the United States · UT Austin
Tried and failed
stacked difference-in-differences with synthetic controls applied to healthcare labor policy impact evaluation. Reason: pre-period parallel trends assumption failed statistical significance test in low baseline states
The Impact of Medicaid Policy Change on People with Disabilities · Harvard
Tried and failed
difference-in-differences across income eligibility threshold applied to earned income tax credit effects. Reason: severe violations of parallel pre-trends between treatment and control groups near threshold
Tried and failed
standard difference-in-differences estimation applied to state income tax reform impacts. Reason: Severe pre-reform downward trends in income and population violated the parallel trends assumption.
ESSAYS IN APPLIED MICROECONOMICS · Cornell
Tried and failed
difference-in-differences with two-way fixed effects applied to evaluating policy effects on crime rates. Reason: violation of parallel trends assumption demonstrated by failed pre-trend placebo tests
Three essays on economics of crime and immigration · Texas Tech
Tried and failed
difference-in-differences estimation applied to heterogeneous firm productivity and investment. Reason: pre-treatment trends differed significantly across firm productivity, capital, and investment levels, violating parallel trends
Tried and failed
difference-in-differences estimation applied to retail experiments. Outcome: no signal. Reason: violation of the parallel trends assumption in over 90% of experiments
Tried and failed
Staggered difference-in-differences design applied to Policy impact on educational outcomes. Reason: Significant pre-existing trends violated the parallel trends assumption across treated and control units.
ESSAYS ON THE HUMAN CAPITAL AND OCCUPATIONAL CHOICES OF ADOLESCENTS · Cornell
Tried and failed
trimming sample outliers in difference-in-differences analysis applied to panel data with heterogeneous treatment. Reason: removing extreme units violated the parallel trends assumption required for valid causal inference
Tried and failed
event study difference-in-differences applied to maternal labor outcomes after childbirth. Reason: pre-existing differential trends violated the parallel trends assumption required for causal identification
Three Essays in Development Economics · Harvard
Tried and failed
difference-in-differences event study applied to aggregate mortality rates after policy expansion. Reason: pre-treatment divergence violated the parallel trends assumption required for causal inference
Essays in health economics · UT Austin
Considered and rejected
Considered and rejected: Rejected standard Difference-in-Differences (DID) due to violation of the parallel trends assumption between state cohorts.
Essays in applied economics · Texas Tech
Considered and rejected
Considered and rejected: Difference-in-Differences (DiD) was rejected for primary causal claims regarding parental labor market participation due to failure of the parallel pre-treatment trends assumption for maternal outcomes.
Socioeconomic implications of adverse birth outcomes · Leibniz Universität Hannover Repository
Considered and rejected
Considered and rejected: Rejected standard Difference-in-Differences (DID) design for evaluating universal FAFSA policies because parallel trends do not hold across states.
Essays in the Economics of Education · Harvard
Considered and rejected
Considered and rejected: Traditional Difference-in-Differences methodology rejected due to violation of parallel trends from staggered rollout and COVID-19 moratorium shocks.
Essays in applied microeconomics · OpenBU
Considered and rejected
Considered and rejected: Rejected standard difference-in-differences using simple 30-city average or 25-year-olds as controls due to violation of parallel trends assumption
Essays on Labor and Development Economics · ResearchWorks
Considered and rejected
Considered and rejected: Rejected difference-in-differences design because non-random DP signings make violation of the parallel trends assumption highly likely
Considered and rejected
Considered and rejected: Naive difference-in-differences without propensity score adjustment was rejected due to severe parallel trends violations from the Great Recession
Essays in Labor and Public Economics · Cornell
Considered and rejected
Considered and rejected: Rejected difference-in-differences comparing corequisite and non-corequisite colleges across the TSI distribution due to failure of the parallel trends assumption outside the eligibility window.
Essays on inference and education economics · UT Austin
Tried and failed
event-study difference-in-differences with placebo validation applied to quarterly sector-level innovation metrics. Reason: placebo tests showed contamination, preventing definitive causal attribution of the post-treatment shift
Tried and failed
event-study difference-in-differences with high-dimensional fixed effects applied to quarterly policy impact on innovation counts. Outcome: no signal. Reason: violation of parallel trends assumptions and failure on placebo falsification tests
Tried and failed
event-study difference-in-differences with fixed effects applied to sector-level quarterly innovation patent counts. Outcome: no signal. Reason: failed parallel trends validation and placebo robustness checks across most treatment groups
Tried and failed
difference-in-differences policy evaluation applied to health IT adoption and hospital outcomes. Outcome: unstable. Reason: initially significant effect estimates vanished under modest violations of parallel trends assumptions
From Theory to Practice: Improving Causal Conclusions from Healthcare Data · MIT
Tried and failed
stacked event study difference-in-differences applied to evaluating policy impacts on mortality rates. Reason: pre-intervention trends between treated and control groups failed the parallel trends assumption
Pressing Challenges in U.S. Health Policy: Polarization, Gun Violence, And Discrimination · Harvard
Staggered rollout and heterogeneous timing produce bias and negative weights in two-way fixed effects
Canonical two-way fixed effects models generated biased estimates because dynamic treatment effects led to forbidden comparisons with already-treated units. Researchers rejected or modified these staggered designs after finding pervasive negative weighting issues and instability across alternative estimators.
Tried and failed
difference-in-differences regression applied to heterogeneous staggered rollout panel data. Reason: Divergent pre-intervention trends and heterogeneous adoption timelines violated standard parallel trends assumptions.
Considered and rejected
Considered and rejected: Standard two-way fixed effects difference-in-differences (DiD) was rejected because ban and non-ban states violated the parallel trends assumption prior to Dobbs and staggered adoption causes bias.
Tried and failed
staggered difference-in-differences with robust estimators applied to evaluating corporate investment policy impacts. Outcome: no signal. Reason: treatment effects disappeared after correcting for staggered adoption bias and outlier winsorization
Essays on the Impact and Effectiveness of Share Repurchase Regulations · Harvard
Considered and rejected
Considered and rejected: Rejected canonical Difference-in-Differences models with period and group fixed effects due to biased estimates arising from 'forbidden' comparisons of already-treated units under heterogeneous treatment effects.
Beyond Trade Losses: Sanctions and the Resilience of Productive Economies · Texas Tech
Considered and rejected
Considered and rejected: Rejected standard two-way fixed effects (TWFE) difference-in-differences due to treatment effect heterogeneity across adoption cohorts and negative weighting issues.
Essays on applied labour economics: personality traits, artificial intelligence and formalisation policies · University of Nottingham Repository
Tried and failed
two-way fixed effects difference-in-differences estimation applied to staggered policy adoption evaluation. Reason: pervasive negative weights on early-treated cohorts and failure of treatment homogeneity assumptions
Considered and rejected
Considered and rejected: Rejected standard Two-Way Fixed Effects (TWFE) with staggered treatment due to time-varying treatment effects leading to biased comparisons with already-treated units, adopting stacked difference-in-differences instead
Essays on infrastructure and urban development in developing countries · OpenBU
Considered and rejected
Considered and rejected: Decided against using traditional Difference-in-Differences estimators because they average exposure-outcome associations over time and can be biased when associations vary as policy generosities change.
Considered and rejected
Considered and rejected: Rejected standard two-way fixed-effects (TWFE) estimators in favor of stacked difference-in-differences to avoid bias from heterogeneous/staggered treatment timing.
Impacts of State Abortion Restrictions on Mental Health and Healthcare · JScholarship
Considered and rejected
Considered and rejected: Rejected staggered difference-in-differences as the primary empirical design due to econometric concerns regarding staggered DiD estimates, relegating it to a robustness check.
Do Private Tax Disclosures Affect the Quality of Public Financial Reporting? · Scholars' Bank
Considered and rejected
Considered and rejected: Rejected staggered difference-in-differences designs due to econometric bias from heterogeneous treatment timing, opting for canonical 2x2 DD models
Considered and rejected
Considered and rejected: Rejected standard two-way fixed-effects (TWFE) difference-in-differences models due to bias caused by dynamic/staggered treatment timing and negative weights comparing early-treated to late-treated units.
Three Essays On Economics of Marriage Law · Texas Tech
Considered and rejected
Considered and rejected: Rejected standard Two-Way Fixed Effects (TWFE) staggered difference-in-differences due to contamination and bias from heterogeneous cohort treatment effects, adopting the Wooldridge (2021) two-way Mundlak regression instead.
Essays on Foreign Direct Investment and International Trade · Research Repository UCD
Omitted time-varying confounders and selection shocks distort estimated treatment effects
Unmeasured confounding shocks, concurrent events, and unobserved selection dynamics generated counterintuitive or spurious treatment point estimates. Models without proper covariate adjustment or controls for time-varying environmental and cohort confounders failed to isolate the true causal intervention.
Considered and rejected
Considered and rejected: Rejected using unweighted difference-in-differences as the sole model without testing multiple group propensity score weighting to address baseline covariate imbalance.
Indirect Effects of Financial Incentives on Physician Behavior in Perinatal Care · JScholarship
Tried and failed
staggered difference-in-differences with robust estimators applied to longitudinal policy evaluation of student performance. Reason: produced spurious long-term negative effects likely driven by unobserved time-varying confounders rather than treatment
Art-tendance: The Effect of Creative Learning on Student Outcomes · UT Austin
Tried and failed
difference-in-differences regression without fixed effects applied to policy intervention impact estimation. Reason: omitted variable bias produced counterintuitive negative treatment coefficient
The Effect of #MeToo on Gender-Related Shareholder Activism · Penn
Tried and failed
difference-in-differences without covariate adjustment applied to continuous spatial policy spillover estimation. Reason: unadjusted confounding produced counterintuitive positive treatment effect estimates at low dose ranges
Causal Inference Methods To Evaluate Health Policies With Spillover · Penn
Tried and failed
staggered difference-in-differences using localized shock applied to environmental quality impact on student performance. Outcome: no signal. Reason: Confounding negative effects from psychological trauma and commuting disruptions masked any environmental benefits.
Three Essays in Applied Microeconomics · Georgia Tech
Tried and failed
difference-in-differences on panel data applied to international trade policy interventions. Reason: unmeasured confounding and anticipatory effects yielded counterintuitive negative causal estimates
Identification and Estimation of Policy-relevant Causal Effects from Observational Data · EPFL
Considered and rejected
Considered and rejected: Rejected relying solely on OLS and Difference-in-Differences with inconsequential units approach due to non-random intermediate small cities and time-varying omitted environmental confounders (e.g., pollution, noise).
Essays on Health and Transportation Economics · DSpace at SUNY Buffalo
Tried and failed
matched difference-in-differences estimation applied to school-level education intervention effects. Reason: unobserved group-by-cohort selection bias caused spurious negative point estimates
Building Communities of Upward Mobility · Harvard
Considered and rejected
Considered and rejected: Rejected relying on standard matched difference-in-differences estimators because treatment was not strictly an absorbing state across cohorts and schools suffered negative selection shocks.
Building Communities of Upward Mobility · Harvard
Estimates are sensitive to variable scaling, functional form, and geographic aggregation
Treatment effects frequently disappeared or degraded when researchers altered expenditure scaling, normalized by baseline growth, or aggregated outcomes across broader geographic units. Authors also rejected standard designs for discrete or bounded outcomes due to baseline mean reversion and out-of-range counterfactual predictions.
Tried and failed
difference-in-differences regression on policy intervention applied to gender disparity in job compensation. Reason: gap reduction stemmed from baseline decline among the advantaged group rather than absolute gains for the disadvantaged
Considered and rejected
Considered and rejected: Rejected standard Difference-in-Differences (DiD) / parallel trends for discrete and bounded outcomes because it confounds baseline mean reversion with treatment effects and can produce counterfactual probabilities outside the [0, 1] range
PROGRAM EVALUATION OF TREATMENT EFFECT HETEROGENEITY: THEORY AND APPLICATIONS · Penn
Tried and failed
Difference-in-differences with alternative variable scaling applied to Corporate expenditure and investment ratios. Outcome: no signal. Reason: Treatment effects lost statistical significance when scaling expenditures by revenue instead of lagged total assets
Tried and failed
difference-in-differences on sector-specific innovation counts applied to evaluating industrial policy impact. Outcome: no signal. Reason: treatment effect disappears after normalizing by total baseline patent growth
Tried and failed
staggered difference-in-differences with robust estimators applied to contraception access policy on education outcomes. Outcome: no signal. Reason: Estimated policy effects on bachelor's completion and major choice were not robust across estimators.
ESSAYS ON THE HUMAN CAPITAL AND OCCUPATIONAL CHOICES OF ADOLESCENTS · Cornell
Considered and rejected
Considered and rejected: Rejected Synthetic Difference-in-Differences (SDiD) because it requires a balanced panel (dropping valid data), is computationally inefficient over long pre-periods, and tends to match on pre-treatment noise.
Essays on Digital Content Strategies: Creation, Diffusion, and Monetization · Harvard
Considered and rejected
Considered and rejected: Rejected Difference-in-Differences (DiD) because it constrains predictor effects to be constant over time rather than accounting for weighted variable changes.
Tried and failed
pooled difference-in-differences estimator applied to downstream property price changes. Outcome: no signal. Reason: pooling all sites obscured heterogeneous treatment effects across locations
Essays on environmental economics · UT Austin
Tried and failed
difference-in-differences regression discontinuity design applied to spillover effects across broader geographic regions. Outcome: no signal. Reason: statistical significance and consistency degraded when aggregating outcomes across the wider subregion
Think Twice: Deterring Transnational Kidnapping through Rescue · Harvard
Universal policy exposure and spatial spillovers violate baseline control and non-interference requirements
Difference-in-differences designs failed when nationwide policy rollout or simultaneous eligibility shocks left no clean, untreated comparison groups. In other empirical settings, spatial spillovers and shifting trade status violated the stable unit treatment value assumption.
Considered and rejected
Considered and rejected: Rejected standard Difference-in-Differences for evaluating the 2017 German corporate tax reform because all German corporate startups were simultaneously affected, adopting Synthetic Control instead.
Essays on the Finance of Startups · Publikationssystem UB Tuebingen
Considered and rejected
Considered and rejected: Rejected difference-in-differences/regression discontinuity due to lack of a clean exogenous shock affecting only MAIs without impacting all patenting firms.
The value of analytics innovations · Iowa State University Digital Repository
Considered and rejected
Considered and rejected: Rejected standard Difference-in-Differences (DiD) in favor of pooled OLS / SAR because the SUTVA assumption failed due to spatial spillovers and shifting net-exporter status.
Spatial Dimensions of Economic Modeling: Interdisciplinary Approaches to Labor, Trade, and Networks · Georgia Tech
Considered and rejected
Considered and rejected: Standard difference-in-differences design between colonias with and without service was rejected due to small sample sizes and gradual rollout across all counties.
Left open by the authors
Problems the authors named and did not get to.
Left open
Estimate the causal effect of continuous per-dollar minimum wage increases on health outcomes using continuous-treatment difference-in-differences methods. Blocker: None
Impact of Social Policies on Health · Harvard
Left open
Modify difference-in-differences minimum wage empirical designs to incorporate cross-state worker migration leakages. Blocker: None
How the Price System Works: Evidence from Supply Chains and Price Controls · Harvard
Left open
Estimate difference-in-differences regressions interacting minimum wage increases with PNTR tariff exposure shocks to measure contemporaneous buffering effects on local crime rates. Blocker: None
Three Essays on the Impacts of Trade Liberalization · Georgia Tech
Left open
Estimate compulsory schooling law impacts on child labor, marriage age, bride prices, sibling allocations, and birth spacing using difference-in-differences. Blocker: None
Essays on China’s Economic Development · Cornell
Left open
Test parallel trends assumptions in ACA Medicaid expansion difference-in-differences models using linked individual microdata rather than aggregate mortality data. Blocker: Requires access to restricted administrative individual microdata (e.g., linked Census/vital statistics records).
Essays in health economics · UT Austin
Left open
Estimate the causal effect of carbon pricing on provincial industrial greenhouse gas emissions and emissions intensity using staggered Difference-in-Differences. Blocker: None
The causal effect of carbon pricing on industrial energy consumption in Canada · MSpace - University of Manitoba
Left open
Estimate difference-in-differences regressions comparing import-reliant and non-reliant sectors to distinguish parallel exchange rate transmission mechanisms. Blocker: None
Exchange Rate Pass-Through from Parallel Foreign Exchange Markets · Harvard
Left open
Perform difference-in-differences analysis testing parallel trends and health trajectory selection on subsequent IFLS survey waves. Blocker: None
Left open
Evaluate longer-term earnings outcomes beyond age 20 using matched difference-in-differences on longitudinal administrative wage and education records. Blocker: Requires access to restricted administrative longitudinal earnings data linked to student records.
Essays on health economics and public policy · UT Austin
Left open
Evaluate whether Science Based Targets participation affects downside beta, drawdowns, and return volatility using matched pair difference-in-differences regression on S&P 500 firms. Blocker: None
Innovation Through a Net Zero Economy and the Impact to an Investors Bottom Line · Harvard
Checking a claim in this area?
We can run the same search on any method or claim. If nothing turns up, we will say so, and that proves nothing on its own.