DS925: Causal Inference for Management Research
Fall 2026
| Instructor | Raviv Murciano-Goroff |
| [email protected] | |
| Time | Tuesday, 3:00–5:45 pm |
| Location | HAR 658 |
| Course pages | ds925.com |
| Office hours | send me an email |
Format
This course is meant to be practical. Therefore, each week:
- Before class — Review the the conceptual reading. Please come to class with questions.
- In class (2.5 hours) — three blocks:
- Concept review. Review of the econometric methods; clear up questions.
- Paper discussion. We’ll consider two applied papers that use the method.
- Research-question lab. Together, we’ll write code and examine data using the method from class.
- After class — work on the problem set (5 across the semester).
Objectives
By the end of the course you should be able to:
- Read an empirical research paper and understand its identifying assumptions.
- Choose between identification strategies given a research question and a data environment.
- Implement RCT analysis, matching, IV, RDD, panel/DiD, synthetic DiD, non-linear models, and ML-augmented estimators in R or Stata.
- Critically evaluate and defend the assumptions behind a chosen design.
- Write the identification section of your own paper.
Prerequisites
This course is designed for students who have previously taken a graduate-level course in econometrics. I strive, however, to make the material accessible to any students with basic knowledge of probability and statistics. The course does not require you to prove theorems, however, in lecture and in the readings there will be technical material. While I will try to focus on the intuition behind the methods and assumptions, the equations can be illuminating, and both the readings and the lectures will spend some time describing them.
Problem sets can be done in Stata, R, or the open-source package of your choice. Most in-class examples will use R largely because most of you already have experience with Stata. Having some familiarity with these tools is great, however, I am also happy to help you get up to speed if you have limited experience with them prior to the course.
Grading
- Participation (10%) — classroom discussion benefits everyone. Ask all your questions.
- Problem sets (65%) — five empirical exercises. You are welcome to discuss the assignments with your fellow students, but the submission and documentation must be done individually.
- Final project (25%) — replicate/extend a published paper, or submit an original research design.
Readings
Each class will have several readings. Some of these are conceptual readings to build understanding of the foundation of a method or assumption. Some of these readings are applied microeconomic papers that apply and grapple with the methods and assumptions. The conceptual readings will help you understand the method theoretically. We’ll discuss the applied papers in class, so you should be prepared to comment on them. You should spend more or less time with each one based on your research interests.
Textbooks (reference)
- Angrist & Pischke, Mostly Harmless Econometrics (MHE) — BU library
- Cunningham, Causal Inference: The Mixtape — mixtape.scunning.com
Auditing and attendance
Auditors who contribute to discussion are welcome; only enrolled students are graded. Attendance is essential.
Diversity and Inclusion Statement
In developing this course, I have aimed to be thoughtful about how identity and culture impact the course content. I invite all students to engage with these sensitive conversations in search of common ground vis-à-vis diversity, racism, equity, and inclusion. If there are topics or conversations that you feel would benefit from incorporation of social context, a differing perspective, or the support of our Questrom’s Office of Diversity & Inclusion, please inform me and I will explore resources and opportunities for us to engage a wide variety of perspectives in our classroom. If you have concerns or ideas about diversity and inclusion at Questrom you can also reach the Questrom Diversity & Inclusion office at [email protected].
Acknowledgements
This course is inspired by classes, slides, and blog posts from Tim Simcoe, Scott Cunningham, Paul Goldsmith-Pinkham, Peter Hull, Chris Conlon, Jon Roth, Frank Wolak, Susan Athey, and Stefan Wagner.
Course outline
Class 1 — Potential Outcomes, Identification, and Endogeneity
Learning objectives
- Use the potential-outcomes framework and be able to articulate the fundamental missing-data problem.
- Distinguish parameter, estimand, and estimate; define identification.
- Explain which variation identifies a coefficient.
- Recognize structural vs. reduced-form econometrics.
- Diagnose types of endogeneity (selection, simultaneity, OVB) and whether controls can fix them.
Conceptual readings
- MHE Ch. 1–2.
- Imbens & Wooldridge (2009) “Recent Developments in the Econometrics of Program Evaluation” pp. 1–12.
- Manski (1995), Identification Problems in the Social Sciences, Harvard University Press: Introduction; Ch. 1, “Extrapolation”; Ch. 2, “The Selection Problem”; Ch. 6, “Simultaneity”; and Ch. 7, “The Reflection Problem.”
- Goldberger (1991), A Course in Econometrics, Ch. 23, “Multicollinearity.”
- Heckman (1981), “Heterogeneity and State Dependence”, through Section 3.1.
Optional conceptual readings
- Lewbel (2019), “The Identification Zoo: Meanings of Identification in Econometrics”, Journal of Economic Literature 57(4): 835–903.
- Gelman (2011), “Causality and Statistical Learning”, American Journal of Sociology 117(3): 955–966.
- Manski (1993), “Identification Problems in the Social Sciences”, Sociological Methodology 23: 1–56.
Empirical paper
- Bertrand & Schoar (2003) “Managing with Style: The Effect of Managers on Firm Policies”, Quarterly Journal of Economics 118(4): 1169–1208.
Optional empirical paper
- LaLonde (1986) “Evaluating the Econometric Evaluations of Training Programs with Experimental Data,” AER.
Slides
- 1.1 — Potential outcomes and the missing-data problem
- 1.2 — Estimands, estimators, estimates, and what “identification” means
- 1.3 — Structural vs. reduced-form econometrics
- 1.4 — Endogeneity I: selection, OVB, measurement error
- 1.5 — Signing OVB bias — and when you can’t
- 1.6 — Endogeneity II: simultaneity and the reflection problem
- 1.7 — Regression refresher, and the Class 1 synthesis
In-class research exercise
A management consultancy claims its “AI-augmented strategy workshops” raise client revenue by 12%. They have before/after revenue data for clients who bought the workshop and for clients who didn’t. What is the estimand they are implicitly claiming to estimate? What is the estimand they are actually estimating? Describe the gap precisely using potential-outcomes notation.
Class 2 — The Experimental Ideal (RCTs)
Learning objectives
- Show how randomization solves selection and the missing-data problem.
- Compare SDO and regression estimation of treatment effects.
- Distinguish intention-to-treat from treatment-on-the-treated.
- Probe heterogeneity; know when regression adjustment helps and when it doesn’t.
- Conduct inference for an RCT and design one with power calculations.
- Design factorial, multi-arm, and dose experiments, and know the power cost of each contrast.
- Pre-specify a primary outcome and control multiplicity across outcomes and subgroups.
- Protect a design ex ante through blocking, and recognize when interference invalidates the estimand.
- State what an RCT cannot deliver, and what limits its external validity.
Conceptual readings
- MHE Ch. 2 or Mixtape Ch. 4.
- Athey & Imbens (2017) “The Econometrics of Randomized Experiments,” sections 1–2, 4–5, 10–11.
Optional conceptual readings
- Lin (2013), “Agnostic Notes on Regression Adjustments to Experimental Data: Reexamining Freedman’s Critique”, Annals of Applied Statistics 7(1): 295–318.
- World Bank, “You Ran a Field Experiment. Should You Then Run a Regression?”.
- Lin’s World Bank summaries of regression adjustment: Part 1 and Part 2.
- Doyle & Feeney, J-PAL, “Quick Guide to Power Calculations”.
- World Bank, “Power Calculations 101: Dealing with Incomplete Take-up”.
- Goldsmith-Pinkham, Hull & Kolesár (2022), “Contamination Bias in Linear Regressions”, NBER Working Paper 30108.
- Duflo, Glennerster & Kremer (2006), “Using Randomization in Development Economics Research: A Toolkit”, NBER Technical Working Paper 333.
Empirical papers
- Bloom, Eifert, Mahajan, McKenzie & Roberts (2013) “Does Management Matter? Evidence from India,” QJE.
- Dai, Kim & Luca (2023) “Frontiers: Which Firms Gain from Digital Advertising? Evidence from a Field Experiment,” Marketing Science.
Slides
- 2.1 — How randomization fixes selection
- 2.2 — SDO vs. regression adjustment
- 2.3 — ITT and the LATE under non-compliance
- 2.4 — Treatment-effect heterogeneity, interactions, and endogenous moderators
- 2.5 — Inference: design-based vs. model-based SEs
- 2.6 — Power calculations and MDE
- 2.7 — Balance, blocking, and re-randomization
- 2.8 — SUTVA, spillovers, and interference
- 2.9 — Adaptive experiments and bandits
- 2.10 — Attrition, differential attrition, and Lee bounds
- 2.11 — Limitations, external validity, and the RCT checklist
- 2.12 — Factorial, multi-arm, and dose designs
- 2.13 — Primary vs. secondary outcomes and outcome multiplicity
In-class research exercise
You are advising a SaaS firm that wants to test whether including an AI-generated meeting summary in their product raises customer retention. They can randomize at the user level or at the workspace level. Design the experiment: which randomization unit, which primary outcome, what’s the MDE you can detect with their 50,000 active workspaces over 8 weeks, and what is the biggest threat to the exclusion-style “randomization-was-actually-random” assumption?
Class 3 — Selection on Observables and Matching
Learning objectives
- Explain how matching addresses the missing-data problem and why weights are needed.
- Distinguish conditional independence and common-support assumptions.
- Compute ATE, ATT, and CATEs; recognize the curse of dimensionality.
- Implement coarsened-exact, nearest-neighbor, and bias-corrected matching.
- Recognize that regression with controls is a matching estimator, state the weights it applies, and explain how it differs from matching on common support and functional form.
- Explain what a propensity score is, why it works, and how to use it (weighting, matching, regression adjustment on the score). Note: “control function” is reserved for the endogeneity-correction sense used in Classes 4-5.
- Recognize when doubly-robust estimation buys you something.
- Quantify how large unobserved confounding would have to be to overturn a selection-on-observables result, and report it (Rosenbaum bounds, Oster, Cinelli-Hazlett).
- Explain how audit and correspondence studies construct matched profiles, identify the causal effect of a randomized signal, and distinguish that narrow effect from broader disparities or downstream outcomes.
Conceptual readings
- Mixtape Ch. 5.
- Imbens & Wooldridge (2009) Ch. 6.
Optional conceptual readings
- Słoczyński (2022), “Interpreting OLS Estimands When Treatment Effects Are Heterogeneous: Smaller Groups Get Larger Weights”, Review of Economics and Statistics 104(3): 501–509.
- Imbens (2015), “Matching Methods in Practice: Three Examples”, Journal of Human Resources 50(2): 373–419, Sections 1–3.
- Oster (2019), “Unobservable Selection and Coefficient Stability: Theory and Evidence”, Journal of Business & Economic Statistics 37(2): 187–204.
Applied readings
- Localization of knowledge
- Jaffe, Trajtenberg & Henderson (1993) “Geographic Localization of Knowledge Spillovers as Evidenced by Patent Citations,” QJE.
- Thompson & Fox-Kean (2005) “Patent Citations and the Geography of Knowledge Spillovers: A Reassessment,” AER.
- Henderson, Jaffe & Trajtenberg (2005), Comment, AER 95(1): 461–464.
- Thompson & Fox-Kean (2005), Reply, AER 95(1): 465–466.
- Sytch, Wohlgezogen & Zajac (2018), “Collaborative by Design? How Matrix Organizations See/Do Alliances”, Organization Science 29(6): 1130–1148.
Slides
- 3.1 — Conditional independence and the selection-on-observables story
- 3.2 — Common support and the curse of dimensionality
- 3.3 — Exact and coarsened matching
- 3.4 — Nearest-neighbor matching with bias correction
- 3.5 — Regression with controls: what OLS is actually doing
- 3.6 — Propensity scores: theorem and three uses
- 3.7 — Doubly-robust estimators
- 3.8 — Sensitivity analysis: how wrong can CIA be?
- 3.9 — Audit and correspondence studies: when the researcher writes the match
In-class research exercise
A retail bank claims that its loyalty-program members spend 23% more annually than non-members. You have demographic and transaction data on all customers. Sketch a matching strategy that would isolate the effect of membership from selection into membership. Be specific: which variables go in, which are off-limits as “bad controls,” and how would you check common support? What does the conditional-independence assumption require here, and what is the most plausible violation?
Class 4 — Instrumental Variables
Learning objectives
- Articulate why IV works; identify “good” variation in a setting.
- State the four IV assumptions (relevance, independence, exclusion, monotonicity).
- Compute first-stage, reduced-form, and Wald estimators; implement 2SLS.
- Interpret LATE in terms of compliers.
- Diagnose weak instruments, monotonicity violations, and exclusion concerns.
Conceptual readings
- MHE Ch. 4.1–4.3.
- Mixtape Ch. 7.
- Angrist & Pischke (2009) “Mostly Harmless Econometrics” Ch. 4 sec on LATE.
- Murray (2006), “Avoiding Invalid Instruments and Coping with Weak Instruments”, Journal of Economic Perspectives 20(4): 111–132.
Optional conceptual readings
- Berry & Haile (2021), “Foundations of Demand Estimation”, Handbook of Industrial Organization, Vol. 4, Ch. 1: 1–62.
- Ackerberg, Benkard, Berry & Pakes (2007), “Econometric Tools for Analyzing Market Outcomes”, Handbook of Econometrics, Vol. 6, Ch. 63: 4171–4276 — the demand-estimation portion.
- Rasmusen, short summary of demand-estimation challenges and methods.
- Conlon, PyBLP implementation guide.
Applied readings
- Bernstein (2015), “Does Going Public Affect Innovation?” Journal of Finance 70(4): 1365–1403.
- Bennedsen, Nielsen, Pérez-González & Wolfenzon (2007), “Inside the Family Firm: The Role of Families in Succession Decisions and Performance”, QJE 122(2): 647–691.
- Babina, Fedyk, He & Hodson (2024), “Artificial Intelligence, Firm Growth, and Product Innovation,” Journal of Financial Economics 151: 103745. Read Section 5.3 and Table 5; consult Appendix A for instrument-construction details. The remainder is optional for this class.
Slides
- 4.1 — Why IV works: good variation, bad variation
- 4.2 — Wald estimator and 2SLS mechanics
- 4.3 — The four assumptions and why each matters
- 4.4 — LATE, Compliers, and Marginal Treatment Effects
- 4.5 — Weak instruments: diagnostics and pitfalls
- 4.6 — Multiple instruments and overidentification
In-class research exercise
A VC firm believes its hands-on board involvement (taking a board seat, monthly check-ins) causes portfolio companies to grow faster. Propose an IV-based identification strategy: candidate instrument, defense of each of the four assumptions, what would constitute a falsification test, and the LATE interpretation — which compliers are you identifying off?
Class 5 — Judge & Examiner Designs
Learning objectives
- Recognize the canonical judge/examiner setup as a leave-one-out leniency IV.
- Confirm monotonicity in this setting and know how it can fail.
- Probe the exclusion restriction when the “treatment” is bundled with other examiner behaviors.
- Implement leave-one-out leniency instruments correctly.
- Read recent critiques of judge IV designs.
Conceptual readings
- MHE Ch. 4.4–4.6.
- Mixtape §7.8 (review the assigned section from the original edition), together with the judge-fixed-effects discussion in Ch. 7.
- Frandsen, Lefgren & Leslie (2023) “Judging Judge Fixed Effects,” AER.
Applied readings
- Bernstein, Colonnelli & Iverson (2019) “Asset Allocation in Bankruptcy,” Journal of Finance.
- Collinson, Humphries, Mader, Reed, Tannenbaum & van Dijk (2024) “Eviction and Poverty in American Cities,” Quarterly Journal of Economics.
- Righi & Simcoe (2019) “Patent Examiner Specialization,” Research Policy.
Slides
- 5.1 — The judge/examiner setup at a glance
- 5.2 — Constructing leave-one-out leniency measures
- 5.3 — Monotonicity in judge designs and how it fails
- 5.4 — The exclusion restriction when bundles are large
- 5.5 — When examiner assignment isn’t actually random
- 5.6 — Recent critiques, leniency-as-LATE, and MTE reinforcement
In-class research exercise
Does access to credit help small firms grow? Using applications for a working-capital loan, construct a leave-one-out loan-officer leniency instrument, estimate the first stage, reduced form, and IV effect on firm outcomes. Interpret the switching margins and examine heterogeneity. Separate in-class scenarios introduce observable and hidden sorting, advice bundled with lending, and crossing officer preferences. We’ll examine what each scenario does to independence, exclusion, or monotonicity and what evidence or redesign would be needed.
Class 6 — Regression Discontinuity and Bunching
Learning objectives
- See how RDD relates to previous methods (matching with continuous controls; IV with interpolation).
- State RDD assumptions: locally randomized cutoff, no manipulation, accurate interpolation.
- Estimate sharp and fuzzy RDDs with local linear regression; choose a bandwidth.
- Read McCrary density tests and covariate-balance checks.
- Recognize kink designs and bunching estimators.
- Assess effect heterogeneity using predetermined moderators, full running-variable interactions, and direct subgroup tests.
Conceptual readings
- Imbens & Lemieux (2008) “Regression Discontinuity Designs: A Guide to Practice”.
- MHE Ch. 6.
Applied readings
- Luca (2016) “Reviews, Reputation, and Revenue: The Case of Yelp.com”.
- Bahar, Choudhury, Kim & Koo (2023) “Innovation on Wings: Nonstop Flights and Firm Innovation in the Global Context,” Management Science.
Slides
- 6.1 — Sharp RDD: the local randomization story
- 6.2 — Estimation: local linear, bandwidth choice, optimal RDD
- 6.3 — McCrary density tests and covariate balance
- 6.4 — Fuzzy RDD as RDD-meets-IV
- 6.5 — Kink designs
- 6.6 — Bunching estimators and what they identify
- 6.7 — Heterogeneity and moderators in RDD
In-class research exercise
A federal grant program awards funding to small firms scoring above a threshold on a proposal-evaluation score. The agency publishes scores. Design an RDD study of the effect of receiving a grant on three-year firm survival. Specify: bandwidth choice procedure, two covariate-balance checks you would run, what you would do if the McCrary density test rejects, and the most plausible reason the threshold is not locally exogenous in the setting.
Class 7 — Panel Data and Fixed Effects
Learning objectives
- Define fixed effects and the panel structures (balanced, strongly balanced).
- Read fixed-effects output: within vs. between variation, rho.
- Recognize what two-way fixed effects control for and what residual variation identifies each coefficient.
- Diagnose attrition, identify singleton groups, and distinguish short-panel information limits from violations of strict exogeneity.
- Combine matching with fixed effects.
- Distinguish event-study estimation from event-study graphs.
Conceptual readings
- MHE Ch. 5.
- Griliches & Mairesse (1995), “Production Functions: The Search for Identification”, NBER Working Paper 5067.
Optional conceptual readings
- Ackerberg, Benkard, Berry & Pakes (2007), “Econometric Tools for Analyzing Market Outcomes”, Handbook of Econometrics, Vol. 6, Ch. 63: 4171–4276 — the production-function-estimation portion.
Applied readings
- Mas & Moretti (2009), “Peers at Work”, American Economic Review 99(1): 112–145.
- Ewens, Peters & Wang (2024) “Measuring Intangible Capital with Market Prices,” Management Science.
Slides
- 7.1 — Fixed effects: within vs. between transformation
- 7.2 — Reading FE output: rho, within R², and what FE actually absorbs
- 7.3 — Two-way fixed effects
- 7.4 — Fixed effects at different nesting levels
- 7.5 — Identifying variation, singletons, and short panels
- 7.6 — Attrition and panel imbalance
- 7.7 — Event studies: estimation vs. plotting
In-class research exercise
You have 12 years of plant-level data on energy use and output for ~3,000 US manufacturing plants. You want to estimate the productivity effect of installing a smart-grid sensor system. What does plant FE buy you? What does year FE buy you? Name a confounder that survives both. Sketch how you would combine matching with two-way FE to address that confounder, and explain which assumption you are now buying.
Class 8 — Shift-Share Instrumental Variables
Learning objectives
- Decompose a shift-share instrument into exogenous shifts and exogenous shares.
- State the Borusyak-Hull-Jaravel exogenous-shocks condition and the Goldsmith-Pinkham-Sorkin-Swift exogenous-shares interpretation.
- Choose the right inference (BHJ vs. ADH standard errors).
- Diagnose shift-share validity in your own setting.
Conceptual readings
- Borusyak, Hull & Jaravel (2022) “Quasi-Experimental Shift-Share Research Designs,” Restud.
- Goldsmith-Pinkham, Sorkin & Swift (2020) “Bartik Instruments: What, When, Why, and How,” AER.
Applied readings
- Autor, Dorn & Hanson (2013) “The China Syndrome: Local Labor Market Effects of Import Competition in the United States,” AER.
- Beaumont, Hebert & Lyonnet (2025) “Build or Buy? Human Capital and Corporate Diversification,” Review of Financial Studies.
- Gagliardi & Sorenson (2026) “Entrepreneurship and Gentrification,” Organization Science.
Slides
- 8.1 — What is a shift-share instrument? Bartik in one picture
- 8.2 — Exogenous shifts (BHJ) vs. exogenous shares (GPSS)
- 8.3 — Constructing Bartik-style instruments in practice
- 8.4 — Inference under shift-share: Adao-Kolesar-Morales
- 8.5 — Identifying the LATE under shift-share
- 8.6 — Common pitfalls and falsification tests
In-class research exercise
A management researcher wants to estimate the effect of local exposure to generative-AI deployment on small-business formation. Propose a shift-share instrument: what are the “shares” (and where do they come from)? What are the “shifts”? Defend either (a) the exogeneity of the shares OR (b) the exogeneity of the shifts. What falsification test would convince a skeptic?
Class 9 — Difference-in-Differences
Learning objectives
- Express DiD as a 2×2 contrast and as a regression with interaction.
- State the parallel-trends assumption; understand its symmetric formulation.
- Test parallel trends with leads and event-study plots.
- Handle covariate-adjusted DiD when control composition differs.
- Recognize that DiD identifies the ATT on the treated, not the ATE.
Conceptual readings
- Mixtape Ch. 9.
- Roth, Sant’Anna, Bilinski & Poe (2023) “What’s Trending in DiD?” sections on canonical 2×2.
- Baker, Callaway, Cunningham, Goodman-Bacon & Sant’Anna (2026) “Difference-in-Differences Designs: A Practitioner’s Guide”, Journal of Economic Literature 64(2): 498–557.
Applied reading
- Card & Krueger (1994) “Minimum Wages and Employment: A Case Study of the Fast-Food Industry in New Jersey and Pennsylvania”, AER.
- Jin & Leslie (2003) “The Effect of Information on Product Quality: Evidence from Restaurant Hygiene Grade Cards”, QJE.
- Greenstone, Oyer & Vissing-Jørgensen (2006) “Mandated Disclosure, Stock Returns, and the 1964 Securities Acts Amendments”, QJE.
Optional applied reading
- Dranove, Kessler, McClellan & Satterthwaite (2003) “Is More Information Better? The Effects of `Report Cards’ on Health Care Providers”, JPE.
Slides
- 9.1 — DiD as 2×2 and as a regression
- 9.2 — Parallel trends: what the assumption really says
- 9.3 — Pre-trend tests and event-study plots
- 9.4 — Triple-difference designs
- 9.5 — Covariate-adjusted DiD and the Sant’Anna-Zhao estimator
- 9.6 — Inference: clustering at the right level
- 9.7 — Matching + DiD: when treatment is a choice and trends are the threat
In-class research exercise
A US state mandates pay-transparency disclosure for all job postings starting January 2024. Design a DiD study of the effect on average posted wages within the state. Specify the control group, the running window, and two falsification tests (one on pre-trends, one on a placebo outcome). What is the strongest threat to parallel trends here? If you find pre-trends do not pass, what do you do?
Class 10 — Staggered & Conditional Difference-in-Differences
Learning objectives
- Show why TWFE breaks under staggered treatment timing (negative weights, forbidden comparisons).
- Implement Callaway-Sant’Anna, Sun-Abraham, de Chaisemartin-D’Haultfœuille, and Borusyak-Jaravel-Spiess estimators.
- Read and produce honest event-study plots in staggered settings.
- Estimate DiD when treatment status is conditional on observables (conditional parallel trends).
Conceptual readings
- Roth, Sant’Anna, Bilinski & Poe (2023) “What’s Trending in DiD?” full paper.
- de Chaisemartin & D’Haultfœuille (2020) “Two-Way Fixed Effects Estimators with Heterogeneous Treatment Effects,” AER.
Papers for discussion
- Azoulay, Graff Zivin & Wang (2010) “Superstar Extinction”, QJE.
- La Forgia & Bodner (2024) “Getting Down to Business: Chain Ownership and Fertility Clinic Performance”, Management Science.
Slides
- 10.1 — Why TWFE goes wrong under staggered adoption
- 10.2 — Callaway & Sant’Anna at a glance
- 10.3 — Sun & Abraham; de Chaisemartin & D’Haultfœuille
- 10.4 — Borusyak, Jaravel & Spiess imputation estimator
- 10.5 — Conditional PT and Continuous-Treatment DiD
- 10.6 — Event-study plots done right
In-class research exercise
Different US states have legalized recreational cannabis at different times between 2012 and 2024. A researcher claims to estimate the effect of legalization on the survival rates of locally headquartered small businesses using two-way FE. Why might this be wrong even with perfectly parallel pre-trends in each state pair? Propose an alternative estimator, justify the choice, and sketch the event-study plot you would expect to see if the effect is positive and constant.
Class 11 — Synthetic DiD
Learning objectives
- See synthetic control as weighted matching on pre-treatment outcomes.
- Implement vanilla synthetic-control and synthetic-DiD estimators.
- Diagnose pre-treatment fit and infer using placebo permutations.
- Choose between synthetic DiD and standard DiD for a given setting.
Conceptual readings
- Abadie (2021) “Using Synthetic Controls: Feasibility, Data Requirements, and Methodological Aspects,” JEL.
- Arkhangelsky, Athey, Hirshberg, Imbens & Wager (2021) “Synthetic Difference-in-Differences,” AER.
Papers for discussion
- Abadie, Diamond & Hainmueller (2010) “Synthetic Control Methods for Comparative Case Studies: Estimating the Effect of California’s Tobacco Control Program”, JASA.
- Lambrecht, Tucker & Zhang (2024) “TV Advertising and Online Sales: A Case Study of Intertemporal Substitution Effects for an Online Travel Platform”, Journal of Marketing Research.
Slides
- 11.1 — The synthetic-control intuition: a weighted control unit
- 11.2 — Abadie-Diamond-Hainmueller estimation
- 11.3 — Synthetic DiD: combining weights with parallel-trends
- 11.4 — Placebo tests, leave-one-out, and inference
- 11.5 — When synthetic methods fail: poor pre-fit, extrapolation
- 11.6 — Synthetic DiD vs. DiD: when to use which
In-class research exercise
One country adopts a sweeping data-portability regulation (analogous to GDPR Article 20) in 2025. You want to estimate the effect on the entry rate of fintech startups in that country. There are 18 plausible donor countries. Design a synthetic-DiD study: which pre-treatment outcomes and covariates would you use as matching variables, how would you assess fit, how would you do inference, and what would convince you the synthetic control is not a credible counterfactual?
Class 12 — Non-Linear Models
Learning objectives
- Diagnose when OLS is fine for limited-dependent-variable outcomes and when it isn’t.
- Compute and interpret marginal effects in logit/probit; recognize the Ai-Norton interaction-term trap.
- Use Poisson and negative-binomial regressions for count outcomes; defend log(1+y) vs. PPML choices.
- Explain the incidental-parameters problem in short nonlinear panels; distinguish dummy-variable logit/probit, conditional logit, and the static FE Poisson exception.
- Distinguish nonlinear slope estimates from probability-scale effects, and recognize the assumptions behind bias corrections.
- Use and interpret hazard (survival/duration) models, including Cox proportional-hazards and parametric models; distinguish hazard rates from survival probabilities and assess censoring, the proportional-hazards assumption, and competing risks.
Conceptual readings
- Ai & Norton (2003) “Interaction Terms in Logit and Probit Models,” Economics Letters.
- Angrist (2001), “Estimation of Limited Dependent Variable Models with Dummy Endogenous Regressors: Simple Strategies for Empirical Practice”, Journal of Business & Economic Statistics 19(1): 2–28.
- Mullahy (1997), “Instrumental-Variable Estimation of Count Data Models: Applications to Models of Cigarette Smoking Behavior”, Review of Economics and Statistics 79(4): 586–593.
- Santos Silva & Tenreyro (2006), “The Log of Gravity”, Review of Economics and Statistics 88(4): 641–658.
- Mullahy (1998), “Much Ado About Two: Reconsidering Retransformation and the Two-Part Model in Health Econometrics”, Journal of Health Economics 17(3): 247–281.
- Cohn, Liu & Wardlaw (2022) “Count (and Count-like) Data in Finance,” JFE.
- Freedman (2008), “Survival Analysis: A Primer”, The American Statistician 62(2): 110–119.
- Jenkins (1995), “Easy Estimation Methods for Discrete-Time Duration Models”, Oxford Bulletin of Economics and Statistics 57(1): 129–136.
Applied readings
- Waguespack & Fleming (2009) “Scanning the Commons? Evidence on the Benefits to Startups Participating in Open Standards Development”, Management Science 55(2): 210–223.
Slides
- 12.1 — Why LPM is usually fine — and when it isn’t
- 12.2 — Logit/probit marginal effects and the interaction trap
- 12.3 — Count data: Poisson, NB, PPML, log(1+y)
- 12.4 — Nonlinear fixed effects and the incidental-parameters problem
- 12.5 — Duration models: Cox, parametric, competing risks
- 12.6 — Identification meets non-linearity: combining with DiD/IV
In-class research exercise
A researcher wants to estimate the effect of a firm’s first patent grant on the count of follow-on patents in the next five years. The outcome is a count with many zeros; the firm has multiple potential first-patents per year. Walk through the modeling choices: (a) OLS on log(1+y), (b) Poisson with firm FE, (c) PPML, (d) hazard model on time-to-second-patent. Which would you use and why? What does each estimand identify?
Class 13 — Machine Learning + Inference / Standard Errors
Learning objectives
- Distinguish prediction from causal inference; know where ML buys causal-inference researchers something.
- Use double/debiased ML for partially linear models.
- Recognize where causal forests and meta-learners help with heterogeneous effects.
- Cluster standard errors at the level of treatment assignment (and know why).
- Use Bayesian-bootstrap / wild-cluster bootstrap when cluster counts are small.
Conceptual readings
- Mullainathan & Spiess (2017) “Machine Learning: An Applied Econometric Approach,” JEP.
- Athey & Imbens (2019) “Machine Learning Methods that Economists Should Know About,” Annu. Rev. Econ..
- Abadie, Athey, Imbens & Wooldridge (2023) “When Should You Adjust Standard Errors for Clustering?” QJE.
Optional conceptual readings
- Wager & Athey (2018), “Estimation and Inference of Heterogeneous Treatment Effects Using Random Forests”, Journal of the American Statistical Association 113(523): 1228–1242.
Applied readings
- Kleinberg, Lakkaraju, Leskovec, Ludwig & Mullainathan (2018) “Human Decisions and Machine Predictions,” QJE.
- Athey & Palikot (2026) “Effective and Scalable Programs to Facilitate Labor Market Transitions for Women in Technology”, NBER Working Paper 34750.
Slides
- 13.1 — Prediction vs. causal inference: where ML helps
- 13.2 — Double/debiased ML for partially linear models
- 13.3 — Causal forests and the GATES
- 13.4 — Clustering standard errors at the right level
- 13.5 — Wild-cluster bootstrap for few clusters
- 13.6 — Pitfalls: sample-splitting, hyperparameter tuning, p-hacking
In-class research exercise
Pick the dataset from one of your previous problem sets in this course. Identify one question in that paper where heterogeneous effects matter. Design an analysis using a causal forest (or another heterogeneous-effects method) to estimate the GATES. State explicitly: (a) what nuisance functions are being learned, (b) what your splitting variables are and why, (c) what the right level of clustering is for your standard errors, and (d) why the wild-cluster bootstrap would or would not be needed.
Last updated 2026-08-31.