https://arxiv.org/api/eU8iwusacW5ox25MQeaRGVEML982026-07-22T00:33:54Z563016515http://arxiv.org/abs/2507.00307v3Orthogonalized Synthetic Controls2026-06-25T00:06:18ZWhen conducting inference for the average treatment effect on the treated with a Synthetic Control Estimator, the vector of control weights is a nuisance parameter that is often constrained, high-dimensional, and may be only partially identified even when the average treatment effect on the treated is point-identified. All three of these features of a nuisance parameter can lead to failure of asymptotic normality for the estimate of the parameter of interest when using standard methods. I provide a new method that yields asymptotic normality for an estimate of average treatment effects, even when all three complications are present. This is accomplished by first estimating the control weights and any other nuisance parameters using a regularization penalty to achieve identification, and then estimating average treatment effects using moment conditions that are orthogonalized with respect to the nuisance parameters. Additionally, I extend results from the fixed-smoothing literature to provide tests that control size without requiring consistent standard errors. I present high-level sufficient conditions applicable to the traditional Synthetic Control Estimator as well as other weighting-based panel data methods, and verify them in an example involving instrumental variables.2025-06-30T22:48:00ZJoseph Fryhttp://arxiv.org/abs/2606.26432v1Embedding Foundation Model Predictions in Discrete-Choice Models with Structural Guarantees2026-06-24T22:44:29ZTabular foundation models achieve strong accuracy on choice prediction tasks, but their predictions often violate the economic logic those tasks require: raising a price can increase predicted demand, implied willingness-to-pay estimates are frequently negative or implausible, and unavailable alternatives receive nonzero probability. We propose a two-stage adapter that takes a foundation model's predicted choice probabilities as a precomputed feature and embeds them inside a multinomial logit's utility. In Stage 1, we fit the multinomial logit's structural coefficients by maximum likelihood with sign constraints; in Stage 2, we freeze those coefficients and fit a small neural correction operating on the foundation model's predictions. We prove that this composition exactly preserves the multinomial logit's marginal rate of substitution, so analytically computable value-of-time becomes a mathematical guarantee rather than an empirical accident. Across three datasets and two foundation models, the adapter gains 6.4 percentage points (pp) of test accuracy on average over the multinomial logit and up to 12.8 pp, maintains 100% cost monotonicity, and produces values of time within the published transportation-economics range on the transportation datasets. Performance degrades gracefully under foundation-model context restriction, retaining at least 6 pp of accuracy gain even at 10% of the original foundation-model context.2026-06-24T22:44:29ZExtends arXiv:2605.26559 (ICML 2026 FMSD Workshop)Yingshuo WangXian SunYanhang LiZhichao FanZexin Zhuanghttp://arxiv.org/abs/2606.15031v2Partial Identification from LLM Prompts2026-06-24T12:18:02ZLarge language models are increasingly used as binary classifiers when the true label is latent. We study partial identification of the prevalence $θ= P(X^* = 1)$ from panels of LLM reports whose errors may be arbitrarily dependent given the truth. The design of replication determines the observable, and hence the identifying content: repeated prompts to one model yield a count, several named models a response vector, and both a response matrix. Cast as a two-component finite mixture, the problem makes the identification failure transparent: absent restrictions that separate the latent components, the prevalence $θ$ is completely unidentified, and weak stochastic-ordering restrictions (first-order dominance, monotone likelihood ratio, mean ordering) leave the identified set at $[0,1]$. Identifying power comes instead from externally calibrated scores and events, which discipline the mixture in the spirit of the misclassification and corrupted-data literature. We characterize the resulting bounds, establishing validity and sharpness, and give an exact account of the identifying information in the full score distribution beyond its mean. When named models are asked repeated versions of the same question, what identifies $θ$ is not the number of positive answers but which models agree across prompts -- a feature a vote count discards. An extension derives implied bounds on regression coefficients when $X^*$ is a regressor of interest that is not directly observed.2026-06-13T00:09:24Z31 PAGES TOTAL. NO FIGURESXiaohong ChenAshesh RambachanElie Tamerhttp://arxiv.org/abs/2606.25688v1Choosing What to Calibrate and What to Estimate in Structural Models2026-06-24T10:55:06ZStructural models often fix (calibrate) some parameters and estimate the rest, but this calibration-estimation partition is usually chosen by convention. This paper treats that choice as an econometric partition-selection problem. For each admissible partition, we construct a scalar sensitivity statistic measuring the local response of a target object -- such as a policy effect, welfare measure, impulse response, or treatment effect -- to perturbations of the calibrated parameters. The selected partition minimizes this statistic and therefore minimizes worst-case local bias from calibration errors. We first illustrate the decision problem in two canonical examples. We then apply it to the New Keynesian model of Nakamura and Steinsson (2018), where the partition choice has large implications for credibility: some partitions remain reliable under sizeable miscalibrations, whereas others generate large bias from small calibration errors. The procedure requires only local derivatives, avoids repeated re-estimation, and applies to a broad class of structural models.2026-06-24T10:55:06ZJoan Alegre Cantonhttp://arxiv.org/abs/2606.25292v1Time-Varying Model Averaging of Multi-layer Network Vector Autoregressions2026-06-24T01:55:13ZIn this paper, we introduce a flexible time-varying multi-layer network vector autoregression (VAR) model framework for large-scale time series, allowing agents in dynamic systems to interact through multiple channels and incorporating multiple adjacency matrices to capture network spillover effects. We propose a penalized model averaging method to determine a time-varying optimal combination of multi-layer network VAR candidate models whose number may be divergent. Under some regularity conditions, the asymptotic properties such as asymptotic optimality and convergence rates of the proposed time-varying weight estimation are derived in the contexts of both the in-sample fitting and out-of-sample prediction. In addition, we extend the conformal prediction method to construct prediction bands for locally stationary time series. Monte-Carlo simulation studies and an empirical application to forecast CPI inflation by combining multiple network information are given to illustrate reliable finite-sample estimation and predictive performance of the developed methodology.2026-06-24T01:55:13ZDegui LiYuying SunBoyao Wuhttp://arxiv.org/abs/2606.22599v2Networked risk perception and behavioral bubbles: the case of a pandemic2026-06-23T20:34:01ZRisk perception is typically modeled as an individual cognitive readout of objective hazard, yet during crises what people judge as risky is shaped by what their peers do. Using weekly mobility data from 313 Massachusetts municipalities over the first year of the COVID-19 pandemic and a pre-pandemic inter-town mobility network that fixes interaction structure before the shock, we estimate two-way fixed-effects panel regressions that separate local case response, inter-town behavioral spillover along the mobility network, and within-town inertia; the pre-shock network and a lagged peer signal address the standard reflection and endogenous-group concerns. Three findings emerge. First, inter-town behavioral spillovers are substantial and localize almost entirely within mobility-defined communities, with effectively no propagation across community boundaries, the empirical referent of behavioral bubbles. Second, the within-community spillover carries behavioral content beyond peer-town case information: when network-exposure-to-cases and network-exposure-to-behavior are raced, the behavioral channel survives and the case-exposure channel goes null. Third, a joint mobility-by-demographic decomposition shows the spillover requires both routine connection and demographic similarity. It concentrates where towns are connected and similar, and vanishes between similar towns that are not connected, ruling out a shared-conditions confound and pointing to an observational and normative channel rather than a purely informational one. These results recast risk perception as a networked phenomenon and identify mobility-defined communities, rather than administrative units, as the operative scale of behavioral response. The pattern should generalize wherever exposure is uncertain, evolving, and socially negotiated, including climate adaptation and financial contagion.2026-06-21T17:16:53ZSepehr IlamiMargherita ComolaSilvia PrinaBabak Heydarihttp://arxiv.org/abs/2606.24867v1Bounds for Standard Errors in Combined Data2026-06-23T17:47:54ZWe propose methods for constructing lower bounds on the standard errors of parameters estimated from moment conditions obtained across different samples. Sharp explicit bounds are derived by exploiting geometric inequalities when no information about correlations across samples is available. Furthermore, we develop computationally tractable sharp bounds for more general settings with no or partial correlation information, which can be obtained by solving a simple semidefinite program. Finally, we illustrate the practical usefulness of our method through three empirical cases: two macroeconomics examples involving menu cost and Heterogeneous Agent New-Keynesian models; and a two sample instrumental variable microeconomic study.2026-06-23T17:47:54Z46 pages, 2 figuresJooyoung ChaYuya SasakiNelson Matthew P. Tanhttp://arxiv.org/abs/2606.24850v1Heterogeneous Peer Effects with Endogenous Network Formation2026-06-23T17:28:01ZThis paper introduces a new econometric framework for modeling social interactions with heterogeneous peer responses, addressing endogenous link formation. Our Selection-corrected Heterogeneous Spatial Autoregressive (SCHSAR) approach jointly models link formation and outcome determination. We incorporate a finite mixture structure to capture heterogeneity in peer effects and account for unobserved individual-specific factors driving both network formation and outcome equations, addressing network endogeneity for credible estimation of heterogeneous spillover effects. We propose a fully Bayesian data augmentation approach for estimation and inference, overcoming challenges posed to standard likelihood-based methods. A simulation study validates our approach. Our empirical application to an innovation network among U.S. firms reveals significant positive, yet heterogeneous, peer effects on corporate R&D investments, after accounting for endogenous network formation. The findings highlight varying firm behaviors in response to exogenous R&D policy shocks and and quantify firm-level direct and spillover effects, offering valuable insights for evidence-based and targeted policy design.2026-06-23T17:28:01ZUploaded for IAAE 2026 conferenceDuong TrinhSantiago Montoya-Blandónhttp://arxiv.org/abs/2606.24785v1Group-Level Treatment Effect Heterogeneity in Difference-in-Differences: A Balanced Approach2026-06-23T16:44:50ZUnderstanding how treatment effects vary across groups is central to policy evaluation. In Difference-in-Differences designs, heterogeneity is often studied using subgroup or triple-difference analyses, which can suffer from conservative inference, reliance on parametric interaction structures, and sensitivity to differences in covariate distributions across groups. We propose the Balanced Group Average Treatment Effect on the Treated (BGATT), a new estimand that isolates heterogeneity in treatment responses from differences in covariate composition and is identified under standard conditional parallel-trends assumptions. BGATT provides a transparent target for comparing group-specific treatment effects. We derive an influence-function representation and develop estimators that are $\sqrt{n}$-consistent and asymptotically normal under flexible machine-learning estimation of high-dimensional nuisance components, enabling valid inference on both group-specific effects and differences across groups. Simulation evidence shows favorable finite-sample performance.2026-06-23T16:44:50ZNora BearthNadja van 't HoffTorben S. D. Johansenhttp://arxiv.org/abs/2606.17165v3Statistical Foundations of LLM-based A/B Testing: A Surrogacy Framework for Human Causal Inference2026-06-23T14:53:08ZOrganizations and researchers show increasing interest in using large language models (LLMs) in place of human participants in A/B tests, in the hope of experimenting faster and at lower cost. We study when a treatment effect estimated on LLM outcomes can recover the effect for the human population of interest. Distributional equivalence between LLM and human outcomes would make any standard estimator valid but is unrealistic. We therefore develop a statistical framework that adapts surrogate endpoint theory to LLMs, showing that calibrating LLM outcomes to human outcomes identifies the average treatment effect under surrogacy and comparability conditions that are jointly weaker than distributional equivalence. We present a falsification test for surrogacy and a bound on the worst-case bias from limited overlap between the LLM and human samples. We further show that the stochasticity inherent to LLMs can weaken surrogacy for identification while also introducing bias and variance during estimation, but that using an average over multiple LLM draws per unit as the surrogate mitigates these issues. Simulations validate the results, and an empirical application to the Upworthy Research Archive dataset shows that raw LLM outputs recover only 39% of the human treatment effect while nonparametric calibration closes the gap. A central takeaway is that A/B testing on LLM responses is correct only by assumption, whereas A/B testing on humans is correct by design, and that the required assumptions are hardest to justify precisely where LLMs promise the greatest benefit. We discuss the choice of LLM, prompting, and temperature as design variables, the compounded challenge posed by long-term outcomes, and how to size human pilot studies for validation.2026-06-15T18:06:20ZJoel PerssonMårten SchultzbergSebastian Ankargrenhttp://arxiv.org/abs/2606.02234v2When Do Treatment Changes Identify Causal Effects?2026-06-23T09:40:46ZThis paper clarifies the identifying assumptions underlying causal inference based on treatment changes rather than levels, and their relationship to conventional identification strategies. We characterize two structural models, with non-nested assumptions, under which treatment-change identification is valid conditional on observed covariates by differencing out time-constant confounders that are additive in the treatment equation. The assumptions underlying treatment changes are generally not nested with those of methods relying on treatment levels, such as selection-on-observables strategies that control for past outcomes, treatments, and covariates, or difference-in-differences approaches that difference outcomes rather than treatments over time. We show, however, that under a random-walk restriction on the treatment process, exploiting treatment changes for identification is equivalent to using treatment levels given lagged treatment. This and other equivalence results motivate overidentification tests based on methods considering treatment levels and changes. Under an alternative model that does not assume a random walk but instead rules out dynamic treatment effects (among other conditions), treatment changes can still be used as an instrument to identify a treatment effect that is constant given covariates. However, without random walk, different identification strategies are generally not nested. In partially linear models, the non-nesting results carry a double robustness implication for two-way fixed effects regression that differences both the outcome and the treatment over time, which under certain conditions remains consistent if either the treatment-change assumption or the parallel-trends assumption holds. We characterize the causal models consistent with each method, run simulations for illustration, and present an empirical application to cigarette demand.2026-06-01T13:26:48ZMartin Huberhttp://arxiv.org/abs/2606.24266v1Semi-nonparametric estimation of spatial dynamic panel data models with nonparametric spatial weights2026-06-23T07:52:07ZWe develop a semi-nonparametric framework for spatial dynamic panel data (SDPD) models with two-way fixed effects when the spatial interaction structure is unknown beyond a distance measure. This is accomplished by modelling spatial weights in the outcome, lagged-outcome, and disturbance channels as unknown functions of underlying economic distances. These enter the SDPD system through matrix-function operators, providing a unified approach that accommodates both spatial autoregressive and matrix exponential spatial specifications. Allowing for unknown heteroskedasticity, we propose sieve GMM estimators based on a stacked set of linear and quadratic moment conditions, and derive a feasible optimal GMM estimator and a more efficient feasible best GMM estimator. As $(n, T) \rightarrow \infty$, the parametric component is $\sqrt{n(T - 1)}$-consistent and asymptotically normal, echoing classical semi-nonparametric results. Monte Carlo experiments indicate excellent finite-sample performance. We apply the method to 'witch' killings as studied by Miguel (2005), and find that economic-geography proximity rather than cultural-geography proximity between communities significantly amplifies spatial dependence in these economic murders.2026-06-23T07:52:07ZAbhimanyu GuptaXi QuJiajun Zhanghttp://arxiv.org/abs/2606.24244v1When Surveys Become Conversations: Adaptive Matrix Validation for AI-Assisted Interviews2026-06-23T07:30:39ZAI-assisted interviews promise to reduce respondent burden in surveys by allowing respondents to describe experiences naturally while an AI system noisily maps those accounts into structured survey variables. That mapping is a measurement process that is fallible, versioned, adaptive, and potentially behaves differently across subgroups. This paper proposes Adaptive Matrix Validation (AMV), a design in which each respondent completes an AI-assisted interview, which is then mapped into tabular data by the AI. Respondents are also asked a small, randomized set of structured questions, which are used for statistical adjustment. The estimator first calibrates the mapped values using validation answers from other respondents, then corrects the remaining error with the validation answers observed for the target respondent. The paper develops estimators for item means, subgroup estimates, and regression coefficients when outcomes, predictors, or both are mapped from interviews. It also gives planning formulas the number of validation questions required and the sample size. A design-calibration simulation, an American Time Use Survey emulation, and a CHAMPS verbal-autopsy narrative study show when sparse validation can improve precision and when it cannot2026-06-23T07:30:39ZTyler H. McCormickhttp://arxiv.org/abs/2606.01706v2Higher-Order Debiased Estimators for General Treatment Models2026-06-23T07:29:23ZIt is now well known that estimators based on influence functions can be sub-optimal in terms of convergence rates in various settings. To address this issue, higher-order influence functions (HOIF) are developed, generalizing the classical semiparametric theory. However, most existing results in this regard focus on treatment effect parameters defined in explicit forms, such as average treatment effects (ATE). In applications, economists are often confronted with tasks of inferring more complex parameters, such as quantile treatment effects (QTE) or effects of complicated treatment regimes/policy. These more complex parameters can often only be implicitly defined as the solution to nonlinear estimating equations, which correspond to M/Z-estimation problems. Our current understanding of these problems is mainly limited to the classical semiparametric theory. Given the foundational role of HOIF for estimating explicit parameters such as ATE, a modest step toward enriching the statistical foundation of econometrics and causal inference is to develop the corresponding higher-order estimators for those more complex parameters. To this end, we consider parameters of a class of non-separable structural models in the econometrics literature and develop a class of higher-order estimators for the target parameters. Statistical properties of these higher-order estimators are derived using recent advances in U-processes theory. Our proposed higher-order estimators relax complexity-reducing assumptions, quantified by Holder smoothness, imposed on the nuisance parameters compared to existing alternative estimators for many important parameters in this class, including QTE and quantile dose-response functions, among others.2026-06-01T05:19:59ZYulin ZhangLin LiuZheng Zhanghttp://arxiv.org/abs/2606.24181v1Visible or Covert? The Causal Effect of Inspector Visibility on Fare Evasion Detection: A Causal Machine Learning and Policy Learning Approach2026-06-23T06:05:39ZFare evasion generates substantial revenue losses for public transport operators and is typically combated through fare inspections, yet little is known about how the mode of inspection-uniformed versus plainclothes-affects detection efficiency. Using a unique dataset of 21,727 inspection records from PostAuto, the largest regional bus operator in Switzerland, we apply causal machine learning to estimate the causal effect of inspector visibility on inspection efficiency, defined as detected fare evaders per inspection hour. Our results indicate that plainclothes inspections are, on average, significantly more effective than uniformed inspections, with an estimated average treatment effect of -0.173 incidents per hour, corresponding to a relative reduction of approximately 26%. Heterogeneity analyses find no evidence of systematic effect variation across contextual characteristics, suggesting that the superiority of plainclothes inspections is robust and pervasive across the PostAuto network. When applying optimal policy learning (based on policy trees) to optimally target subgroups by one or the other treatment depending on relative effectiveness, plainclothes inspections are recommended for the large majority of contexts (83.3%), with uniformed inspections suggested only for lines characterised by a below-median share of foreign residents and above-median population size.2026-06-23T06:05:39ZHannes WallimannCédric BrütschMartin Huber