https://arxiv.org/api/Ya+/5dVXQCqE5Yvf4Xnu2a4k4VQ 2026-07-21T21:22:29Z 5630 120 15 http://arxiv.org/abs/2607.02385v1 Inference for Group Interaction Experiments 2026-07-02T16:22:28Z A common experimental research design is one in which individuals are randomly allocated into groups that then interact under different group-level treatment conditions. We develop design-based inference for such "group interaction" experiments, covering scenarios in which groups are either fixed or randomly formed and in which potential outcomes are either fixed relative to others' group assignments or subject to interference. For each scenario, we characterize the causal estimand that the design targets and the inferential strategy appropriate to it. Working in a sparse-sampling asymptotic regime, we show that cluster-robust inference remains consistent and accounts for dependencies from various sources when interference is present, delivering valid inference on marginalized exposure effects. When interference is absent and groups are formed randomly, the design reduces to an individually randomized experiment, and individual-level heteroskedasticity-robust inference suffices for the average treatment effect. Our results on the asymptotic distribution of commonly used estimators rely on a novel coupling strategy that may be useful for design-based inference in other complex experiments. 2026-07-02T16:22:28Z Jiawei Fu Cyrus Samii Ye Wang http://arxiv.org/abs/2211.01921v7 Principal Component Analysis for High-Dimensional Approximate Factor Models in Time Series: Assumptions, Asymptotic Theory, and Identification 2026-07-02T12:46:00Z We consider estimation of large approximate factor models in high-dimensional panels of stationary time series using Principal Component Analysis (PCA). We review the key results establishing the necessary and sufficient conditions for consistency and asymptotic normality of the estimators which hold when both the cross-sectional dimension $n$ and the sample size $T$ tend to infinity. Special emphasis is placed on identification. First, we show that the common and idiosyncratic components are identified only in the limit $n\to\infty$. Second, we discuss the restrictions required to uniquely determine factors and loadings and examine their consequences for statistical inference. 2022-11-03T16:01:49Z Matteo Barigozzi http://arxiv.org/abs/2607.02095v1 Granular Instrumental Variables in Large Panels: Identification and Inference Across Strong, Nearly Weak, and Weak GIV 2026-07-02T12:32:28Z I develop the asymptotic theory of instrument strength for Granular Instrumental Variables (GIV) in large panels with both $N$ and $T$ growing. The strength of the GIV depends on the presence of dominant units. I formalise what dominance means and characterise three regimes of instrument strength. When a few units dominate the aggregate, the instrument is strong. The GIV estimator is consistent and asymptotically normal at the standard $\sqrt{T}$ rate. When large units stand out but do not dominate, the instrument weakens. But I show that the parameter of interest remains recoverable. The GIV estimator remains consistent and asymptotically normal, now at a rate slower than $\sqrt{T}$. When units are comparable in size and none stands out, the instrument is weak in the standard sense. The GIV estimator is inconsistent and has a non-standard distribution. Wald inference is reliable only outside the weak regime. When the instrument is weak, I recommend Anderson-Rubin confidence sets. In practice, the instrument must be constructed in a first stage. I show that the feasible estimator attains the same rate, but its asymptotic variance picks up an additional term from the first-stage estimation. Valid inference must use standard errors that account for this term. I apply the GIV estimator with the correct standard errors to recover the short-run demand elasticities of three commodities: refined copper, crude oil, and natural gas. 2026-07-02T12:32:28Z Job market paper. 129 pages, 2 figures. JEL: C33, C36, C55, C38, C12, Q41 Gokul Gopalan Ramachandran http://arxiv.org/abs/2607.01905v1 Measuring Opportunity Cost with Stock Lifetime Value 2026-07-02T09:05:08Z Measuring the long-term opportunity cost of interventions remains a critical challenge in e-commerce A/B testing. While strategic levers (such as dynamic pricing, ranking algorithms, and promotional campaigns) trigger shifts in consumer behaviour that persist over months, operational constraints necessitate fast decision-making cycles that are typically limited to weekly experimental windows. Standard metrics like revenue and conversion are inherently short-sighted, biasing decisions toward immediate gains. We introduce Stock Lifetime Value (SLV), a stock-centric metric that captures long-term opportunity cost within short experiments by aggregating expected profit from current inventory through the end of its selling lifecycle. We develop the methodology in the context of fashion e-commerce at Zalando, where stock constraints and seasonal lifecycles make the trade off between short-term and long-term outcomes particularly relevant. SLV aggregates the expected profit from current inventory through the end of its selling lifecycle, providing a way to evaluate interventions against their true profit impact. We discuss three applications: (a) SLV efficiency as a metric for article-level and customer-level A/B tests, validated against realized 18-month lifecycle outcomes; (b) SLV as an optimization target for pricing algorithms, aligning the metric used for measurement with the objective used for decision-making; and (c) a framework for annualizing treatment effects into financial reporting metrics required by business stakeholders. While our empirical setting is fashion retail, the framework applies broadly to any inventory-constrained environment where value decays over time or interventions shift demand across periods. 2026-07-02T09:05:08Z Geoffrey Decrouez Tobias Huelden Paresh Nakhe Dominik Prugger http://arxiv.org/abs/1911.04696v3 Extended MinP Tests for Global and Multiple testing 2026-07-02T03:38:39Z Empirical economic studies often involve multiple propositions or hypotheses, with researchers aiming to assess both the collective and individual evidence against these propositions or hypotheses. To rigorously assess this evidence, practitioners frequently employ tests with quadratic test statistics, such as $F$-tests and Wald tests, or tests based on minimum/maximum type test statistics. This paper introduces a combination test that merges these two classes of tests using the minimum $p$-value principle. The proposed test capitalizes on the global power advantages of both constituent tests while retaining the benefits of the stepdown procedure from minimum/maximum type tests. 2019-11-12T06:12:56Z 44 pages, 5 figures, 4 tables Zeng-Hua Lu http://arxiv.org/abs/2607.01429v1 A condition for the identification of multivariate models with binary instruments -- with Corrigendum and Addendum 2026-07-01T19:47:16Z This article introduces an empirical condition for the nonparametric point-identification of multivariate instrumental variable models with continuous endogenous variables using binary instruments. Verifying this condition can confirm point-identification in settings in which traditional approaches are not applicable. In particular, it shows that nonlinear instrumental variable models with general heterogeneity can be point-identified with only a binary instrument. This generalizes existing identification results which either restrict the unobserved heterogeneity substantially or require the instrument to have a large support. The main assumption on the instrumental variable model is cyclic monotonicity of its first stage, a multivariate generalization of the classical rank-invariance assumption for univariate models. Asymptotic convergence results for the empirical observable distributions are derived that allow to check the condition in practice. The identification rests on a fixed-set convergence result of cyclically monotone maps between quasi-concave functions. The corrigendum corrects the proof of Lemma 1. The proof given there incorrectly identifies preservation of distributional level sets with preservation of the underlying probability measure via Brenier maps. We replace that argument by one based on inverse Brenier maps, which play the role of multivariate ranks. The corrected argument applies to a different but significantly more flexible class of distributions than the quasi-concave class considered in the original paper. In particular, it allows for smooth non-quasi-concave and multimodal densities on compact supports, provided the associated rank fixed set satisfies a nondegeneracy condition. Moreover, it is generically satisfied for smooth parmetric classes of distributions. 2026-07-01T19:47:16Z Original paper published in the Journal of Econometrics including corrigendum and addendum Journal of Econometrics 235.1 (2023): 220-238 Florian Gunsilius http://arxiv.org/abs/2607.01377v1 Liquidity Premium and Investment Horizons 2026-07-01T18:41:20Z We estimate Kyle's (1985) price-impact coefficient $λ$ directly from daily equity order flow and test its ability to forecast the cross-section of subsequent stock returns. Using CRSP data from 2020 to 2025, we construct firm-month measures of signed order flow and two estimators of $\hatλ_{it}$: a within-month price-impact regression and an Amihud-style ratio. Signed order flow strongly predicts contemporaneous and one-month-ahead returns, while volume volatility predicts lower subsequent returns, consistent with widening price impact degrading price discovery. Fama-MacBeth regressions confirm that our order-flow signal carries significant cross-sectional return information after Newey--West adjustment. Theoretically, we resolve the liquidity premium puzzle of Constantinides (1986) through an adverse-selection mechanism: low order flow widens $λ$ and depresses prices today; subsequent normalization restores prices, generating the illiquidity premium without risk-based compensation. 2026-07-01T18:41:20Z 20 pages Irene Aldridge http://arxiv.org/abs/2606.03665v2 Sparse Tree-Based Aggregation for Time Series Regressions 2026-07-01T13:52:02Z High-dimensional time series regressions are often regularized to produce sparse coefficients. We show that temporal aggregation provides a powerful alternative to reduce dimensionality in high-order autoregressions and mixed-frequency regressions. To this end, we propose StarTime (Sparse Tree-based Aggregation for Time Series), a convex penalization method that uses a temporal tree to arrange lags hierarchically from high to low frequency. StarTime then flexibly selects coefficients to be aggregated at possibly varying frequencies, sparse or a combination thereof. We provide new error bounds for StarTime, demonstrate improved estimation accuracy and recovery of aggregation and sparsity in simulations relative to benchmarks, and illustrate StarTime's relevance for financial and macroeconomic applications. 2026-06-02T13:50:32Z Marie Corillon Stephan Smeekes Ines Wilms http://arxiv.org/abs/2212.04814v3 The Generalized Falsification Adaptive Set for Violations of the Exclusion Restriction and Exogeneity 2026-07-01T10:35:46Z The falsification adaptive set (FAS) as proposed by Masten and Poirier (2021) provides an identified set for a treatment effect when the baseline model is falsified, assuming invalid instruments violate exclusion only. We show that whether an invalid instrument is a confounder or collider has important consequences: incorrect treatment can cause the FAS to exclude the true parameter. We derive pattern-specific falsification adaptive sets for each combination of violations and propose a generalized FAS as their union, containing the true parameter value if any instrument is valid. We illustrate our results with the roads and trade application of Duranton et al. (2014). 2022-12-09T12:42:17Z Nicolas Apfel Frank Windmeijer http://arxiv.org/abs/2501.17455v3 A Uniform Confidence Band for the Marginal Treatment Effect Function 2026-07-01T07:14:54Z This paper presents a method for constructing uniform confidence bands for the marginal treatment effect (MTE) function. The shape of the MTE function provides insight into how the unobserved propensity to receive treatment relates to the treatment effect. Our approach visualizes the statistical uncertainty of an estimated function, facilitating inferences about the function's shape. The proposed method is computationally inexpensive and requires only minimal information: sample size, standard errors, kernel function, and bandwidth. We derive a Gaussian approximation for a local quadratic estimator and consider the approximation of the distribution of its supremum in polynomial order. Monte Carlo simulations demonstrate that our bands provide the desired coverage and are less conservative than those based on the Gumbel approximation. An empirical application based on the rural electrification program is included. 2025-01-29T07:32:42Z Toshiki Tsuda Yanchun Jin Ryo Okui http://arxiv.org/abs/2607.00312v1 Post-selection inference for network structure 2026-07-01T01:29:07Z Researchers often use the density of connections between groups of agents, such as communities, blocs, or markets, to characterize the structure of a social or economic network. In many cases, these groups are selected using the network data, making conventional fixed-group inference procedures potentially invalid. To address this issue, we develop two new confidence intervals that are universally valid post-selection in the sense that they guarantee simultaneous coverage asymptotically over all pairs of groups whose relative sizes do not vanish. Our first interval builds on a strategy of \cite{berk2013valid}. Our second interval is based on a Talagrand-type concentration inequality for empirical processes. Both intervals are simple to compute and scalable to large networks, but a key technical contribution of our paper is show that only the second interval achieves the best-possible width asymptotically up to a constant factor. Three empirical illustrations show that accounting for selection can matter in practice. Some evidence for homophily in a social network and a hub-and-spoke structure in a trade network survives our correction, while evidence for disjoint market segments in a worker transition network does not. 2026-07-01T01:29:07Z Eric Auerbach Jonathan Auerbach Sidonia McKenzie http://arxiv.org/abs/2607.00280v1 Understanding Guest Preferences and Optimizing Two-sided Marketplaces: Airbnb as an Example 2026-07-01T00:11:25Z Airbnb is a community based on connection and belonging -- many hosts on Airbnb are everyday people who share their worlds to provide guests with the feeling of connection and being at home; Airbnb strives to connect people and places. Among our efforts to connect guests and hosts, we provide tools to enable hosts to set competitive prices, which helps improve affordability for guests while helping hosts get more bookings. We also personalize the guest experience to show them the listings that match their needs. To help inform these efforts, we combine economic modeling and causal inference techniques to understand how guests book stays based on the prices hosts set, among other factors, and how that preference varies across different guests and listings. Such understanding helps us identify opportunities for Airbnb to support the marketplace and better connect guests and hosts. For example, understanding how much guests respond to different prices helps optimize the tools that we provide to hosts, in order to enable hosts to choose and set competitive prices that further balance demand and supply. As another example, understanding heterogeneity in guest preferences helps us personalize the guest experience and better match them with the listings that meet their needs, based on how much they respond to different prices and other factors. 2026-07-01T00:11:25Z 5 pages, 3 figures. Presented at the KDD 2024 Workshop on Two-Sided Marketplace Optimization, Barcelona, Spain Yufei Wu Daniel Schmierer http://arxiv.org/abs/2006.16997v8 Inference in Difference-in-Differences with Few Treated Units and Spatial Correlation 2026-06-30T22:00:50Z We consider the problem of inference in Difference-in-Differences (DID) when there are few treated units and errors are spatially correlated. We first show that, when there is a single treated unit, some existing inference methods designed for settings with few treated and many control units remain asymptotically valid when errors are weakly dependent. However, these methods may be invalid with more than one treated unit. We propose a menu of alternatives that are asymptotically valid in this setting, even when the relevant distance metric across units is unavailable. These alternatives vary in terms of the length of the resulting confidence intervals and the strength of the required assumptions. Our methods are also valid for comparison-of-means estimators and for construction of prediction intervals for counterfactual imputation methods. 2020-06-30T17:58:43Z Luis Alvarez Bruno Ferman http://arxiv.org/abs/2607.00219v1 Asymptotic Properties of Empirical Quantile-Based Estimators 2026-06-30T21:52:28Z We consider inference for parameters of the form $θ_0 = E[F_Y^{-1}\circ F_Z(X)]$ for some variables $X$, $Y$ and $Z$. Such parameters appear, in particular, in the ``changes-in-changes'' model of \cite{AtheyImbens2006}. We first establish that $\widehatθ$, a plug-in estimator of $θ_0$, is root-$n$ consistent and asymptotically normal under weaker conditions than those previously available, allowing in particular for unbounded variables. Next, we propose a new estimator of the asymptotic variance of $\widehatθ$ and show its consistency, also allowing for unbounded variables. Monte Carlo simulations suggest that the conditions for root-$n$ consistency and asymptotic normality are, in some sense, minimal. These simulations highlight that our variance estimator also leads to more accurate inference than some alternative approaches. 2026-06-30T21:52:28Z Julien Chhor Xavier D'Haultfœuille Jérémy L'Hour Martin Mugnier http://arxiv.org/abs/2606.31936v1 Coupling and Maximal Inequalities for Graph-Dependent Empirical Processes 2026-06-30T16:42:30Z We develop maximal inequalities for empirical processes indexed by graph-dependent observations. Our bounds separate the complexity of the indexing class from two features specific to graph dependence: the geometry of the underlying graph and the cost of coupling graph-separated blocks to independent copies. The coupling construction combines a novel graph-adapted dependence coefficient with a coloring of a block partition. We specialize the results to graphs with polynomial and exponential growth and to directed dyadic graphs. We then derive Glivenko--Cantelli results and characterize the associated effective sample size. A central implication is that graph-dependent empirical processes need not exhibit a generic root-$n$ rate: convergence is jointly determined by function-class complexity, graph geometry, and the decay of dependence with graph distance. Finally, we apply the results to obtain uniform laws of large numbers for network autoregressive models, nonlinear local-propagation models, and treatment-interference settings. 2026-06-30T16:42:30Z Mengsi Gao Demian Pouzo