https://arxiv.org/api/Ya+/5dVXQCqE5Yvf4Xnu2a4k4VQ2026-07-21T21:22:29Z563012015http://arxiv.org/abs/2607.02385v1Inference for Group Interaction Experiments2026-07-02T16:22:28ZA common experimental research design is one in which individuals are randomly allocated into groups that then interact under different group-level treatment conditions. We develop design-based inference for such "group interaction" experiments, covering scenarios in which groups are either fixed or randomly formed and in which potential outcomes are either fixed relative to others' group assignments or subject to interference. For each scenario, we characterize the causal estimand that the design targets and the inferential strategy appropriate to it. Working in a sparse-sampling asymptotic regime, we show that cluster-robust inference remains consistent and accounts for dependencies from various sources when interference is present, delivering valid inference on marginalized exposure effects. When interference is absent and groups are formed randomly, the design reduces to an individually randomized experiment, and individual-level heteroskedasticity-robust inference suffices for the average treatment effect. Our results on the asymptotic distribution of commonly used estimators rely on a novel coupling strategy that may be useful for design-based inference in other complex experiments.2026-07-02T16:22:28ZJiawei FuCyrus SamiiYe Wanghttp://arxiv.org/abs/2211.01921v7Principal Component Analysis for High-Dimensional Approximate Factor Models in Time Series: Assumptions, Asymptotic Theory, and Identification2026-07-02T12:46:00ZWe consider estimation of large approximate factor models in high-dimensional panels of stationary time series using Principal Component Analysis (PCA). We review the key results establishing the necessary and sufficient conditions for consistency and asymptotic normality of the estimators which hold when both the cross-sectional dimension $n$ and the sample size $T$ tend to infinity. Special emphasis is placed on identification. First, we show that the common and idiosyncratic components are identified only in the limit $n\to\infty$. Second, we discuss the restrictions required to uniquely determine factors and loadings and examine their consequences for statistical inference.2022-11-03T16:01:49ZMatteo Barigozzihttp://arxiv.org/abs/2607.02095v1Granular Instrumental Variables in Large Panels: Identification and Inference Across Strong, Nearly Weak, and Weak GIV2026-07-02T12:32:28ZI develop the asymptotic theory of instrument strength for Granular Instrumental Variables (GIV) in large panels with both $N$ and $T$ growing. The strength of the GIV depends on the presence of dominant units. I formalise what dominance means and characterise three regimes of instrument strength. When a few units dominate the aggregate, the instrument is strong. The GIV estimator is consistent and asymptotically normal at the standard $\sqrt{T}$ rate. When large units stand out but do not dominate, the instrument weakens. But I show that the parameter of interest remains recoverable. The GIV estimator remains consistent and asymptotically normal, now at a rate slower than $\sqrt{T}$. When units are comparable in size and none stands out, the instrument is weak in the standard sense. The GIV estimator is inconsistent and has a non-standard distribution. Wald inference is reliable only outside the weak regime. When the instrument is weak, I recommend Anderson-Rubin confidence sets. In practice, the instrument must be constructed in a first stage. I show that the feasible estimator attains the same rate, but its asymptotic variance picks up an additional term from the first-stage estimation. Valid inference must use standard errors that account for this term. I apply the GIV estimator with the correct standard errors to recover the short-run demand elasticities of three commodities: refined copper, crude oil, and natural gas.2026-07-02T12:32:28ZJob market paper. 129 pages, 2 figures. JEL: C33, C36, C55, C38, C12, Q41Gokul Gopalan Ramachandranhttp://arxiv.org/abs/2607.01905v1Measuring Opportunity Cost with Stock Lifetime Value2026-07-02T09:05:08ZMeasuring the long-term opportunity cost of interventions remains a critical challenge in e-commerce A/B testing. While strategic levers (such as dynamic pricing, ranking algorithms, and promotional campaigns) trigger shifts in consumer behaviour that persist over months, operational constraints necessitate fast decision-making cycles that are typically limited to weekly experimental windows. Standard metrics like revenue and conversion are inherently short-sighted, biasing decisions toward immediate gains. We introduce Stock Lifetime Value (SLV), a stock-centric metric that captures long-term opportunity cost within short experiments by aggregating expected profit from current inventory through the end of its selling lifecycle. We develop the methodology in the context of fashion e-commerce at Zalando, where stock constraints and seasonal lifecycles make the trade off between short-term and long-term outcomes particularly relevant. SLV aggregates the expected profit from current inventory through the end of its selling lifecycle, providing a way to evaluate interventions against their true profit impact. We discuss three applications: (a) SLV efficiency as a metric for article-level and customer-level A/B tests, validated against realized 18-month lifecycle outcomes; (b) SLV as an optimization target for pricing algorithms, aligning the metric used for measurement with the objective used for decision-making; and (c) a framework for annualizing treatment effects into financial reporting metrics required by business stakeholders. While our empirical setting is fashion retail, the framework applies broadly to any inventory-constrained environment where value decays over time or interventions shift demand across periods.2026-07-02T09:05:08ZGeoffrey DecrouezTobias HueldenParesh NakheDominik Pruggerhttp://arxiv.org/abs/1911.04696v3Extended MinP Tests for Global and Multiple testing2026-07-02T03:38:39ZEmpirical economic studies often involve multiple propositions or hypotheses, with researchers aiming to assess both the collective and individual evidence against these propositions or hypotheses. To rigorously assess this evidence, practitioners frequently employ tests with quadratic test statistics, such as $F$-tests and Wald tests, or tests based on minimum/maximum type test statistics. This paper introduces a combination test that merges these two classes of tests using the minimum $p$-value principle. The proposed test capitalizes on the global power advantages of both constituent tests while retaining the benefits of the stepdown procedure from minimum/maximum type tests.2019-11-12T06:12:56Z44 pages, 5 figures, 4 tablesZeng-Hua Luhttp://arxiv.org/abs/2607.01429v1A condition for the identification of multivariate models with binary instruments -- with Corrigendum and Addendum2026-07-01T19:47:16ZThis article introduces an empirical condition for the nonparametric point-identification of multivariate instrumental variable models with continuous endogenous variables using binary instruments. Verifying this condition can confirm point-identification in settings in which traditional approaches are not applicable. In particular, it shows that nonlinear instrumental variable models with general heterogeneity can be point-identified with only a binary instrument. This generalizes existing identification results which either restrict the unobserved heterogeneity substantially or require the instrument to have a large support. The main assumption on the instrumental variable model is cyclic monotonicity of its first stage, a multivariate generalization of the classical rank-invariance assumption for univariate models. Asymptotic convergence results for the empirical observable distributions are derived that allow to check the condition in practice. The identification rests on a fixed-set convergence result of cyclically monotone maps between quasi-concave functions. The corrigendum corrects the proof of Lemma 1. The proof given there incorrectly identifies preservation of distributional level sets with preservation of the underlying probability measure via Brenier maps. We replace that argument by one based on inverse Brenier maps, which play the role of multivariate ranks. The corrected argument applies to a different but significantly more flexible class of distributions than the quasi-concave class considered in the original paper. In particular, it allows for smooth non-quasi-concave and multimodal densities on compact supports, provided the associated rank fixed set satisfies a nondegeneracy condition. Moreover, it is generically satisfied for smooth parmetric classes of distributions.2026-07-01T19:47:16ZOriginal paper published in the Journal of Econometrics including corrigendum and addendumJournal of Econometrics 235.1 (2023): 220-238Florian Gunsiliushttp://arxiv.org/abs/2607.01377v1Liquidity Premium and Investment Horizons2026-07-01T18:41:20ZWe estimate Kyle's (1985) price-impact coefficient $λ$ directly from daily equity order flow and test its ability to forecast the cross-section of subsequent stock returns. Using CRSP data from 2020 to 2025, we construct firm-month measures of signed order flow and two estimators of $\hatλ_{it}$: a within-month price-impact regression and an Amihud-style ratio. Signed order flow strongly predicts contemporaneous and one-month-ahead returns, while volume volatility predicts lower subsequent returns, consistent with widening price impact degrading price discovery. Fama-MacBeth regressions confirm that our order-flow signal carries significant cross-sectional return information after Newey--West adjustment. Theoretically, we resolve the liquidity premium puzzle of Constantinides (1986) through an adverse-selection mechanism: low order flow widens $λ$ and depresses prices today; subsequent normalization restores prices, generating the illiquidity premium without risk-based compensation.2026-07-01T18:41:20Z20 pagesIrene Aldridgehttp://arxiv.org/abs/2606.03665v2Sparse Tree-Based Aggregation for Time Series Regressions2026-07-01T13:52:02ZHigh-dimensional time series regressions are often regularized to produce sparse coefficients. We show that temporal aggregation provides a powerful alternative to reduce dimensionality in high-order autoregressions and mixed-frequency regressions. To this end, we propose StarTime (Sparse Tree-based Aggregation for Time Series), a convex penalization method that uses a temporal tree to arrange lags hierarchically from high to low frequency. StarTime then flexibly selects coefficients to be aggregated at possibly varying frequencies, sparse or a combination thereof. We provide new error bounds for StarTime, demonstrate improved estimation accuracy and recovery of aggregation and sparsity in simulations relative to benchmarks, and illustrate StarTime's relevance for financial and macroeconomic applications.2026-06-02T13:50:32ZMarie CorillonStephan SmeekesInes Wilmshttp://arxiv.org/abs/2212.04814v3The Generalized Falsification Adaptive Set for Violations of the Exclusion Restriction and Exogeneity2026-07-01T10:35:46ZThe falsification adaptive set (FAS) as proposed by Masten and Poirier (2021) provides an identified set for a treatment effect when the baseline model is falsified, assuming invalid instruments violate exclusion only. We show that whether an invalid instrument is a confounder or collider has important consequences: incorrect treatment can cause the FAS to exclude the true parameter. We derive pattern-specific falsification adaptive sets for each combination of violations and propose a generalized FAS as their union, containing the true parameter value if any instrument is valid. We illustrate our results with the roads and trade application of Duranton et al. (2014).2022-12-09T12:42:17ZNicolas ApfelFrank Windmeijerhttp://arxiv.org/abs/2501.17455v3A Uniform Confidence Band for the Marginal Treatment Effect Function2026-07-01T07:14:54ZThis paper presents a method for constructing uniform confidence bands for the marginal treatment effect (MTE) function. The shape of the MTE function provides insight into how the unobserved propensity to receive treatment relates to the treatment effect. Our approach visualizes the statistical uncertainty of an estimated function, facilitating inferences about the function's shape. The proposed method is computationally inexpensive and requires only minimal information: sample size, standard errors, kernel function, and bandwidth. We derive a Gaussian approximation for a local quadratic estimator and consider the approximation of the distribution of its supremum in polynomial order. Monte Carlo simulations demonstrate that our bands provide the desired coverage and are less conservative than those based on the Gumbel approximation. An empirical application based on the rural electrification program is included.2025-01-29T07:32:42ZToshiki TsudaYanchun JinRyo Okuihttp://arxiv.org/abs/2607.00312v1Post-selection inference for network structure2026-07-01T01:29:07ZResearchers often use the density of connections between groups of agents, such as communities, blocs, or markets, to characterize the structure of a social or economic network. In many cases, these groups are selected using the network data, making conventional fixed-group inference procedures potentially invalid. To address this issue, we develop two new confidence intervals that are universally valid post-selection in the sense that they guarantee simultaneous coverage asymptotically over all pairs of groups whose relative sizes do not vanish. Our first interval builds on a strategy of \cite{berk2013valid}. Our second interval is based on a Talagrand-type concentration inequality for empirical processes. Both intervals are simple to compute and scalable to large networks, but a key technical contribution of our paper is show that only the second interval achieves the best-possible width asymptotically up to a constant factor. Three empirical illustrations show that accounting for selection can matter in practice. Some evidence for homophily in a social network and a hub-and-spoke structure in a trade network survives our correction, while evidence for disjoint market segments in a worker transition network does not.2026-07-01T01:29:07ZEric AuerbachJonathan AuerbachSidonia McKenziehttp://arxiv.org/abs/2607.00280v1Understanding Guest Preferences and Optimizing Two-sided Marketplaces: Airbnb as an Example2026-07-01T00:11:25ZAirbnb is a community based on connection and belonging -- many hosts on Airbnb are everyday people who share their worlds to provide guests with the feeling of connection and being at home; Airbnb strives to connect people and places. Among our efforts to connect guests and hosts, we provide tools to enable hosts to set competitive prices, which helps improve affordability for guests while helping hosts get more bookings. We also personalize the guest experience to show them the listings that match their needs.
To help inform these efforts, we combine economic modeling and causal inference techniques to understand how guests book stays based on the prices hosts set, among other factors, and how that preference varies across different guests and listings. Such understanding helps us identify opportunities for Airbnb to support the marketplace and better connect guests and hosts. For example, understanding how much guests respond to different prices helps optimize the tools that we provide to hosts, in order to enable hosts to choose and set competitive prices that further balance demand and supply. As another example, understanding heterogeneity in guest preferences helps us personalize the guest experience and better match them with the listings that meet their needs, based on how much they respond to different prices and other factors.2026-07-01T00:11:25Z5 pages, 3 figures. Presented at the KDD 2024 Workshop on Two-Sided Marketplace Optimization, Barcelona, SpainYufei WuDaniel Schmiererhttp://arxiv.org/abs/2006.16997v8Inference in Difference-in-Differences with Few Treated Units and Spatial Correlation2026-06-30T22:00:50ZWe consider the problem of inference in Difference-in-Differences (DID) when there are few treated units and errors are spatially correlated. We first show that, when there is a single treated unit, some existing inference methods designed for settings with few treated and many control units remain asymptotically valid when errors are weakly dependent. However, these methods may be invalid with more than one treated unit. We propose a menu of alternatives that are asymptotically valid in this setting, even when the relevant distance metric across units is unavailable. These alternatives vary in terms of the length of the resulting confidence intervals and the strength of the required assumptions. Our methods are also valid for comparison-of-means estimators and for construction of prediction intervals for counterfactual imputation methods.2020-06-30T17:58:43ZLuis AlvarezBruno Fermanhttp://arxiv.org/abs/2607.00219v1Asymptotic Properties of Empirical Quantile-Based Estimators2026-06-30T21:52:28ZWe consider inference for parameters of the form $θ_0 = E[F_Y^{-1}\circ F_Z(X)]$ for some variables $X$, $Y$ and $Z$. Such parameters appear, in particular, in the ``changes-in-changes'' model of \cite{AtheyImbens2006}. We first establish that $\widehatθ$, a plug-in estimator of $θ_0$, is root-$n$ consistent and asymptotically normal under weaker conditions than those previously available, allowing in particular for unbounded variables. Next, we propose a new estimator of the asymptotic variance of $\widehatθ$ and show its consistency, also allowing for unbounded variables. Monte Carlo simulations suggest that the conditions for root-$n$ consistency and asymptotic normality are, in some sense, minimal. These simulations highlight that our variance estimator also leads to more accurate inference than some alternative approaches.2026-06-30T21:52:28ZJulien ChhorXavier D'HaultfœuilleJérémy L'HourMartin Mugnierhttp://arxiv.org/abs/2606.31936v1Coupling and Maximal Inequalities for Graph-Dependent Empirical Processes2026-06-30T16:42:30ZWe develop maximal inequalities for empirical processes indexed by graph-dependent observations. Our bounds separate the complexity of the indexing class from two features specific to graph dependence: the geometry of the underlying graph and the cost of coupling graph-separated blocks to independent copies. The coupling construction combines a novel graph-adapted dependence coefficient with a coloring of a block partition. We specialize the results to graphs with polynomial and exponential growth and to directed dyadic graphs. We then derive Glivenko--Cantelli results and characterize the associated effective sample size. A central implication is that graph-dependent empirical processes need not exhibit a generic root-$n$ rate: convergence is jointly determined by function-class complexity, graph geometry, and the decay of dependence with graph distance. Finally, we apply the results to obtain uniform laws of large numbers for network autoregressive models, nonlinear local-propagation models, and treatment-interference settings.2026-06-30T16:42:30ZMengsi GaoDemian Pouzo