https://arxiv.org/api/wVL7/InWAv908hbtKmQYkCtRdiI2026-07-21T20:24:16Z563010515http://arxiv.org/abs/2607.04257v1Randomization Tests in Randomized Saturation Designs2026-07-05T12:13:20ZRandomized saturation designs are widely used to study spillover effects in clustered populations. In these designs, clusters are first assigned to treatment saturation levels, and units are then randomized within clusters according to the assigned saturation. This paper develops randomization tests for such experiments under several null hypotheses that arise naturally in spillover analysis. For a fixed pair of saturation levels, we first study two individual-level hypotheses: a partially sharp null of no spillover effect for every untreated unit and a bounded null that restricts individual spillover effects by a prespecified constant. Both hypotheses can be tested using a common conditional randomization framework, with finite-sample validity obtained by combining the same focal-unit relabeling distribution with null-specific statistics. We then study weak average-spillover nulls and show that, although these nulls do not yield finite-sample exact conditional tests, studentized relabeling statistics deliver asymptotically valid randomization-based inference. Finally, for multiple ordered saturation levels, we develop a finite-sample valid unconditional pairwise-imputation test for global monotonicity of spillover effects. Simulations and an application to the Zomba Cash Transfer experiment illustrate the finite-sample behavior and practical implementation of the methods.2026-07-05T12:13:20ZJizhou LiuAzeem M. ShaikhLiang Zhonghttp://arxiv.org/abs/2602.16376v3Two-way Clustering Robust Variance Estimator in Quantile Regression Models2026-07-05T09:28:00ZWe study inference for linear quantile regression with two-way clustered data. Using a separately exchangeable array framework and a projection decomposition of the quantile score, we characterize regime-dependent convergence rates and establish a self-normalized Gaussian approximation. We propose a two-way cluster-robust sandwich variance estimator with a kernel-based density ``bread'' and a projection-matched ``meat'', and prove consistency and validity of inference in Gaussian regimes. We also show an impossibility result for uniform inference in a non-Gaussian interaction regime.2026-02-18T11:35:18ZUlrich HounyoJiahao Linhttp://arxiv.org/abs/2108.06473v3Evidence Aggregation for Treatment Choice2026-07-05T07:01:26ZConsider a planner who has limited knowledge of the policy's causal impact on a certain local population of interest due to a lack of data, but does have access to the publicized intervention studies performed for similar policies on different populations. How should the planner make use of and aggregate this existing evidence to make her policy decision? Following Manski (2020; Towards Credible Patient-Centered Meta-Analysis, \textit{Epidemiology}), we formulate the planner's problem as a statistical decision problem with a social welfare objective, and solve for an optimal aggregation rule under the minimax-regret criterion. We investigate the analytical properties, computational feasibility, and welfare regret performance of this rule. We apply the minimax regret decision rule to decide whether to enact an active labor market policy based on 14 randomized control trial studies.2021-08-14T05:53:42ZTakuya IshiharaToru Kitagawahttp://arxiv.org/abs/2607.03718v1Remote Work: Driver or Deterrent of Digital Product Innovation2026-07-04T05:49:23ZAs firms adopt divergent policies regarding work-from-home (WFH), the implications of remote work for collaborative and interdependent outcomes such as digital product innovation remain uncertain. This study examines how remote work adoption affects continuous digital product innovation using a panel dataset of mobile applications. We identify firm-level remote work adoption from job postings data and estimate its effects on app innovation using a staggered difference-in-differences design. We find that remote work significantly increases both major releases and new feature introductions per app, indicating enhanced digital product innovation performance. To assess whether these gains come at the expense of originality, we distinguish between novel and imitative feature introductions and show that remote work does not reduce the originality of digital product innovation. Moreover, improvements in digital product innovation translate into greater market success, as reflected in increased app downloads. The positive effects of remote work are stronger for app development teams with prior modular collaboration experience through open-source participation, suggesting that teams with greater experience coordinating modular work can better leverage remote work arrangements. We also find that remote work enables teams to expand their workforce and increase their collective skill capacity, both of which are associated with improved digital product innovation outcomes. In contrast, reductions in commuting time and app maturity do not explain the observed digital product innovation gains. Overall, our findings suggest that remote work can enhance continuous digital product innovation at the team level without compromising innovation novelty.2026-07-04T05:49:23ZFangchen SongYixuan LiuAshish Agarwalhttp://arxiv.org/abs/2607.03669v1Split-Session Cluster GARCH for Overnight and Intraday Returns: The Role of Tail Heterogeneity2026-07-04T02:45:02ZWe propose the Split-Session Cluster GARCH model for heavy-tailed multivariate dependence among asset returns decomposed into overnight and intraday components. The model uses convolution-$t$ distributions to allow tail behavior to differ across clusters defined by trading sessions and, within each session, by economic sectors. It also accommodates block-structured conditional correlation matrices, preserving parsimony and scalability in high-dimensional settings. The resulting likelihood remains tractable and yields a score-driven specification for dynamic correlations. We apply the model to U.S. equity returns in six-asset and 100-asset applications. The results reveal pronounced tail heterogeneity between overnight and intraday returns. Model comparisons show that session-specific tail parameters substantially improve fit relative to a common multivariate-$t$ specification, while sector-level tail partitioning delivers additional gains concentrated mainly in the overnight component. In the 100-asset application, asset-level tail heterogeneity delivers the strongest out-of-sample likelihood and global minimum-variance (GMV) portfolio performance.2026-07-04T02:45:02ZXinxian ChenPeter Reinhard HansenChen Tonghttp://arxiv.org/abs/2607.03665v1A Pseudo Panel Difference-in-Differences (DiD) Analysis of Online Shopping Behavior in the Puget Sound Regional Council (PSRC) Region2026-07-04T02:34:44ZOnline shopping is a growing trend, particularly following the COVID-19 lock-downs enacted by many cities. Understanding these trends requires robust panel dat methods. However, panel data (i.e., with repeated measurements of the same observational units) are often unavailable. In this study, we use a propensity score weighting (PSW) approach to adjust repeated cross-sectional travel diary surveys collected in the Puget Sound Regional Council (PSRC) region. The most common home delivery is packages (2.09 days per week in 2023), followed by food (0.36 days per week in 2023). We find that single-family detached and townhouse residents tend to receive more home deliveries than those living in apartments. We also find a difference in several patterns pre- and post-COVID pandemic. Vehicle deficient households did not exhibit the same increase in home delivery frequency as other households pre-COVID, but the pattern reversed in the post-COVID period - i.e., vehicle deficient households saw a relative increase in delivery frequency. Overall, this study demonstrates the pseudo-panel approach to causal inference by leveraging differences in PSW-weighted in delivery frequency across treatment groups.2026-07-04T02:34:44ZJason HawkinsUsman AhmedOmid Armantalabhttp://arxiv.org/abs/2202.04154v5Dynamic Heterogeneous Distribution Regression Panel Models, with an Application to Labor Income Processes2026-07-03T20:01:48ZWe introduce a dynamic distribution regression panel data model with heterogeneous coefficients across units. The objects of primary interest are functionals of these coefficients, including predicted one-step-ahead and stationary cross-sectional distributions of the outcome variable. Coefficients and their functionals are estimated via fixed effect methods. We investigate how these functionals vary in response to counterfactual changes in initial conditions or covariate values. We also identify a uniformity problem related to the robustness of inference to the unknown degree of coefficient heterogeneity, and propose a cross-sectional bootstrap method for uniformly valid inference on function-valued objects. We showcase the utility of our approach through an empirical application to individual income dynamics. Employing the annual Panel Study of Income Dynamics data, we establish the presence of substantial coefficient heterogeneity. We then highlight some important empirical questions that our methodology can address. First, we quantify the impact of a negative labor income shock on the distribution of future labor income. Second, we demonstrate the existence of heterogeneity in income mobility, and its implications for an individuals' incidence to be trapped in poverty. Simulation evidence confirms that our procedures work well in small samples.2022-02-08T21:30:54ZIvan Fernandez-ValWayne Yuan GaoYuan LiaoFrancis Vellahttp://arxiv.org/abs/2604.13188v2Is Productivity Advantage of Cities Really Down To Shift, Dilation, and Truncation?2026-07-03T17:35:21ZFirms in denser areas are more productive, owing to agglomeration and selection. To disentangle these channels, Combes et al. (2012, ECTA) assume that total factor productivity (TFP) distributions in denser and less dense areas are identical up to shift, dilation, and truncation. Using Spanish firm-level data and methods robust to noisy TFP estimates, we find that TFP distributions are indeed statistically identical up to these parameters, validating such decompositions. Furthermore, shifts and dilations alone are sufficient to capture distributional differences, at least in Spain. This suggests that policymakers should focus on agglomeration policies.2026-04-14T18:12:15ZVladislav MorozovAndrea Syhttp://arxiv.org/abs/2607.05440v1Retrieval over Reasoning: A Cost-Controlled Benchmark of Language Models for Energy-Retrofit Recommendation2026-07-03T16:52:55ZRecommending the correct set of energy conservation measures (ECMs) for a building is a structured, multi-label prediction problem in which a task-specific supervised model has weak training signal and a general language model has no grounding in the local building stock. We study this problem on 10,422 real New York City Local Law 87 (LL87) energy-audit records, taking as ground truth the set of ECM categories that certified auditors actually recommended. We make four contributions. First, we establish that energy-use-intensity (EUI) prediction - the upstream task - is effectively solved by tree ensembles: across fifteen trained models, a stacking ensemble reaches a coefficient of determination R^2 = 0.757, and every one of six neural architectures is outperformed by gradient-boosted trees. Second, we show that the framing of the recommendation task dominates model choice: recasting ECM recommendation as 19-way multi-label classification rather than single-label categorization lifts a gradient-boosted-tree baseline from a previously reported 25.9% accuracy to a micro-F1 of 0.571. Third, we benchmark eight large language models (LLMs) from four providers in a 2x2 design that independently toggles retrieval grounding and explicit reasoning, scoring each arm on per-label F1, U.S.-dollar cost per building, and latency; retrieval-augmented generation (RAG) improves micro-F1 by +0.11 to +0.20 on every model, while explicit reasoning yields no measurable accuracy change (-0.018 to +0.010) at up to 8.4x the cost. Fourth, we show LLMs systematically over-recommend - high recall, low precision - and that retrieval closes the gap chiefly by improving precision. A 70-billion-parameter open-weight model with a fifteen-line nearest-neighbor retrieval step reaches 0.511 micro-F1 at $0.00032 per building, comparable to a frontier model at roughly 10.1x lower cost.2026-07-03T16:52:55ZEliseo Curciohttp://arxiv.org/abs/1910.04610v3Robust Likelihood Ratio Tests for Incomplete Economic Models2026-07-03T16:10:15ZEconomic models with multiple equilibria, self-selection, or weak behavioral restrictions often make set-valued predictions and therefore do not imply a unique likelihood. This paper develops robust likelihood-ratio tests for structural hypotheses in such incomplete models. We evaluate tests by their power guarantee, defined as the smallest rejection probability over all selection mechanisms compatible with the alternative, while requiring uniform size control over all null selections. Using the Huber--Strassen theory of least favorable pairs, we construct finite-sample minimax likelihood-ratio tests. The main result shows that, in repeated experiments, the least favorable pair is the product of the single-experiment least favorable pairs whenever the latent variables are independent across experiments. This product structure holds even though unrestricted selection may induce arbitrary heterogeneity and dependence in the observed outcomes. It reduces a high-dimensional robust testing problem to single-experiment calculations and yields exact finite-sample critical values and Gaussian approximations. For directed one-sided alternatives, we provide conditions under which conditioning on a selection-invariant statistic delivers exact conditional uniformly most powerful (UMP) tests with respect to the power guarantee. In the examples, these optimal tests are simple and interpretable, using selection-invariant features of the data that are directly tied to the hypothesis of interest. Monte Carlo experiments in entry-game and Roy-model designs illustrate size control, power, and the role of robust testability.2019-10-10T14:41:10ZHiroaki KaidoYi Zhanghttp://arxiv.org/abs/2604.26088v3Stochastic Frontier meets Breakdown Frontier2026-07-03T15:24:22ZThis paper studies sensitivity analysis in stochastic frontier models by developing relaxations of the baseline assumptions imposed on the latent inefficiency and noise components, and characterize bounds for a benchmark technical-efficiency object under such relaxations. We then derive the associated breakdown frontier for conclusions about conditional technical efficiency and illustrate the procedure using a well-known dataset. We show the estimation and inference of the breakdown frontier. We also extend the analysis for widely used alternative specifications of the stochastic frontier error structure, that is, the Normal-Truncated Normal, Normal-Exponential, and Normal-Half Normal cases. Finally, we suggest avenues for extending this analysis under heteroskedasticity, the relaxation of input exogeneity, and applications using panel data. Code for empirical implementation is also provided.2026-04-28T20:02:18ZSantiago AcerenzaFrancisco Rosashttp://arxiv.org/abs/2607.03331v1When Does Heteroskedasticity Matter? A Contrast-Specific Theory of Robust Inference2026-07-03T13:51:22ZConventional heteroskedasticity diagnostics ask whether the conditional variance of the regression disturbance varies with covariates. This paper asks a different question: when does that variation matter for inference on the estimand of interest? The paper develops a contrast-specific theory characterizing when covariance perturbations are inferentially relevant. We show that, for any linear contrast $a'β$ in a linear regression, the difference between the heteroskedasticity-robust variance and the pooled fixed-design variance is governed by the empirical covariance between conditional error variance and a contrast-specific leverage score. Thus, heteroskedasticity may be present in the model yet first-order irrelevant for a particular coefficient or linear combination. Conversely, modest heteroskedasticity may have a large inferential effect if it is concentrated on observations that are highly informative for the contrast of interest. We characterize the effect exactly through a heteroskedasticity relevance ratio and a standard-error inflation factor, relate the result to pairs and residual bootstrap procedures, and extend the decomposition to general covariance structures, where off-diagonal dependence contributes a separate contrast-specific term. The results provide a unified way to understand why robust, clustered, and bootstrap standard errors can differ across coefficients in the same regression.2026-07-03T13:51:22ZUlrich Hounyohttp://arxiv.org/abs/2511.16029v3Possibilistic Instrumental Variable Regression with Potentially Invalid Instruments2026-07-03T09:46:11ZInstrumental variable regression is a common approach for causal inference in the presence of unobserved confounding. However, identifying valid instruments is often difficult in practice. In this paper, we propose a novel method based on possibility theory that performs posterior inference on the treatment effect, conditional on a user-specified set of potential violations of the instrument exogeneity assumption. Our method can provide valid results even when only a single, potentially invalid, instrument is available. Crucially, and in contrast with existing methods, we prove a finite-sample coverage guarantee for the exactly calibrated (validified) uncertainty intervals when the violation set contains the true value, and we provide practical MC/$χ^2$ approximations. Simulation experiments and real-data applications indicate strong performance of the proposed approach.2025-11-20T04:20:42ZGregor SteinerJeremie HoussineauMark F. J. Steelhttp://arxiv.org/abs/2607.03124v1Open Bitcoin Metrics: Verifiable Full-Node-Derived Bitcoin Time Series for Economic Research2026-07-03T09:10:56ZBitcoin research increasingly relies on on-chain indicators to study network activity, monetary issuance, transaction demand, miner incentives, coin-age behavior, and long-run monetary dynamics. However, many commonly used Bitcoin metrics are dispersed across commercial platforms, subject to heterogeneous definitions, or not fully reproducible from primary blockchain data. This manuscript introduces Open Bitcoin Metrics (OBM), a reproducible, full-node-derived dataset and reference guide for Bitcoin on-chain time series designed for economic and econometric research. The dataset provides documented daily series covering block production, block-space usage, transaction counts, supply, issuance, fees, miner revenue, mining difficulty, estimated hashrate, Bitcoin Days Destroyed, dormancy, liveliness, UTXO counts, spent output value, and related UTXO-age indicators. Metrics are reconstructed from a locally maintained Bitcoin Core full node, a persistent spent-output indexer, or deterministic transformations of previously generated OBM series. Each series is accompanied by open-source Python code, stable identifiers, explicit definitions, metadata, validation procedures, interpretive caveats, and comparisons with the closest publicly available metrics. The dataset is intended to support transparent empirical research, replication, teaching, and comparative analysis across monetary economics, financial economics, and blockchain studies.2026-07-03T09:10:56Z200 pages, 1 table, reference guide in appendix; dataset version OBM v0.1.0 archived on Zenodo; code and rolling updates available on GitHubDiego R. Llanoshttp://arxiv.org/abs/2607.02898v1Conformalized Lee Inference: Distribution-Free Individual Treatment Effect Intervals under Monotone Sample Selection2026-07-03T02:50:13ZEmpirical studies often observe outcomes only for selected units, and treatment may change who is observed. This paper studies prediction in randomized studies with one-sided selection. Standard prediction intervals can fail because treated selected observations are not the same group as selected controls. The paper asks how to predict missing treated outcomes and individual treatment effects for always-observed units. The proposed conformalized Lee procedure uses treated selected observations to train and check any prediction rule, then adjusts the cutoff using the observed treatment-control selection gap. For selected controls, the missing treated-outcome interval is shifted by the observed untreated outcome to produce an individual treatment-effect interval. The method provides reliable coverage without requiring the prediction rule to be correctly specified. The key result shows that the proposed adjustment uses the exact amount of uncertainty implied by the monotone selection logic of Lee [2009]. In simulations, ordinary conformal prediction demonstrates a lower coverage rate under selection-induced distribution shift, while the Lee-adjusted methods achieve the desired coverage rate. The results show that the proposed selection correction method can support reliable counterfactual prediction, while retaining practical implementation with modern prediction tools.2026-07-03T02:50:13ZJung Hyub Lee