https://arxiv.org/api/fHFpxRRMlUDXKInkbChK+5dGqdo 2026-10-02T17:27:12Z 5936 0 15 http://arxiv.org/abs/2609.10615v2 ACT, WAIT, or EXPERIMENT: A Causal Governance Framework for Retail Price Optimization Under Abstentions 2026-10-01T16:23:44Z This paper presents a causal decision-making framework for estimating price elasticity in retail channels, a process typically confounded by promotions, competitor movements, and market frictions. Rather than forcing a calculation when data is ambiguous, the system introduces decision abstention (\textsc{wait}) as an active diagnostic tool rather than an estimation failure. Combining Double Machine Learning and conformal prediction, the tool evaluates whether reliable conditions exist to adjust prices or if pausing the decision is preferable. When the system abstains, it exhaustively classifies the reason for the pause, identifying which products require designed pricing experiments or whether aggregating data to the brand level restores usable estimates. Tested on controlled synthetic data, the model shows that this operational discipline drastically reduces estimation error (lowering RMSE from 0.571 to 0.159) and offers a practical, secure alternative to blind estimation in thin-data retail environments. 2026-09-08T15:42:31Z Pedro Cadahia http://arxiv.org/abs/2610.01935v1 Pragmatic DML with AI-Learned Representations 2026-10-01T16:05:18Z Text, images, and other rich covariates are increasingly compressed into AI-learned representations and then used as controls in causal analysis. We study when this approach is valid and develop a practical framework for causal inference with learned representations. For a broad class of estimands, an imperfect representation distorts the target causal parameter by the product of two representation errors: one in the outcome regression and one in the balancing weight (or Riesz representer). This yields three constructive results. First, cross-fitted double machine learning (DML) provides valid Wald inference for the representation-dependent target. When representation errors are small, the same interval covers the causal parameter, and it can even attain the semiparametric efficiency bound. Second, fold-wise representation learning (or fine-tuning) is compatible with DML inference for the causal parameter. To this end, we develop convex- and star-aggregation pipelines for learning and combining representations. Third, when representation errors are substantial, we can provide interpretable sensitivity regions and root-$n$ inference for their endpoints. In a multi-modal demand application, seven representation-specific estimates and their star aggregate all imply a negative near-unit elasticity for rank-based price response, and the result remains robust over the reported sensitivity grid. 2026-10-01T16:05:18Z Andres Aradillas Fernandez Victor Chernozhukov Carlos Cinelli Sven Klaassen Whitney Newey Martin Spindler Jan Teichert-Kluge Suhas Vijaykumar http://arxiv.org/abs/2305.10256v3 Nowcasting using regression on signatures 2026-10-01T13:50:45Z We introduce a new method of nowcasting using regression on path signatures. Path signatures capture the geometric properties of sequential data. Because signatures embed observations in continuous time, they naturally handle mixed frequencies and missing data. We prove theoretically, and demonstrate with simulations, that regression on signatures both subsumes the linear Kalman filter and has desirable consistency properties. Nowcasting with signatures is more robust to disruptions in data series than previous methods, making it useful in stressed times (for example, during COVID-19). This approach is performant in nowcasting US GDP growth, and in nowcasting UK unemployment. 2023-05-17T14:44:06Z An early version of this paper was the result of a collaboration with Silvia Lui, Will Malpass, Andrew Reeves, Craig Scott, Emma Small from the UK Office for National Statistics. An earlier implementation of our algorithm in Python is available at https://github.com/alan-turing-institute/Nowcasting_with_signatures Samuel N. Cohen Giulia Mantoan Lars Nesheim Áureo de Paula Arthur Turrell Lingyi Yang http://arxiv.org/abs/2610.01654v1 Identifying Panel Conditioning with Refreshment Samples: Sharp Bounds and Design Assumptions 2026-10-01T13:17:10Z Refreshment samples are the standard remedy for panel attrition, and the identification results behind them maintain that participation does not change measurement. We characterize what a refreshment sample identifies about panel conditioning, modelled as a deterministic monotone map at reinterview, when attrition is unrestricted. A candidate map is consistent with the data if and only if the retention-scaled distribution of the stayers' implied latent outcomes is setwise dominated by the refreshment distribution; every such map is rationalized by an explicit attrition process. Without attrition the map is identified on the latent-outcome support; with attrition, a density-ratio condition governs the identified set, and tail behaviour alone does not determine it. For an unrestricted map the survivors' mean effect has the familiar trimming bounds; for an item with all categories reported, the model reduces to a test of no conditioning. Within a cohort, other waves, dropout patterns and entry-wave items leave the set unchanged unless restrictions link selection across waves. Under explicit selection restrictions, survival matching, symmetric matching and entry-wave correction identify survivor effects. We derive their biases and give rank conditions under which refreshment schedules identify curvature in the conditioning path. A Japanese panel illustrates the results. 2026-10-01T13:17:10Z 89 pages, 1 figure. Companion to arXiv:2609.28871. Replication archive: https://github.com/sokubo/paper-refreshment-designs-replication (tag paper-v1.5) Shoki Okubo http://arxiv.org/abs/2605.03997v2 Uncertainty Quantification in Forecast Comparisons 2026-10-01T11:23:56Z Skill scores, which measure the relative improvement of a forecasting method over a benchmark via consistent scoring functions and proper scoring rules, are a standard tool in forecast evaluation, yet their sampling uncertainty is rarely rigorously quantified. With modern forecasting applications being increasingly multivariate and involving evaluations across multiple horizons, variables, spatial locations, and forecasting methods, standard tools like the pairwise Diebold-Mariano forecast accuracy test or pointwise confidence intervals fail to account for the multiple comparison problem, leading to inflated Type I error rates and invalid joint inference. To address the lack of a coherent, statistically rigorous framework for quantifying uncertainty across these multi-dimensional evaluation problems, we introduce simultaneous confidence bands for expected scores and skill scores. Our framework provides a versatile tool for joint inference that is applicable to any forecast type from mean and quantile to full distributional forecasts. We develop a bootstrap implementation and show that our bands are valid under multivariate extensions of the classical Diebold-Mariano assumptions. We demonstrate the practical utility of the approach in two case studies by quantifying the benefits of time-varying parameter models for macroeconomic forecasting, and by comparing data-driven and physics-based models in probabilistic weather forecasting. 2026-05-05T17:20:48Z Marc-Oliver Pohle Tanja Zahn Sebastian Lerch http://arxiv.org/abs/2610.01465v1 A distributional modelling approach with application to electricity price forecasting 2026-10-01T10:59:16Z The increasing volatility of electricity prices driven by renewable energy integration, market shocks, and regulatory changes has reinforced the need for forecasting methods that go beyond point predictions and accurately describe the full conditional price distribution. This paper applies the Generalised Additive Models for Location, Scale and Shape (GAMLSS) framework to forecast Spanish day-ahead electricity prices using hourly data from 2020 to 2024. Alternative specifications based on Normal, Johnson's SU (JSU), and Sinh-Arcsinh (SHASH) distributions are considered, allowing the location, scale, and shape parameters to vary with market fundamentals, including electricity demand, renewable generation, seasonal effects, and regulatory and geopolitical risk factors. Forecasts are generated using a rolling-window approach and evaluated through the mean absolute error (MAE), pinball loss, and Diebold-Mariano tests. The results show that flexible distributional specifications improve forecasting performance relative to a naive benchmark and the standard normal specification. While SHASH and JSU specifications provide the lowest point forecasting errors, the hourly analysis reveals substantial intraday variation in relative performance across specifications. JSU specification with all four parameters driven by covariates achieves the best probabilistic forecasting performance, particularly in the tails of the distribution. Diebold-Mariano tests confirm the statistical significance of these improvements. These findings highlight the importance of modelling time-varying shape distributional parameters and demonstrate the value of GAMLSS models for forecasting and risk management in increasingly volatile electricity markets. 2026-10-01T10:59:16Z 25 pages, 2 figures, under review Aitor Ciarreta Peru Muniain Ainhoa Zarraga http://arxiv.org/abs/2610.01464v1 Inference after data-driven control-unit selection in difference-in-differences with estimated covariance 2026-10-01T10:57:28Z In difference-in-differences (DiD), researchers may use pre-treatment trends to select a control group for which the parallel-trends assumption appears plausible, with the aim of estimating the average treatment effect on the treated (ATT). Our earlier paper,Nakano and Hoshino (2016), and the present paper jointly provide the first selective-inference approach to the ATT that explicitly accounts for this control selection. We generalize our exact Gaussian procedure with known covariance to allow the covariance matrix to be estimated from the same individual-level data used for control selection and DiD estimation. We use this estimate to compute the variance, conditioning direction, residual, and truncation set. With fixed numbers of regions and periods, we establish uniform conditional coverage for selection events with probabilities bounded away from zero, and marginal coverage of the selected target without that restriction. We allow unequal regional sample sizes, heterogeneous covariances, ties in population fit, and regional sample shares that converge to zero. We establish asymptotic equivalence between the plug-in and known-covariance interval endpoints and derive rates for interval length. For staggered adoption, the control pools may differ across cohorts and periods, controls may be not yet treated, observations may be reused, and treatment effects may be heterogeneous. We also construct inference conditional on unions of selection paths that leave the reported parameter unchanged, together with simultaneous confidence bands for finitely many event-time effects. Under parallel trends and the other identifying conditions, the coverage results apply to the ATT. We give sufficient sampling conditions for individual panels and independent repeated cross-sections. 2026-10-01T10:57:28Z Ryoya Nakano Takahiro Hoshino http://arxiv.org/abs/2408.09187v3 Externally Valid Selection of Experimental Sites via the k-Median Problem 2026-10-01T01:05:06Z We present a decision-theoretic justification for viewing the question of how to best choose where to experiment in order to optimize external validity as a $k$-median problem, a popular problem in computer science and operations research. In particular, when treatment effect heterogeneity across experimental and policy-relevant sites is substantial (in a sense we make precise), we present conditions under which minimizing the worst-case, welfare-based regret among all nonrandom schemes that select $k$ sites to experiment is equivalent to solving a $k$-median problem. The connection costs in the relevant $k$-median problem are given by ex-ante bounds on worst-case voltage effects between sites, and minimizing the sum of worst-case voltage effects can be cast as a linear integer program. Two empirical applications illustrate the theoretical and computational benefits of the suggested procedure. 2024-08-17T12:56:12Z This version is a substantial update José Luis Montiel Olea Brenda Prallon Chen Qiu Jörg Stoye Yiwei Sun http://arxiv.org/abs/2610.00581v1 The Anatomy of Commodity Risk: Micro, Market, and Economy-Wide Sources 2026-09-30T18:48:10Z We study the anatomy of commodity risk by distinguishing micro, market-level, and economy-wide sources. We develop a two-stage "divide-and-conquer" framework that allows sensitivities to these risk sources to vary across commodities while treating economy-wide risk as latent. The first stage uses defactored instrumental-variable estimation to recover commodity-specific sensitivities to micro and market conditions. The second combines principal components with high-dimensional variable selection to identify an observable representation of macro-financial risk. We then construct Risk Intensity Indices (RIIs), which combine estimated sensitivities with prevailing risk conditions to quantify the relative importance of each risk source on a common scale. Market risk is the largest component on average, accounting for about two fifths of total risk intensity and more than half for energy commodities. Risk intensity is also highly concentrated across individual commodities: the top 20% account for approximately half of micro and market risk intensity, whereas macro risk is more broadly dispersed. The composition of risk varies substantially across sectors and over time, with market risk becoming particularly prominent during episodes of commodity-market stress. Micro and market RIIs also contain information about future volatility and absolute returns. These findings provide investors, risk managers, and policymakers with a diagnostic of where commodity risk is concentrated, which risk layers are most important, and how their importance changes over time. More broadly, our divide-and-conquer framework provides a flexible approach to decomposing layered risk in settings where common risk is latent. 2026-09-30T18:48:10Z Nektarios Aslanidis Aurelio Bariviera George Kapetanios Vasilis Sarafidis Alexia Ventouri http://arxiv.org/abs/2609.40156v1 Partial identification with entropy regularized optimal transport 2026-09-30T16:55:57Z In many statistical settings, the available data and maintained assumptions do not suffice to uniquely identify the model parameters of interest. In such cases, one can only identify sets which are guaranteed to contain the true parameters. These are often characterized through linear programs that optimize over models compatible with the observed data. These programs can be infinite-dimensional in the optimizer and the number of constraints. We provide a unified way to characterize and solve such optimization problems by phrasing them as optimal transport problems on path spaces. This allows us to regularize the problem with an entropy penalty, recasting it as a multi-marginal entropic optimal transport problem, which can be solved efficiently via Sinkhorn iterations. In addition, it allows us to establish convergence of the regularized value to the sharpest bound, derive consistency rates for a plug-in estimator, and obtain asymptotic distribution for approximate bounds. The method is general and accommodates settings ranging from instrumental variable models with continuous variables to welfare estimation in heterogeneous demand models. We verify the statistical and computational properties in simulations and provide an application to demand estimation. 2026-09-30T16:55:57Z Bruno N. Costa Florian F. Gunsilius http://arxiv.org/abs/2409.14202v4 Mining Causality: AI-Assisted Search for Instrumental Variables 2026-09-30T16:20:32Z The instrumental variables (IVs) method is a leading empirical strategy for causal inference. Finding IVs is a heuristic and creative process, and justifying their validity---especially exclusion restrictions---is largely rhetorical. We propose using large language models (LLMs) to search for new IVs through narratives and counterfactual reasoning, similar to how a human researcher would. The stark difference, however, is that LLMs can dramatically accelerate this process and explore an extremely large search space. We introduce a discovery pipeline that searches for potentially novel and valid IVs using two-step and role-playing prompting strategies. We contend that these strategies simulate the endogenous decision-making of economic agents and social actors and ground language models in real-world scenarios, thereby masking the IV discovery task itself. We apply our method to three canonical areas in economics: returns to schooling, demand and supply, and peer effects. We then introduce an evaluation pipeline in which we conduct expert surveys and train an LLM judge on the survey data. This evaluation reveals that the IV candidates discovered by our method appear both novel and likely valid, whereas a more direct search tends to return instruments already established in the literature. 2024-09-21T17:19:29Z Sukjin Han http://arxiv.org/abs/2106.12886v3 Constrained Classification and Policy Learning 2026-09-30T14:43:51Z Modern machine learning approaches to classification, including AdaBoost, support vector machines, and deep neural networks, utilize surrogate loss techniques to circumvent the computational complexity of minimizing empirical classification risk. These techniques are also useful for causal policy learning problems, since estimation of individualized treatment rules can be cast as a weighted (cost-sensitive) classification problem. Consistency of the surrogate loss approaches studied in Zhang (2004) and Bartlett et al. (2006) relies on the assumption of correct specification, which means that the specified set of classifiers is rich enough to contain a first-best classifier. This assumption is, however, less credible when interpretability or fairness constraints restrict the set of classifiers. Consequently, the applicability of surrogate-loss-based algorithms in such second-best scenarios remains unknown. This paper studies the consistency of surrogate loss procedures under a constrained set of classifiers without assuming correct specification. We show that in settings where the constraint restricts the classifier's prediction set only, hinge losses (i.e., $\ell_1$-support vector machines) are the only surrogate losses that preserve consistency in second-best scenarios. If the constraint additionally restricts the functional form of the classifier, consistency of a surrogate loss approach is not guaranteed, even with hinge loss. We therefore characterize conditions on the constrained set of classifiers that can guarantee consistency of hinge-risk-minimizing classifiers. Exploiting our theoretical results, we develop robust and computationally attractive hinge-loss-based procedures for a monotone classification problem. 2021-06-24T10:43:00Z Toru Kitagawa Shosei Sakaguchi Aleksey Tetenov http://arxiv.org/abs/2309.09299v5 Bounds on Average Effects in Discrete Choice Panel Data Models 2026-09-30T14:17:55Z In discrete choice panel data, estimation of average effects is crucial for quantifying the effect of covariates, and for policy evaluation and counterfactual analysis. However, in short panels with individual-specific effects, challenges arise due to partial identification and the incidental parameter problem. In particular, estimating the sharp identified set of average effects becomes impractical when covariates have large support sets, such as when they are continuous. This paper proposes a method for estimating outer bounds on the identified set of average effects, which are easy to construct, converge at the parametric rate, and remain computationally feasible even for moderately large samples. Asymptotically valid confidence intervals are also provided. 2023-09-17T15:25:14Z Cavit Pakel Martin Weidner http://arxiv.org/abs/2210.08698v4 A General Design-Based Framework and Estimator for Randomized Experiments 2026-09-30T11:41:22Z We describe a widely applicable design-based framework for drawing causal inference in randomized experiments. Causal effects are defined as linear functionals evaluated at unit-level potential outcome functions. Assumptions about the potential outcome functions are encoded as function spaces. This makes the framework expressive, allowing experimenters to formulate and investigate a wide range of causal questions that previously could not be investigated with design-based methods. The framework is particularly suited for complex, non-discrete interventions and causal interference. We describe a class of estimators for estimands defined using the framework and investigate their properties. We provide necessary and sufficient conditions for unbiasedness and consistency. We also describe a class of conservative variance estimators, which facilitate the construction of confidence intervals. In order to demonstrate the value of our approach in practice, we provide several illustrative examples of causal investigations which can be handled within our framework, but that could not be addressed using conventional design-based methods. 2022-10-17T02:19:11Z Christopher Harshaw Fredrik Sävje Yitan Wang http://arxiv.org/abs/2311.07243v2 Optimal Estimation of Large-Dimensional Nonlinear Factor Models 2026-09-30T07:56:37Z This paper studies optimal estimation of large-dimensional nonlinear factor models. The key challenge is that the observed variables are possibly nonlinear functions of some latent variables, with the functional forms left unspecified. A local principal component analysis method combining K-nearest neighbors matching and principal component analysis is proposed to estimate the factor structure and recover information on latent variables and latent functions. Large-sample properties are established, including a sharp bound on the matching discrepancy of nearest neighbors, sup-norm error bounds for estimated local factors and factor loadings, and the uniform convergence rate of the factor structure estimator. Under mild conditions our estimator of the latent factor structure can achieve the optimal rate of uniform convergence for nonparametric regression. The method is illustrated with a Monte Carlo experiment and an empirical application studying the effect of tax cuts on economic growth. 2023-11-13T11:25:00Z Yingjie Feng