https://arxiv.org/api/hYsPBjA9eXsWIFfNb8Utz98ujSM 2026-07-21T22:16:18Z 5630 135 15 http://arxiv.org/abs/2606.31935v1 Delegation Rights: Property, Agency, and Investment Incentives in the Age of AI Agents 2026-06-30T16:39:04Z AI agents increasingly operate inside digital accounts by exercising privileges that users already hold, raising a new control question: whether an existing account entitlement must be exercised manually or may be exercised through a user-authorized automated proxy. We define \emph{delegation rights} as the revocable, identity-preserving, scope-limited, and mode-specific authority of an account holder to authorize such proxy execution. We develop a three-party incomplete-contracts model with a User, an AI Agent provider, and a Platform. The contested object is not platform ownership, account transferability, data portability, or unrestricted API access, but residual control over the mode of account execution. Under Platform Control, the platform can protect infrastructure, identity systems, privacy boundaries, and third parties, but its discretionary veto weakens the User--Agent coalition's disagreement payoff and depresses relationship-specific investment. Under User Control, hold-up is reduced, but security, privacy, congestion, and third-party risks may remain insufficiently internalized. We then analyze \emph{Certified Delegation}, under which access protection is conditional on verifiable authorization, revocability, auditability, rate-limit compliance, data minimization, and risk mitigation. Certification is therefore not merely a technical safety screen; it is a conditional allocation of residual control. Illustrative mechanism simulations show how this regime can reduce deadweight loss by restoring delegation incentives while bounding residual risk. 2026-06-30T16:39:04Z Yukun Zhang Kemu Xu http://arxiv.org/abs/2606.31930v1 Quasi-Bayesian Hierarchical Models 2026-06-30T16:36:58Z We develop the Quasi-Bayesian Hierarchical Model (QBHM) for grouped GMM settings. The framework combines Bayesian hierarchical modelling with Laplace-type estimation: it preserves each group-specific objective function, while introducing a pooling term for economically comparable parameters. When the number of studies is fixed, the QBHM estimator-the quasi-posterior mean-has the same asymptotic distribution as GMM when estimating strongly identified study parameters. For weakly identified studies, we analyze the asymptotic properties of the method via a weak-GMM limit experiment: an asymptotic approximation in which the sample-moment criterion remains a random function over the weak parameter space, and the upper-level pooling relation induces a family of priors over weak values. In this experiment, the weak-limit QBHM rule is a Bayes rule under squared loss for the hierarchy-induced weak-limit prior, which provides a decision-theoretic justification for our procedure. We also extend our results to mixed within-study blocks, allowing a single study to contain both strongly and weakly identified parameters. Pooling can also reduce the pointwise asymptotic mean squared error (MSE) relative to unpooled estimation when the bias--variance tradeoff is favorable. Gaussian likelihood, nonlinear weak-GMM, and weak-IV calculations show when this happens, while simulations and a microenterprise application illustrate the method. 2026-06-30T16:36:58Z Desmond Fairall Thomas Glinnan http://arxiv.org/abs/2606.31685v1 Design-Based Inference for Time-Series GMM 2026-06-30T14:00:09Z This paper studies inference for time-series GMM when uncertainty comes from shock assignment within a realized historical episode. Rather than treating the data as one random draw from a population of hypothetical economies, the framework conditions on the historical environment and considers alternative realizations of shocks and instruments. For locally correctly specified GMM estimators, the centered moment has design long-run variance $Ω_R$, which determines the sandwich covariance for the finite-history estimand. Conventional HAC estimators instead converge to $Ω_R^+=Ω_R+Ω_μ$, where $Ω_μ\succeq0$ is the long-run variance of the centered mean-moment path. HAC inference is therefore conservative for scalar functions of the finite-history estimand. Projection adjustment using predetermined covariates can reduce this HAC variance limit in Loewner order and, under an additional long-run orthogonality condition, yields a tighter conservative bound on the corresponding asymptotic covariance. Monte Carlo evidence shows when the distinction is quantitatively important. In a monetary-policy application, standard-error reductions from rich macro covariates provide a diagnostic for economically meaningful predictable variation in the mean-moment path. 2026-06-30T14:00:09Z Thomas Glinnan http://arxiv.org/abs/2602.03981v2 DeXposure-FM: A Time-series, Graph Foundation Model for Credit Exposures and Stability on Decentralized Financial Networks 2026-06-30T12:22:31Z Credit exposure in Decentralized Finance (DeFi) is often implicit and token-mediated, creating a dense web of inter-protocol dependencies. Thus, a shock to one token may result in significant and uncontrolled contagion effects. As the DeFi ecosystem becomes increasingly linked with traditional financial infrastructure through instruments, such as stablecoins, the risk posed by this dynamic demands more powerful quantification tools. We introduce DeXposure-FM, the first time-series, graph foundation model for measuring and forecasting inter-protocol credit exposure on DeFi networks, to the best of our knowledge. Employing a graph-tabular encoder, with pre-trained weight initialization, and multiple task-specific heads, DeXposure-FM is trained on the DeXposure dataset that has 43.7 million data entries, across 4,300+ protocols on 602 blockchains, covering 24,300+ unique tokens. The training is operationalized for credit-exposure forecasting, predicting the joint dynamics of (1) protocol-level flows, and (2) the topology and weights of credit-exposure links. The DeXposure-FM is empirically validated on two machine learning benchmarks; it consistently outperforms the state-of-the-art approaches, including a graph foundation model and temporal graph neural networks. DeXposure-FM further produces financial economics tools that support macroprudential monitoring and scenario-based DeFi stress testing, by enabling protocol-level systemic-importance scores, sector-level spillover and concentration measures via a forecast-then-measure pipeline. Empirical verification fully supports our financial economics tools. The model and code have been publicly available. Model: https://huggingface.co/EVIEHub/DeXposure-FM. Code: https://github.com/EVIEHub/DeXposure-FM. 2026-02-03T20:10:11Z Aijie Shu Wenbin Wu Gbenga Ibikunle Fengxiang He http://arxiv.org/abs/2107.07942v6 Flexible Covariate Adjustments in Regression Discontinuity Designs 2026-06-30T11:20:24Z Empirical regression discontinuity (RD) studies often include covariates in their specifications to increase the precision of their estimates. In this paper, we propose a novel class of estimators that use such covariate information more efficiently than existing methods and can accommodate many covariates. Our estimators are simple to implement and involve running a standard RD analysis after subtracting a function of the covariates from the original outcome variable. We characterize the function of the covariates that minimizes the asymptotic variance of these estimators. We also show that the conventional RD framework gives rise to a special robustness property which implies that the optimal adjustment function can be estimated flexibly via modern machine learning techniques without affecting the first-order properties of the final RD estimator. We demonstrate our methods' scope for efficiency improvements by reanalyzing data from a large number of recently published empirical studies. 2021-07-16T15:00:06Z Claudia Noack Tomasz Olma Christoph Rothe http://arxiv.org/abs/2603.00248v2 Targeted Local Projections 2026-06-30T10:56:30Z Local projection (LP) and structural vector autoregression (SVAR) are commonly employed to estimate dynamic causal effects of macroeconomic policies at multiple horizons. With enough lags as controls, LP estimators have little bias but their variance can increase with the horizon due to accumulating additional shocks. Because they typically employ fewer lags or suffer from local misspecification, SVAR estimators typically incur higher bias, but their variance decreases with the horizon due to exponentiation. We propose to target the LP estimators towards their SVAR counterparts - constructed with fewer lags than LP at each horizon - to reduce their variance at the cost of incurring some bias. The resulting targeted LP estimator is a linear combination of the LP and SVAR estimators. We propose choosing this linear combination optimally to minimize the mean-squared error of the new estimator. Our simulations show that, under a locally misspecified SVAR model, targeting substantially reduces the LP variance at longer horizons while maintaining near-nominal coverage in small samples when a double bootstrap is employed. 2026-02-27T19:02:54Z Aleksei Nemtyrev Otilia Boldea http://arxiv.org/abs/2408.06624v4 Estimation and Inference on Average Treatment Effect in Percentage Points under Heterogeneity 2026-06-30T02:35:27Z In semi-logarithmic regressions, treatment coefficients are often interpreted as approximations of an average treatment effect (ATE) in percentage points. This paper highlights the overlooked bias of this approximation under treatment effect heterogeneity, arising from Jensen's inequality. The issue is particularly relevant for difference-in-differences designs with log-transformed outcomes and staggered treatment adoption, where treatment effects may vary across groups and periods. This paper proposes new estimation and inference methods for an estimand that accounts for heterogeneity across observable subgroups and can improve upon conventional measures. The estimand provides a lower bound on the ATE in percentage points for the relevant target (sub)population, and coincides with it in the absence of within-group heterogeneity. I establish the methods' large-sample properties and study their finite-sample performance through Monte Carlo experiments, which reveal substantial discrepancies between conventional and proposed measures when systematic heterogeneity is large. Two empirical applications further underscore the practical importance of these methods. 2024-08-13T04:23:55Z Ying Zeng http://arxiv.org/abs/2606.30999v1 Estimating Supply Incrementality in Two-sided Marketplaces: A Causal Machine Learning Approach 2026-06-30T00:23:00Z In two-sided marketplaces with heterogeneous products, it is important to understand the causal relationship between additional supply and marketplace outcomes, such as the total quantity transacted or transaction value in the marketplace. This paper studies a causal machine learning approach to estimating this relationship across product segments. We use the Airbnb marketplace as an example, focusing on the impact of additional listing supply on total bookings, but the methodology applies to other two-sided marketplaces. Our approach combines double/debiased machine learning with a hierarchical Bayesian framework that leverages pre-existing knowledge as priors. We construct tractable and informative features for the model by leveraging measures of product segment similarity from the geospatial literature. We find that such a model provides plausible estimates of the marketplace returns to additional supply and strong out of sample performance. 2026-06-30T00:23:00Z 5 pages, 3 figures. Accepted at the KDD 2025 Workshop on Causal Inference and Machine Learning in Practice (not presented) Yufei Wu Daniel Schmierer Dan Zylberglejd http://arxiv.org/abs/2011.08174v10 Policy design in experiments with unknown interference 2026-06-29T15:34:40Z This paper studies experimental designs for estimation and inference on policies with spillover effects. Units are organized into a finite number of large clusters and interact in unknown ways within each cluster. First, we introduce a single-wave experiment that, by varying the randomization across cluster pairs, estimates the marginal effect of a change in treatment probabilities, taking spillover effects into account. Using the marginal effect, we propose a test for policy optimality. Second, we design a multiple-wave experiment to estimate welfare-maximizing treatment rules. We provide strong theoretical guarantees and an implementation in a large-scale field experiment. 2020-11-16T18:58:54Z Davide Viviano Jess Rudder http://arxiv.org/abs/2606.30040v1 The Shape of Macroeconomic Beliefs 2026-06-29T09:35:50Z Macroeconomic expectations are usually observed through point forecasts or through asset prices whose mapping into beliefs is model-dependent. This paper uses prediction-market prices to recover high-frequency distributions of short-run macroeconomic beliefs. We construct a panel of Kalshi-implied distributions for CPI and core CPI releases by converting adjacent threshold contracts into probability mass over inflation outcomes. The data reveal market-implied means, uncertainty, and upper-tail probabilities from 30 days to one hour before each release. The market-implied mean contains meaningful forecast information, especially for headline CPI, but the main signal is distributional. Lagged Reuters Poll surprises do not predict systematic deviations of Kalshi means from the current Reuters consensus. By contrast, large lagged surprises are associated with higher implied uncertainty, and positive lagged surprises raise the probability assigned to fixed high-inflation outcomes. In the baseline specification with variable-by-horizon fixed effects, a 0.1 percentage point positive lagged surprise raises the probability of monthly inflation above 0.3 percent by about 4.7 percentage points, even after controlling for the current consensus forecast. In release-level validation tests, Kalshi upper-tail probabilities also predict the realization of high-inflation states, including episodes in which the market-implied mean remains close to the Reuters consensus. The evidence suggests that prediction markets can provide real-time information about inflation risk that is missed by point forecasts. 2026-06-29T09:35:50Z Giovanni Angelini http://arxiv.org/abs/2604.10845v3 Learning Preferences from Conjoint Data: A Hybrid Structural Deep Learning Approach 2026-06-29T07:40:50Z Conjoint experiments randomize multidimensional profiles, yet political science applications typically report only nonparametric averages that do not recover counterfactual choices or individual tradeoffs. We develop a hybrid structural approach for recovering individual preferences from conjoint data. The estimator combines a flexible machine-learning mean preference function, via a deep netural network in our applications, with respondent-level empirical-Bayes updating in a logistic random utility model, allowing preferences to vary with observed characteristics while learning residual heterogeneity from repeated choices. Double/debiased machine learning delivers valid inference for population-average preference parameters with any sufficiently accurate first-stage learner. Across three applications, the method reveals heterogeneity reduced-form averages obscure: opposition to undemocratic behavior is broad but uneven in intensity, progressive tax preferences are widespread across partisan subgroups, and partisan polarization offsets the average gender effect in candidate choice. The framework opens the door to core theoretical questions in political science by recovering substantively interpretable structural parameters. 2026-04-12T22:35:04Z Avidit Acharya Jens Hainmueller Yiqing Xu http://arxiv.org/abs/2606.29833v1 Sensitivity, Informativeness, and Misspecification in GMM Estimation 2026-06-29T06:18:55Z This paper develops misspecification-robust sensitivity and informativeness diagnostics for GMM estimators, evaluated at pseudo-true values. The sensitivity matrix nests that of Andrews, Gentzkow, and Shapiro (2017) under correct specification. The informativeness $Δ$ measures the share of an estimator's asymptotic variance explained by sampling variation in the moments, a notion of structural efficiency that equals one under correct specification and can fall below one under misspecification, even when the Hansen $J$-test does not reject. We derive influence-function representations for one-step, two-step, iterated, and continuously updating GMM. We show that in minimum-distance estimation, estimating the optimal weight matrix adds estimator variance that the moments do not explain, lowering informativeness, while simpler weight matrices largely avoid it. The choice of weight matrix therefore involves a trade-off between classical efficiency and informativeness. In applications to the automobile demand model of Berry, Levinsohn, and Pakes (1995), the consumption insurance model of Blundell, Pistaferri, and Preston (2008), and the income-and-democracy regressions of Acemoglu, Johnson, Robinson, and Yared (2008), misspecification reorders sensitivity rankings, simpler weights preserve the informativeness that the optimal weight loses, and $Δ$ detects structural-efficiency losses that the $J$-test does not. 2026-06-29T06:18:55Z Fangzhou Yu Seojeong Lee http://arxiv.org/abs/2606.29784v1 HERO: Improving the Reliability and Sensitivity of Generative Model Evaluation Using Historical Data 2026-06-29T05:04:57Z Reliable generative AI models critically rely on expert human annotations to evaluate output quality, yet these "gold" labels are expensive to collect and limited in quantity. Organizations thus often turn to collecting vast but noisy "silver" labels from crowdsourced workers or vendor annotators as proxies for gold labels. Because gold remains the evaluation target, naively aggregating noisy silver labels may introduce bias, and estimators built on sparsely observed gold labels may have high variance to resolve the model performance gaps that guide practical decisions. Model evaluation has become an ongoing operational practice rather than a one-time exercise, with evaluation rounds repeating across model versions, releases, and content domains. A natural question is whether the previous historical evaluation data can be used to improve each new round of evaluation. We introduce HERO (History Enhanced RObust model evaluation), a novel framework that uses historical data to suppress bias (improve reliability) and reduce variance (improve sensitivity) in model performance evaluation. HERO calibrates silver labelers' performance learned from historical gold annotations, and stabilizes the resulting estimator by anchoring it to covariate information measured with high precision in the historical data. HERO can be broadly applied across multiple common evaluation tasks, and remains valid when only a subset of historical labelers appears in the current round. We establish conditions under which the bias and variance reductions hold, showcase HERO's performance in simulation studies, and demonstrate its effectiveness on real-world model evaluation benchmarking datasets. 2026-06-29T05:04:57Z 30 pages, 6 figures Xinrui Ruan Zhenyu Zhao Waverly Wei Yueshan Zhang Zeyu Zheng Sui Huang Jingshen Wang http://arxiv.org/abs/2105.09254v4 Multiply Robust Causal Mediation Analysis with Continuous Treatments 2026-06-29T04:42:27Z In many applications, researchers are interested in the direct and indirect causal effects of a treatment or exposure on an outcome of interest. Mediation analysis offers a rigorous framework for identifying and estimating these causal effects. For binary treatments, efficient estimators for the direct and indirect effects are presented by Tchetgen Tchetgen and Shpitser (2012) based on the influence function of the parameter of interest. These estimators possess desirable properties such as multiple-robustness and asymptotic normality while allowing for slower than root-n rates of convergence for the nuisance parameters. However, in settings involving continuous treatments, these influence function-based estimators are not readily applicable without making strong parametric assumptions. In this work, utilizing a kernel smoothing approach, we propose an estimator suitable for settings with continuous treatments inspired by the influence function-based estimation strategy. Our proposed approach employs cross-fitting, relaxing the smoothness requirements on the nuisance functions and allowing them to be estimated at slower rates than the target parameter. Additionally, similar to influence function-based estimators, our proposed estimator is multiply robust and asymptotically normal, allowing for inference in settings where parametric assumptions may not be justified. 2021-05-19T16:58:57Z Yizhen Xu AmirEmad Ghassami Numair Sani Ilya Shpitser http://arxiv.org/abs/2606.29756v1 Modeling Mode and Departure Time Responses to Congestion Pricing: A Spatial and Behavioral Analysis Using Cross-Nested Logit Model 2026-06-29T04:00:45Z Effective congestion management strategies require a detailed understanding of how travellers respond to different pricing interventions. This paper presents an in-depth analysis of traveller behaviour under congestion pricing scenarios, focusing specifically on mode and departure time decisions. Utilizing stated preference survey data from commuters in Calgary, Canada, three discrete choice models including Multinomial Logit, Nested Logit, and Cross-Nested Logit are developed and compared. Results indicate that the Cross-Nested Logit model provides superior behavioural realism and flexibility by capturing simultaneous substitutions across modes and departure times. Spatial analysis and elasticity assessments reveal substantial geographic variation in traveller sensitivity to pricing, particularly highlighting stronger responses among commuters travelling to high-demand central locations and during peak travel periods. Further elasticity analyses clarify behavioural patterns, identifying traveller groups with varying degrees of flexibility. Policy analyses underscore the effectiveness of targeted, dynamic tolling, particularly cordon-based pricing combined with time-specific toll adjustments, in reducing congestion levels. Additionally, the findings highlight the necessity of complementary measures, including improved transit services and targeted discounts, to ensure equitable outcomes. The findings offer targeted insights into how specific pricing strategies such as cordon, distance, and travel time-based tolls can be used to influence travel behaviour, reduce peak-period congestion, and guide equitable policy design in urban transportation planning. 2026-06-29T04:00:45Z Mohammad Amin Ashena Adam Weiss Jason Hawkins Lina Kattan