https://arxiv.org/api/nP/tfEnT+of//XBbndd383JuPM42026-09-11T21:03:59Z377406015http://arxiv.org/abs/2609.09516v1Differential Privacy Guarantees in Small Area Estimation2026-09-08T22:57:14ZStatistical agencies increasingly rely on small area estimation to produce reliable estimates for subpopulations with limited sample sizes. These estimates are built from individual survey responses, so agencies must ensure that releasing them does not reveal information about any single respondent. We show that when a single draw from the posterior distribution of the Bayesian Fay-Herriot model is released, pure $\varepsilon$-differential privacy is unattainable, but the release satisfies formal privacy guarantees under Rényi differential privacy and zero-concentrated differential privacy without any noise being added, provided we treat the variance components as fixed. The key insight is that the posterior draw equals the posterior mean plus the Gaussian noise whose variance equals the posterior variance. The guarantee is thus governed by the sensitivity of the direct survey estimate and the posterior variance, and applies equally to a release of the posterior mean with that amount of noise added. For binary outcomes estimated with the Hájek estimator, the sensitivity equals the largest survey weight in the area divided by the sum of the weights. For the intercept-only model we derive exact coefficients describing how a change in one record propagates to every area's posterior mean, giving finite-sample per-area guarantees and a joint guarantee for releasing all areas at once that exceeds the largest per-area guarantee by at most a few percent in our applications. Two applications, poverty prevalence across 2,462 Public Use Microdata Areas in the American Community Survey and smoking prevalence across 52 substrata in the Washington state Behavioral Risk Factor Surveillance System, show that the guarantee is driven far more by the inequality of the survey weights than by the sample size, and that the shrinkage of the model tightens it substantially.2026-09-08T22:57:14ZSoumojit DasJörg Drechslerhttp://arxiv.org/abs/2603.04199v2Bayesian Adversarial Privacy2026-09-08T22:33:11ZTheoretical and applied research into privacy encompasses an incredibly broad swathe of differing approaches, emphases and aims. This work introduces a novel quantitative notion of privacy that is both contextual and specific. Building on and extending ideas from statistical disclosure control and differential privacy, our aim is to model the implications of a disclosure decision in an adversarial setting. Our definition relies on concepts inherent to standard Bayesian decision theory, while departing from them in several important respects. In particular, (i) inference about the data itself becomes meaningful and (ii) the party controlling the release of sensitive information should make disclosure decisions from the prior viewpoint, rather than conditional on the data, which is a feature shared with Bayesian design. Illuminating toy examples are exploited towards highlighting the specificities of the method.2026-03-04T15:46:24ZCameron BellTimothy JohnstonAntoine LucianoChristian P Roberthttp://arxiv.org/abs/2512.13473v2Parsimonious Ultrametric Manly Mixture Models2026-09-08T21:49:50ZA family of parsimonious ultrametric mixture models with the Manly transformation is developed for clustering high-dimensional data where the clusters may be asymmetric. While advances in Gaussian mixture modeling sufficiently handle high-dimensional data, they often struggle with the common presence of cluster skewness. To address this, we incorporate the extended ultrametric covariance structure and the Manly transformation, resulting in the parsimonious ultrametric Manly mixture model family. The ultrametric covariance structure reduces the number of free parameters while identifying latent groups of variables within a nested hierarchy. This phenomenon enables the visualization of hierarchical relationships within clusters, improving cluster interpretability. Additionally, as with many classes of mixture models, model selection remains a fundamental challenge; to this end, a two-step model selection procedure is proposed herein. Through simulation studies and real data analyses, we demonstrate improved model selection via the proposed two-step method, as well as the effective clustering performance.2025-12-15T16:09:32ZAlexa A. SochaniwskyPaul D. McNicholashttp://arxiv.org/abs/2501.10886v3Linear scaling causal discovery from high-dimensional time series by dynamical community detection2026-09-08T21:35:48ZUnderstanding which parts of a dynamical system cause each other is extremely relevant in fundamental and applied sciences. However, inferring causal links from observational data, namely without direct manipulations of the system, is still computationally challenging, especially if the data are high-dimensional. In this study we introduce a framework for constructing causal graphs from high-dimensional time series, whose computational cost scales linearly with the number of variables. The approach is based on the automatic identification of dynamical communities, groups of variables which mutually influence each other and can therefore be described as a single node in a causal graph. These communities are efficiently identified by optimizing the Information Imbalance, a statistical quantity that assigns a weight to each putative causal variable based on its information content relative to a target variable. The communities are then ordered starting from the fully autonomous ones, whose evolution is independent from all the others, to those that are progressively dependent on other communities, building in this manner a community causal graph. We demonstrate the computational efficiency and the accuracy of our approach on time-discrete and time-continuous dynamical systems including up to 80 variables.2025-01-18T21:50:43ZMatteo AllioneVittorio Del TattoAlessandro Laio10.1103/kd73-93cghttp://arxiv.org/abs/2609.09436v1MiNCE: Nonparametric, Strongly Consistent Confidence Envelopes for Band-Limited Functions and their Smoothed Spectra2026-09-08T20:37:27ZMinimum-norm confidence envelope strategies offer a nonparametric approach to constructing nonasymptotic, simultaneous confidence regions for band-limited functions, exploiting the theory of Reproducing Kernel Hilbert Spaces (RKHS). While the finite-sample coverage guarantees of these envelopes have been established, their consistency has not been analyzed so far. In this paper, we study this construction, here termed the Minimum-Norm Confidence Envelope (MiNCE) framework, and establish the strong uniform consistency of the resulting bands, both for noise-free and noisy observation models, under mild assumptions on the measurement noises. We further extend this formulation to the frequency domain, deriving nonasymptotic, simultaneous, strongly uniformly consistent confidence bands for the smoothed spectra. Numerical experiments in nonparametric regression and spectral estimation empirically confirm our theoretical results, illustrating the contraction of the confidence envelopes toward the target function as the sample size increases.2026-09-08T20:37:27ZBalázs Csanád CsájiBálint Horváthhttp://arxiv.org/abs/2609.09401v1Semiparametric Receiver Operating Characteristic Analysis in the Presence of an Imperfect Reference Standard via a Box-Cox Density Ratio Model2026-09-08T19:57:16ZReceiver operating characteristic (ROC) analysis is commonly used to evaluate the diagnostic accuracy of continuous biomarkers. In practice, the true disease status may be unavailable and only a nominal disease status provided by an imperfect reference standard is observed. Existing nonparametric methods have been developed for ROC analysis in this setting, but may suffer from reduced estimation efficiency, numerical instability, or sensitivity to the choice of biomarker scale. We propose a semiparametric method based on a Box-Cox density ratio model, which links the biomarker distributions of the truly healthy and diseased populations while leaving the baseline distribution unspecified. A key feature of the proposed method is that the transformation parameter is estimated from the data rather than specified in advance, allowing the density-ratio structure to adapt to different transformation scales. We develop an empirical likelihood approach for estimation and an expectation-maximization algorithm for computation. We establish the asymptotic distributions of estimators of the ROC curve, area under the curve, Youden's index, and the sensitivity and specificity at the Youden-optimal cutoff, and develop bootstrap confidence intervals and a goodness-of-fit test. Simulation studies demonstrate that the proposed method provides accurate and numerically stable estimation and inference across a range of distributional settings without requiring the transformation scale to be specified in advance. The proposed method is illustrated using data from a malaria study.2026-09-08T19:57:16ZYi ChangSiyan LiuQinglong TianPengfei Lihttp://arxiv.org/abs/2609.09389v1A Counterfactual Framework for Estimating Infectious Disease Prevalence under Repeated Testing with Symptomatic and Contact-Tracing Components2026-09-08T19:40:48ZThis paper addresses the problem of estimating infectious disease prevalence under longitudinal testing programs that include scheduled, symptomatic, and contact-tracing testing. Our study is motivated by data from The Ohio State University, where a mandatory once-per-week COVID-19 testing and isolation program was implemented during the Fall 2020 semester, supplemented by additional testing for symptomatic individuals and identified contacts. In this setting, the probability of being tested depends on symptoms or contact-tracing status, creating a complex observation process. We develop a counterfactual framework that links the observation process to a hypothetical process in which infection is prevented. This formulation enables unbiased estimation of disease prevalence by modeling the testing process, possibly nonparametrically, without requiring explicit modeling of transmission dynamics, even though the testing and infection processes are jointly dependent.2026-09-08T19:40:48ZAccepted for publication in The Annals of Applied StatisticsJeongjin LeeJunke YangGrzegorz A. RempalaPatrick M. Schnellhttp://arxiv.org/abs/2609.09377v1Estimating Causal Treatment Effects in Placebo-Controlled Randomized Clinical Trials When High Placebo Response is Anticipated2026-09-08T19:21:30ZIn placebo-controlled randomized clinical trials (RCTs), the placebo response significantly modifies treatment effects and diminishes the intention-to-treat (ITT) treatment effect, $Δ_{ITT}$. This study presents a novel two-stage framework for estimating the standardized causal treatment effect, $Δ_{STD}$, among the ITT population, under the assumption that their placebo responses are similar to the levels of self-administering medication at home. The first stage employs a real-world, pragmatic, single-blinded placebo lead-in to measure placebo responses to levels expected during routine at-home use. This is achieved by preserving the participants' expectations and controlling for trial-related factors that inflate the responses. In the second stage, a double-blinded randomized phase is used to estimate the conditional average treatment effect (CATE) as a function of placebo response levels and other important effect modifiers. To facilitate CATE estimation, the prognostic scores, defined as the expected placebo responses, are used for dimension reduction. The causal estimand $Δ_{STD}$ is computed by integrating the CATE function over the distribution of the expected placebo response levels from stage one and other modifiers. We further derive theoretical values for $Δ_{ITT}-Δ_{STD}$ to quantify the underestimated treatment benefit due to high placebo responses. The validity and statistical performance of the proposed framework are evaluated through comprehensive simulations.2026-09-08T19:21:30ZYang SongYuezhe QianChanmin KimGheorghe Doroshttp://arxiv.org/abs/2603.25806v3Context Tree Prior Distributions based on Node Weighting with exact Bayes Factors2026-09-08T18:31:24ZVariable-length Markov chains (VLMCs) are a flexible class of higher-order Markov models that admit a natural representation as context trees. Existing Bayesian methods for specifying prior distributions on trees rely on branching processes, but these suffer from a fundamental limitation: the connection between node-branching probabilities and the structural properties of the induced tree distribution is not straightforward, making it difficult to encode specific structural beliefs. We address this issue by introducing a novel representation of prior distributions on tree spaces, characterized by assigning weights to individual contexts through a function on nodes. In this way, our approach provides an intuitive mechanism for incorporating structural hypotheses into the prior while preserving computational tractability, allowing marginal likelihoods and posterior mode trees to be computed exactly via generalizations of the Context Tree Weighting (CTW) and Context Tree Maximizing (CTM) algorithms. By enabling exact Bayes factor calculations, our methodology provides a principled framework for model comparison over structural priors. We demonstrate the flexibility and effectiveness of our approach by comparing different prior specifications through simulation studies and an application to financial markets.2026-03-26T18:12:06Z35 pages, 6 figuresThiago PaulichenVictor Fregugliahttp://arxiv.org/abs/2609.09326v1Real-time and adaptive anomaly detection algorithm for cyclostationary models2026-09-08T18:11:57ZThis article introduces PeriodicCALM, an effective real-time anomaly detection framework designed for cyclostationary data streams. While classical cyclostationary processes feature periodically time-varying statistical properties, real-world signals often contain recurring impulsive components that conceal abnormal behavior. Existing real-time methods for struggle with these dynamics, frequently misinterpreting phase-dependent variability as non-cyclic anomalies and causing excessive false alarms. To address this, PeriodicCALM incorporates cycle-dependent variability to systematically ignore regular cyclic impulses while accurately isolating genuine anomalies. Operating in real time with continuous retraining capabilities, the method adapts dynamically to evolving signal characteristics. Comparative evaluations against the baseline CALM framework using simulated data demonstrate significant improvements in detection accuracy and training efficiency, alongside a reduction in prediction latency. Furthermore, the practical utility of PeriodicCALM is validated on real-world vibration signals collected from a compressor monitoring system.2026-09-08T18:11:57Z16 pagesJustyna WitulskaTomasz BarszczIreneusz JabłońskiAgnieszka Wyłomańskahttp://arxiv.org/abs/2609.09325v1Kalman Filtering and Smoothing for Improving Precision in Horvitz--Thompson Estimation of Infectious Disease Prevalence2026-09-08T18:10:48ZHorvitz--Thompson (HT) estimators can provide unbiased daily estimates of infectious disease prevalence under repeated surveillance by correcting for nonrandom testing induced by scheduled, symptom-based, and contact-tracing components. However, because each HT estimate is based on the testing data available for that day and may involve highly variable inverse probability weights, it can be noisy, have precision that varies over time, and become unavailable during temporary interruptions in testing. The daily HT estimator is modeled as a noisy observation of an underlying prevalence process, with day-specific observation variances estimated using a delete-a-group jackknife. Our primary specification is a joint local linear trend state-space model that extends the standard level-only random walk by adding a latent slope. The Kalman filter improves precision by borrowing information from past estimates. At each time $t$, the process variances are estimated or carried forward using only observations available through time $t$, so the resulting filtered estimate is available in real time. We also describe the corresponding Kalman smoother as a retrospective extension based on the full observed series. When daily HT estimates are missing, the Kalman filter proceeds through prediction-only updates, whereas the corresponding smoother retrospectively reconstructs those periods using later observations. In simulations, the joint Kalman filter substantially improves precision relative to the raw daily HT estimator while preserving the main temporal pattern, and the smoother provides a more stable retrospective summary. In The Ohio State University's fall 2020 SARS-CoV-2 surveillance data, the filter and smoother provide estimates on no-testing days, when the HT estimator provides neither point nor interval estimates, and yield narrower confidence intervals than HT intervals on observed days.2026-09-08T18:10:48ZAbstract shortened to comply with arXiv's 1,920-character limitJeongjin LeeGrzegorz A. RempalaPatrick M. Schnellhttp://arxiv.org/abs/2609.09089v1Sieve Estimation of Optimal Transport Maps from Paired Data in Gaussian Spaces2026-09-08T17:34:55ZWe estimate optimal transport maps on an infinite-dimensional Hilbert space with a Gaussian reference measure, from noisy paired observations. A source draw is seen together with a noisy evaluation of its image, rather than through independent unpaired samples. The estimator is a cylindrical sieve of Cameron--Martin gradient maps, restricted to a compact parameter set; individual sieve elements need not be transport maps. It yields a finite regression contrast even though the noise has infinite Cameron--Martin norm, and reduces estimation to finite-dimensional empirical risk minimization. We establish a nonasymptotic oracle inequality separating approximation error, stochastic error and the local conditioning of the parametrization, together with a minimax lower bound of order $N^{-s/(2s+1)}$ under weighted coordinate regularity of order $s$. Output regularity alone does not deliver cylindrical approximation; for general Sobolev potentials an input-regularity index does, via a conditional Gaussian Poincaré argument, and for potentials of bounded chaos degree the degree bound plays that role, an orthogonal Hermite sieve then attaining the same rate with the interaction order replacing unity in the exponent. That rate is minimax on a diagonal Gaussian class and on a nonlinear block class whose interaction survives every fixed orthogonal change of coordinates.2026-09-08T17:34:55ZXin JinKit ChanRiddhi Pratim Ghoshhttp://arxiv.org/abs/2607.06330v2Estimation of Linear Functionals in Multilayer Panels under Staggered Adoption2026-09-08T17:18:24ZWe study the estimation of bilinear forms from noisy, partially observed multilayer data. The signal follows a Tucker2 model, with shared unit and time factors across tensor layers and slice-specific cores. The missingness pattern is structured and motivated by staggered adoption designs, which are common in causal inference and related applications. We first analyze the four-block missingness pattern, the basic building block for general staggered adoption, and propose a spectral algorithm that pools information across layers and targets the functional directly. We prove a non-asymptotic mean squared error bound that exhibits a phase transition in the number of layers, showing when pooling improves estimation, and match it with a local minimax lower bound up to constants when ranks and logarithmic factors are treated as constant-order quantities. We then extend the construction to general staggered adoption designs via an anchored four-block reduction, and derive analogous theoretical guarantees. Finally, we validate our theoretical findings using synthetic and real-world data on Castle Doctrine laws and COVID-19 policies2026-07-07T14:26:36Z108 pages, 12 figures, 3 tablesAlberto BordinoThomas B. BerrettOlga Klopphttp://arxiv.org/abs/2609.09053v1Quadratic Point Estimate Method for Uncertainty Quantification with Dependent Non-Gaussian Inputs2026-09-08T17:10:28ZAs an extension of the Point Estimate Method (PEM) to evaluate probabilistic moments of quantities of interest (QoI) in general $n$-dimensional spaces, the Quadratic Point Estimate Method (QPEM) has been recently developed. This new method is defined to fully represent up to fifth-order input moments in the Gaussian space, providing general analytical expressions for sample locations and weights, without requiring any numerical optimization. The QPEM can significantly improve the estimation accuracy of the output QoI moments, in relation to PEM-based methods whose numbers of sigma points grow linearly with the problem dimension, while at the same time having an affordable and competitive computational cost up to a considerable number of dimensions. The QPEM is further enhanced in this work by enabling copula integration into the framework, which enables effective modeling of the joint input probability density function by estimating marginals and the dependence structure of the involved random variables. The validity and efficient performance of the copula-based QPEM are showcased against numerous other sampling methods in various examples considering two practical scenarios: (i) when the joint dependence structure can be inferred from data, and (ii) when only marginal distributions and correlation matrices are known.2026-09-08T17:10:28ZMinhyeok KoKonstantinos G. Papakonstantinouhttp://arxiv.org/abs/2609.09039v1Covariate Adjustment in Randomized Experiments: A Unified Framework for Decision and Practice2026-09-08T17:03:57ZShould researchers adjust for covariates in randomized experiments, and if so, how? The literature offers three distinct prescriptions: do not adjust because randomization guarantees unbiasedness; adjust for outcome-prognostic covariates to improve precision; or adjust for covariates imbalanced between treatment arms. These competing prescriptions create confusion and uncertainty. We develop a unified framework for decision and practice. Given available information, we show that the optimal correction is what we call ex-post bias. The only relevant criterion for adjustment is prognosticity for ex-post bias; neither raw covariate imbalance nor outcome prognosticity is sufficient by itself. We also show that correcting imbalance and improving precision are two sides of the same decision problem. We develop two estimation approaches, one of which recovers familiar adjustment estimators and provides a new theoretical justification for them. Simulations compare alternative covariate-selection and adjustment strategies. Overall, our framework provides a unified foundation for covariate adjustment in randomized experiments.2026-09-08T17:03:57ZJiawei FuDonald P. Green