https://arxiv.org/api/HUfCgUaUcim0yEtua0Heoyx9f5o 2026-09-11T19:06:26Z 37740 15 15 http://arxiv.org/abs/2609.11173v1 Hierarchical Clustering Can Jointly Satisfy Richness, Consistency, and Scale Invariance 2026-09-10T07:20:10Z Despite its ubiquity, clustering lacks a universally accepted definition of what is a cluster. Kleinberg's Impossibility Theorem formalizes this difficulty by showing that no flat clustering method can simultaneously satisfy three natural axioms: scale invariance, richness, and consistency. In this paper, we ask whether this impossibility persists when the output is a hierarchy rather than a single partition. We show that, in contrast to the flat clustering setting, the hierarchical analog of these axioms are jointly satisfiable. In fact, there exist uncountably many hierarchical clustering methods satisfying these axioms, which we call admissible. We explicitly construct several admissible methods, including methods based on well-separated clusters and a non-binary version of single linkage. For certain pairs of admissible methods, the hierarchy produced by one always refines that produced by the other. This refinement relation defines a partial order on the class of admissible methods. This partially ordered set has no greatest element and contains uncountably many pairwise incompatible maximal elements, revealing substantial diversity among admissible methods. Nevertheless, this diversity is constrained: every admissible method contains a hierarchy of sufficiently well-separated clusters, and every finite collection of admissible methods shares such a nontrivial common backbone. 2026-09-10T07:20:10Z 51 pages, 3 figures Daichi Kuroda Maximilien Dreveton Matthias Grossglauser Patrick Thiran http://arxiv.org/abs/2609.11136v1 Beyond Tweedie's Formula: Conditional Score Modeling for Empirical Bayes Inference 2026-09-10T06:27:27Z We propose conditional f-modeling (Cf-modeling), a framework for empirical Bayes inference with covariates. A central identity shows that the conditional marginal score function determines not only the posterior mean through Tweedie's formula, but also the posterior moment-generating function, providing a basis for recovering posterior quantities without explicit prior modeling. Motivated by this observation, we treat the conditional marginal score as the primary object of inference and estimate it directly using an energy-based representation and Hyvärinen score matching, thereby avoiding potentially intractable covariate-dependent normalizing constants. The resulting framework flexibly accommodates covariate effects and heteroscedasticity and provides a practical approach to posterior moment estimation and uncertainty quantification. We demonstrate the effectiveness of the proposed method through simulations and an RNA-seq application. 2026-09-10T06:27:27Z 28 pages (main) + 6 pages (supplement) Shonosuke Sugasawa Zhigen Zhao http://arxiv.org/abs/2508.08324v3 Consistent Bayesian Spatial Domain Partitioning Using Predictive Spanning Tree Methods 2026-09-10T06:07:05Z Bayesian model-based spatial clustering methods are widely used for their flexibility in estimating latent clusters with an unknown number of clusters while accounting for spatial proximity. Many existing methods are designed for clustering finite spatial units, limiting their ability to make predictions, or may impose restrictive geometric constraints on the shapes of subregions. Furthermore, the posterior clustering consistency theory of spatial clustering models remains largely unexplored in the literature. In this study, we propose a Spatial Domain Random Partition Model (Spat-RPM) and demonstrate its application for spatially clustered regression, which extends spanning tree-based Bayesian spatial clustering by partitioning the spatial domain into disjoint blocks and using spanning tree cuts to induce contiguous domain partitions. Under an infill-domain asymptotic framework, we introduce a new distance metric to study the posterior concentration of domain partitions. We show that Spat-RPM achieves a consistent estimation of domain partitions, including the number of clusters (which may go to infinity), and derive posterior concentration rates for partition, parameter, and prediction. We also establish conditions on the hyperparameters to achieve consistency, offering important practical guidance for hyperparameter selection. Finally, we examine the asymptotic properties of our model through simulation studies and apply it to Atlantic Ocean data. 2025-08-09T16:36:44Z Kun Huang Huiyan Sang http://arxiv.org/abs/2609.11091v1 Is the Linear Threshold Good Enough? A Scale-Free Parameter and Adequacy Test for Curvature-Induced Threshold Displacement 2026-09-10T05:02:21Z Applied work often locates a threshold by linearizing a smooth function about a reference point and solving for the crossing. When the function is curved, the linear crossing can be substantively displaced even when standard errors are valid. We introduce the curvature-overstatement parameter \(Θ_{COT}=\log(|h_2^*|/|h_1^*|)\), the log ratio of second- to first-order threshold displacement. On the quadratic branch continuous with the linear solution, \(|h_2^*/h_1^*|=2/(1+\sqrt{1-u})\), where \(u=2qa/b^2\) is a dimensionless index formed from the local gap, slope, and curvature. Thus \(Θ_{COT}\) is scale-free, depends on the local parameters through one scalar, and has regular-domain range \((-\infty,\log 2)\), with boundary limit \(\log 2\) at tangency. We derive regular asymptotic inference, characterize local-to-tangency and weak-slope failures, and give a remainder bound linking the second-order crossing to the true threshold. The main practical contribution is an adequacy test that can affirm that the linear threshold is accurate within a prespecified proportional tolerance, rather than treating failure to detect curvature as evidence of adequacy. Monte Carlo results confirm regular-case calibration, the predicted nonstandard behavior near tangency, and failure under weak slope, while bootstrap diagnostics identify regimes in which regular inference should not be used. COT therefore provides an effect-size scale, adequacy test, and diagnostics for deciding whether a first-order threshold is accurate enough to report. 2026-09-10T05:02:21Z 29 pages,4 tables, 2 figures Subir Hait http://arxiv.org/abs/2511.18432v2 Change-Point Detection With Multivariate Repeated Measures 2026-09-10T04:14:07Z Graph-based methods have shown particular strengths in change-point detection (CPD) tasks for high-dimensional nonparametric settings. However, existing CPD research has rarely addressed data with repeated measurements or local group structures. A common treatment is to average repeated measurements, which can result in the loss of important within-individual information. In this paper, we propose a new graph-based method for detecting change-points in data with repeated measurements or local structures by incorporating both within-individual and between-individual information. Analytical approximations to the significance of the proposed statistics are derived, enabling efficient computation of p-values for the combined test statistic. We also establish consistency of the proposed test and the estimated change-point location. The proposed method effectively detects change-points across a wide range of alternatives, particularly when within-individual differences are present. The new method is illustrated through an analysis of the New York City taxi dataset. 2025-11-23T12:53:44Z Serim Han Jingru Zhang Hoseung Song http://arxiv.org/abs/2609.11032v1 Bayesian Variable Selection for High-Dimensional Predictors with Missing Psychometric Outcomes 2026-09-10T03:27:54Z High-dimensional, multimodal predictors and partially observed multivariate outcomes are common in psychometric research. However, existing regularization methods often do not accommodate hierarchical predictor structures and are primarily designed for univariate outcomes. We propose SHIM, a Bayesian framework for structured variable selection that combines hierarchical horseshoe shrinkage with a Bayesian treatment of missing outcomes. The framework jointly accommodates predictor hierarchies, dependence among outcomes, and incomplete multivariate responses. We establish theoretical properties of the proposed prior specification and evaluate SHIM through simulation studies. The results demonstrate that SHIM balances sensitivity with false-positive control while yielding accurate coefficient estimates and well-calibrated uncertainty quantification. We further apply SHIM to data from an Alzheimer's disease cohort to characterize associations between multimodal neuroimaging measures and multivariate neuropsychological outcomes and to generate posterior-based multiple imputations for downstream analyses of the relationships between fluid biomarkers and cognition. An R package, shim, is publicly available to facilitate implementation. 2026-09-10T03:27:54Z Zongyue Teng Shujie Ma Timothy J. Hohman Angela L. Jefferson Panpan Zhang http://arxiv.org/abs/2609.11012v1 Discretization in covariate-adaptive randomization: gains and losses 2026-09-10T02:44:11Z Covariate-adaptive randomization(CAR) is widely implemented in clinical trials to balance prognostic covariates across treatment arms. Continuous covariates are often discretized into strata in practice, yet their consequences are not clearly understood. This paper provides a comprehensive study of the impact of discretization on both the CAR design process and the inferential results thereafter. We establish the asymptotic properties of both imbalance measures and treatment effect estimators under discretized and non-discretized settings. Practical recommendations are given on when and how discretization should be employed. We show that discretization in design is generally recommended, as it enhances robustness against model misspecification. However, if the true model is known, the most efficient strategy is to balance covariates according to that model in the design. The theoretical results are corroborated by extensive simulation studies and an empirical application to a diabetes trial dataset. Together, the results clarify the gains and losses of discretization in CAR and pave the way for learning impact of discretization to other designs and beyond. 2026-09-10T02:44:11Z Zixuan Zhao Feifang Hu http://arxiv.org/abs/2609.09471v2 Covariate-localized False Discovery Rates 2026-09-10T01:30:29Z We introduce a flexible model for covariate-dependent multiple testing which can be encoded using a nonparametric Gaussian mixture model. Weight-localized predictive recursion (PRx), a new development in the methodology of Newton's predictive recursion algorithm, is then leveraged to estimate the components of this mixture model, allowing for recovery of the covariate-localized false discovery rate $\text{Pr}(H_i = 0|z_i,x_i)$ using a single, unified algorithm. This quantity represents the most direct extension of Efron's local false discovery rate to the covariate-dependent setting, and admits provable Bayesian FDR control properties under simple rejection rules. We introduce several procedures for estimating and thresholding the local false discovery rate, and show using various simulations and a real-data example that our procedures lead to increased power, tighter Bayesian FDR control, and more interpretable rejections. We furthermore show that this holds for fixed and randomized hypothesis labels, indicating that our proposed methods perform well under both frequentist and Bayesian interpretations of multiple testing. 2026-09-08T21:39:57Z Jonathan Lin Surya Tokdar http://arxiv.org/abs/2608.18365v2 Posterior Convergence without Force Convergence: Resolution-Stable Sampling for Rough Bayesian Inverse Problems 2026-09-10T01:19:59Z Bayesian targets may converge under model refinement even when the exact sensitivities used by gradient-based samplers do not. We study this probability--sensitivity mismatch and its consequences for Metropolized Hamiltonian proposals. A vanishing-amplitude wiggly-energy model first gives the basic analytic obstruction: the potential perturbation tends to zero while its classical derivative is of order $r_\varepsilon/\varepsilon$. We then show that the same scaling arises naturally in a periodic elliptic inverse problem, where homogenization makes the forward map and Gaussian likelihood converge while differentiation with respect to a microscopic scale parameter retains an $O(1)$ oscillatory contribution. This provides a PDE origin for the single-scale wiggly mechanism. The main construction concerns a more demanding nested Weierstrass hierarchy, interpreted as an analytically tractable prototype for repeated corrector contributions across geometrically separated scales. There all previously resolved scales persist, adjacent classical-force increments grow geometrically like $(ab)^N$, and the limiting rough component may fail to possess a classical derivative. In this self-similar setting the matched Jackson quotient is structurally adapted to the refinement through dilation covariance and exact finite closure. Uniform negative-log-likelihood approximation yields explicit total-variation, Hellinger, and bounded quantity-of-interest bounds. Measurable kick--drift--kick maps remain exact after Metropolis correction, local field convergence propagates to fixed-length proposals and kernels, and the first classical HMC half-kick can have no fixed-step refinement limit. Numerical experiments on scale-structured inverse problems test the resulting resolution-stability mechanism across one- and two-dimensional inverse problems. 2026-08-18T22:35:17Z Zhiliang Deng Xiaomei Yang http://arxiv.org/abs/2609.10933v1 Generalized Ridge Refitting for the Lasso and Prediction Improvement Bounds 2026-09-10T00:39:44Z We study a class of Lasso based estimators obtained by applying a quadratic correction on the Lasso equicorrelation set. The penalty matrix determines both the magnitude and geometry of the correction and contains, among other cases, the isotropic Lasso--Ridge correction, least squares refitting, Gram proportional interpolation between the Lasso and least squares, and coordinate specific penalties. We first derive a closed form representation and isolate the positive gain component of the resulting prediction improvement. We then control the remaining stochastic linear term in expectation by localizing the random signed equicorrelation model around a deterministic reference support. This yields a finite sample expectation bound that explicitly accounts for the randomness induced by Lasso model selection. The resulting decomposition provides a unified framework for understanding when Lasso based quadratic corrections can improve prediction. 2026-09-10T00:39:44Z 28 pages Guo Liu Waseda University http://arxiv.org/abs/2311.12252v4 Partial identification and unmeasured confounding with multiple treatments and multiple outcomes 2026-09-10T00:32:07Z Estimating the health effects of multiple air pollutants is a crucial problem in public health, but one that is difficult due to unmeasured confounding bias. Motivated by this issue, we develop a framework for partial identification of causal effects in the presence of unmeasured confounding in settings with multiple treatments and multiple outcomes. Under a factor confounding assumption, we show that joint partial identification regions for multiple estimands can be more informative than considering partial identification for individual estimands one at a time. We show how assumptions related to the strength of confounding or magnitude of plausible effect sizes for one estimand can reduce the partial identification regions for other estimands. As a special case of this result, we explore how negative control assumptions reduce partial identification regions and discuss conditions under which point identification can be obtained. We develop novel computational approaches to finding partial identification regions under a variety of these assumptions. We then estimate the causal effect of PM$_{2.5}$ components on a variety of public health outcomes in the United States Medicare cohort, where we find that the detrimental effect of certain air pollutants are robust to the potential presence of unmeasured confounding bias. 2023-11-21T00:19:25Z Suyeon Kang Alexander Franks Heejun Shin Michelle Audirac Danielle Braun Joseph Antonelli http://arxiv.org/abs/2609.10927v1 Surprise Reduction and Nullification in Bayesian and Inverse Bayesian Inference under Ambiguous Prediction-Error Attribution 2026-09-10T00:20:41Z In non-stationary environments, prediction errors may signal environmental change or transient outliers, and adaptive systems must track such changes without overreacting to outliers. We distinguish surprise reduction, which updates beliefs to fit observations, from surprise nullification, which weakens constraints imposed by the predictive structure, and formalize both within Bayesian and inverse Bayesian (BIB) inference. Belief and likelihood updates are derived from variational objectives sharing a nullification strength, determined endogenously by minimizing surprise under the candidate post-update predictive distribution. In the Gaussian case, nullification expands belief and likelihood variances by a common factor relative to standard Bayesian updating, leaving the ratio unchanged. BIB thus defers attribution of the prediction error, committing to neither latent-state change nor observation-process uncertainty. The nullification strength is carried over as a candidate and is maintained or released according to the predictive surprise of the next observation. In a mean estimation task with outliers and changepoints, no scanned parameter setting of a Sage-Husa-type adaptive Kalman filter, fixed-strength BIB variant, or belief-forgetting-only variant outperforms BIB in both changepoint tracking and post-outlier stability. An oracle-informed reduced Bayesian model tracks changepoints better but is less stable after outliers. Although BIB maintains no explicit hypotheses about changepoints or outliers, it generates event-dependent dynamics. The learning rate increases after changepoints, whereas after outliers, nullification is released, and this increase is suppressed. Deferring attribution and letting subsequent observations differentiate the responses may constitute a principle of adaptive inference in non-stationary environments. 2026-09-10T00:20:41Z Shuji Shinohara Daiki Morita Yoshihiro Nakajima Takeshi Takano Masakazu Higuchi Ung-il Chunge Yukio-Pegio Gunji http://arxiv.org/abs/2609.10904v1 Agnostic Model-Assisted Estimation with Machine Learning for Survey Data 2026-09-09T23:16:21Z Model-assisted estimation uses prediction rules to improve the efficiency of estimators of finite population parameters while retaining design-based inference. Although flexible prediction methods have been considered, existing theoretical results are largely method-specific. We develop a learner-agnostic framework that replaces separate analyses for individual learners with general conditions on the sampling design and prediction error. We connect design-aware and design-agnostic cross-fitting and characterize the sampling designs under which they yield conditional independence across folds. Under suitable conditions, conditional weighting gives exact design-unbiasedness. We establish first-order equivalence to oracle estimators, leading to design consistency and asymptotic normality, and clarify when conditional and original inclusion probabilities yield the same first-order behavior. We propose consistent variance estimators based on cross-fitted residuals and construct asymptotically valid confidence intervals. Under additional model and regularity conditions, we establish asymptotic optimality through attainment of the Godambe--Joshi lower bound. Simulations show that cross-fitting substantially reduces finite-sample bias and improves variance estimation and coverage with adaptive learners. 2026-09-09T23:16:21Z Ziming An Mehdi Dagdoug David Haziza Yves Tillé http://arxiv.org/abs/2410.02727v6 Regression Discontinuity Designs Under Interference 2026-09-09T22:17:20Z We extend the continuity-based framework to Regression Discontinuity Designs (RDDs) to identify and estimate causal effects under interference when units are connected through a network. Assignment to an "effective treatment," combining the individual treatment and a summary of neighbors' treatments, is determined by the unit's score and those of interfering units, yielding a multiscore RDD with complex, multidimensional boundaries. We characterize these boundaries and derive assumptions to identify boundary causal effects. We develop a distance-based nonparametric estimator and establish its asymptotic properties under restrictions on the network degree distribution. We show that while direct effects converge at the standard rate, the rate for indirect effects depends on the number of scores fixed at the cutoff. Finally, we propose a variance estimator accounting for network correlation and apply our method to PROGRESA data to estimate the direct and indirect effects of cash transfers on school attendance. 2024-10-03T17:48:18Z Elena Dal Torrione Tiziano Arduini Laura Forastiere http://arxiv.org/abs/2609.10865v1 The Elliptically Optimal Confidence Interval: A Bivariate Extension of Wilson's Score Method 2026-09-09T22:06:19Z Constructing a confidence interval for the difference between two independent binomial proportions involves a nuisance direction that is not identified by the estimand. The one-sample Wilson score interval inverts a scalar score test, but has no direct bivariate analogue isolating the difference: inverting the joint normal approximation yields an elliptical region in the unit square, whereas the estimand \(p_1-p_2\) is one-dimensional. We define the Elliptically Optimal (EO) confidence interval as the range of \(p_1-p_2\) over this region and solve the resulting optimization problem in closed form, obtaining explicit bounds in six mutually exclusive and exhaustive cases. The solution admits a compact characterization: the EO interval is the score interval obtained by maximizing over the nuisance variance rather than estimating it. It is therefore the shortest interval obtained by projecting the elliptical region, and inherits its coverage guarantee. We derive the exact coverage excess, \(2[Φ(z\mathcal{R})-Φ(z)]\), where \(\mathcal{R}\) is the ratio of the least-favourable to the true standard deviation. The excess vanishes on an explicit line through the parameter space, is bounded by \(α\), and is invariant under proportional scaling of the sample sizes. Exact enumeration of the binomial coverage shows that the Wald interval, whose variance estimator is downward biased by a factor \(1-1/n\) under balanced allocation, falls below nominal coverage almost everywhere. The EO interval never under-covers under the normal approximation and always yields admissible, non-degenerate bounds. Its price is over-coverage when both proportions are extreme, which we quantify exactly. 2026-09-09T22:06:19Z Nawaf Mohammed