https://arxiv.org/api/5TrmVOLKs7O/W9/owGiSZ49hExs2026-10-02T21:11:37Z1067527015http://arxiv.org/abs/2608.20601v1A Comprehensive Bayesian Approach to Entity Resolution for Data with Multiple Truths2026-08-20T22:41:06ZIn many applications, from government to ecology, integrating data from diverse and noisy sources is critical for downstream inference. However, a unique identifier to link records cleanly from the same entity may not exist. Entity resolution (also referred to as de-duplication or record linkage) merges such databases to identify duplicates, allowing for more complete data for inference. A multitude of methods from statistics and computer science in recent years assume a single, immutable true value for each variable used in linkage. We argue this restrictive assumption is often violated in practice, leading researchers to discard useful data and biasing downstream inference. For example, a respondent's education status may truly change between two different surveys, yet existing methods would treat at least one observation as a distorted version of the truth, even though both are correct. In this paper, we propose a novel entity resolution model that comprehensively accommodates "multiple truths" by introducing a mixture of exponential family distributions that further handles multiple data types. We provide options to fit the model with Markov chain Monte Carlo routines and variational inference for massive datasets. We demonstrate the method's value via simulation and linking a longitudinal survey of Italian household wealth.2026-08-20T22:41:06Z74 pagesHyungjoon KimAndee KaplanMatthew D. Koslovskyhttp://arxiv.org/abs/2508.11461v2Importance Sampling Approximation of Sequence Evolution Models with Site-Dependence2026-08-20T19:46:31ZWe consider models for molecular sequence evolution in which the transition rates at each site depend on the local sequence context, giving rise to a time-inhomogeneous Markov process in which sites evolve under a complex dependency structure. We introduce a randomized approximation algorithm for the marginal sequence likelihood under these models using importance sampling, and provide matching order upper and lower bounds on the finite sample approximation error. Given two sequences of length $n$ with $r$ observed mutations, we show that for practical regimes of $r/n$, the complexity of the importance sampler does not grow exponentially in $n$, but rather in $r$, making the algorithm practical for many applied problems. We demonstrate the use of our techniques to obtain problem-specific complexity bounds for a well-known dependent-site model from the phylogenetics literature.2025-08-15T13:21:06ZJoseph MathewsScott C. Schmidlerhttp://arxiv.org/abs/2511.07736v2Improved Bounds for Context-Dependent Evolutionary Models Using Sequential Monte Carlo2026-08-20T19:34:39ZStatistical inference in evolutionary models with site-dependence is a long-standing challenge in phylogenetics and computational biology. We consider the problem of approximating marginal sequence likelihoods under dependent-site models of biological sequence evolution. We prove an upper bound on the mixing time for a Markov chain Monte Carlo algorithm that samples the conditional distribution over latent sample paths, when the chain is initialized with a warm start. We then introduce a sequential Monte Carlo (SMC) algorithm for approximating the marginal likelihood, and show that our mixing time bound can be combined with recent importance sampling and finite-sample SMC results to obtain bounds on the finite sample approximation error of the resulting estimator. Our results show that the proposed SMC algorithm yields an efficient randomized approximation scheme for many practical problems of interest, and offers a significant improvement over a recently developed importance sampler for this problem. Our approach combines recent innovations in obtaining bounds for MCMC and SMC samplers, and may prove applicable to other problems of approximating marginal likelihoods and Bayes factors.2025-11-11T01:40:40ZJoseph MathewsScott C. Schmidlerhttp://arxiv.org/abs/2609.27884v1Uniform efficiency of the Ulrich-Wood sampler for the von Mises-Fisher distribution2026-08-20T18:13:44ZThe standard approach to stochastic simulation from the von Mises-Fisher distribution is a rejection sampler proposed by Ulrich and Wood. This note provides a theoretical justification of its efficiency, lower-bounding its acceptance probability uniformly over the dimension and concentration parameter. A novel interpretation of the proposal is given which provides some intuition for this efficiency.2026-08-20T18:13:44Z4 pages, no figuresSam Powerhttp://arxiv.org/abs/2608.20475v1A Complexity Bound for the Kent-Ganeiber-Mardia Sampler for the Bingham Distribution2026-08-20T18:07:38ZThe Bingham distribution is a family of antipodally symmetric distributions on the unit sphere, characterised by an exponential-of-quadratic change of measure with respect to the uniform distribution. Kent, Ganeiber and Mardia proposed a rejection sampler for generating samples from Bingham distributions using proposals from an angular central Gaussian (ACG) distribution. Their empirical results suggest that the least efficient regime is the high-concentration limit, where the acceptance probability is of order $d^{-1/2}$ in dimension $d$, implying a polynomial complexity guarantee.
In this note, we verify this dimension-dependent prediction, establishing the uniform guarantee $\inf\{α_D:D=D^\top\in\mathbb{R}^{d\times d}\}\ge c_\star/\sqrt{d}$, where $c_\star=0.759\ldots$. A one-dimensional high-concentration limit demonstrates that the $d^{-1/2}$ rate is unimprovable and that even the constant $c_\star$ cannot be improved beyond $0.857\ldots$. The proof relies on a novel interpretation of the acceptance probability and a comparison principle for weighted sums of chi-squared random variables, which may be of independent interest.2026-08-20T18:07:38Z8 pages, no figuresSam Powerhttp://arxiv.org/abs/2608.20279v1Robustness of random-walk Metropolis for steep potentials2026-08-20T17:13:48ZIn Markov chain Monte Carlo sampling, light-tailed target distributions present something of a poisoned chalice: their light tails offer good confinement, and tend to imply good mixing properties for natural continuous-time dynamics, but the steepness of their tail decay means that they often fall outside of the scope of modern quantitative convergence theory. For usual gradient-based samplers, this reflects a genuine instability issue, whereby Metropolis acceptance rates can degrade badly. In this work, we study the gradient-free random-walk Metropolis sampler, and show that for a wide range of light-tailed targets, the acceptance probability remains stable for reasonable choices of proposal variance, from which effective and favourable mixing time estimates can be deduced. The analysis relies on a simple relationship between the first and second derivatives of the log-density of the target distribution.2026-08-20T17:13:48Z11 pages, no figuresSam Powerhttp://arxiv.org/abs/2609.27879v1Covariance Kernels on Unordered Pair Spaces: Theory and Applications to Network-Valued Data2026-08-20T16:46:34ZMany scientific problems are relational: the quantity of interest is a connection between two objects, while information about similarity is available for the objects themselves. We develop a covariance framework for unordered relationships that transfers object-level geometry to the relations they form while preserving endpoint identity and invariance to ordering. Building on symmetric pairwise-kernel representations, we develop theory for the loop-free domains used in undirected networks. We establish spectral interlacing and trace-loss results after self-pairs are removed, connect the relational spectrum to regularization and risk, derive an exact inferential error for spectral truncation, and quantify how perturbations of the underlying geometry propagate to pair covariance and estimation. Simulations show when structured borrowing improves estimation and how geometric misspecification can erode that benefit. We apply the framework to autism neuroimaging using resting-state functional-connectivity data from the Autism Brain Imaging Data Exchange (ABIDE). With a 116-region parcellation and 6,670 unique connections, the application shows that a large connectome can have a much smaller effective covariance dimension. It also demonstrates that high explained covariance alone is insufficient for choosing a low-rank representation when inferential accuracy is the goal. The framework provides a principled foundation for covariance and regularization when the statistical units are unordered relationships.2026-08-20T16:46:34Z33 pages; theoretical results, simulation studies, and an application to ABIDE autism neuroimagingMontserrat Fuenteshttp://arxiv.org/abs/2608.20243v1A Bayesian Edge-Space Framework for Whole-Connectome Inference in Multisite Autism Neuroimaging2026-08-20T16:36:56ZAutism spectrum disorder (ASD) is associated with heterogeneous alterations across distributed brain systems, creating challenges for whole-connectome inference. The difficulty arises not only from the large number of connections, but also from dependence among effects indexed by anatomically and functionally related region pairs. We introduce a Bayesian Edge-Space regression framework that treats each participant's connectome as a network-valued response and models the adjusted ASD effect over unordered brain-region pairs. The main methodological contribution is a positive-semidefinite covariance construction defined directly on connections. Anatomical and diagnosis-blind functional similarities are lifted from regions to edge space through a symmetrized endpoint-matching operation that preserves endpoint identity and is invariant to endpoint ordering. An additive Bayesian hierarchy estimates anatomical, functional, and interaction contributions together with multisite adjustments and connection-specific effects. Theoretical results establish covariance validity and continuous nesting of the structured components. Low-rank kernel representations and an exact sufficient-statistic reduction enable whole-connectome computation without preliminary edgewise estimation. Simulations show improved recovery of the effect surface, particularly under weak signals. In the Autism Brain Imaging Data Exchange, the framework identifies widespread reductions together with localized increases in ASD-associated connectivity. This pattern supports heterogeneous reorganization across distributed neural systems rather than uniform hyper- or hypoconnectivity. Under the fitted parameterization, the functional component has the largest structural scale, indicating organization beyond anatomical proximity alone.2026-08-20T16:36:56Z49 pages, 4 main-text figures, 2 main-text tables; supplementary material includes theoretical results, simulation studies, computational diagnostics, and additional figuresMontserrat FuentesVeronica B. Pattersonhttp://arxiv.org/abs/2607.21847v2Distributional Determinantal Point Process for Repulsive Clustering of Distributions2026-08-20T14:37:33ZWe introduce the distributional determinantal point process (dDPP) as a novel repulsive point process whose atoms are probability distributions rather than points in a real space. The dDPP is constructed via an L-ensemble with a sliced Wasserstein (SW) kernel between distributions. We show its validity as a well-defined point process. In the discrete setting, we derive concentration results for plug-in estimators of the L-ensemble, the correlation kernel, and their determinants given i.i.d. samples from the distributional atoms. Leveraging this framework, we propose a distribution-valued random partition model by way of a repulsive generalized Bayesian mixture model. The model places a dDPP prior over the atoms of the mixing measure and defines a generalized likelihood based on SW distance. To summarize posterior inference, we develop a decision-theoretic approach to report a point estimate of the mixing measure as a Bayes rule under a hierarchical optimal transport utility function. The latter is a natural choice given that the mixing measure is itself a distribution over distributions. We use the proposed framework for inference with single-cell gene expression data and human epilepsy data, producing interpretable and well-separated clusters that reflect meaningful structure in the data.2026-07-23T22:32:00Z61 pages, 15 figuresKhai NguyenYang NiElizabeth Juarez-ColungaPeter Muellerhttp://arxiv.org/abs/2606.26804v2Structured Secant Methods to Select Smoothing Parameters for General Smooth Models2026-08-20T14:32:20ZGeneral smooth models replace parameters of a regular likelihood with additive models. The models can include parametric terms, Gaussian random effects, and smooth functions of covariates. The latter are parameterized via a reduced-rank spline basis and regularized via weighted quadratic penalties placed on the basis coefficients. Estimates for these weights (i.e., smoothing parameters) can be obtained by optimizing the Laplace-approximate Bayesian marginal likelihood. Existing (second-order) methods require the Hessian of the log-likelihood to solve this optimization problem approximately - exact optimization requires up to fourth order derivatives - which can be difficult to derive and expensive to evaluate. To address these problems, we present a quasi-Newton variant of the second-order Extended Fellner-Schall (EFS) optimization method. Our qEFS method relies on structured limited-memory secant approximations to the Hessian of the log-likelihood and is principally first-order. However, the approximation can also be accumulated for a sub-block of the Hessian, with the remaining columns being constrained to match those of the actual Hessian. The exact columns then provide additional structure for the sub-block approximation, which becomes more accurate as a result. We show that the qEFS method converges to the EFS method under certain conditions and continues to provide good estimates beyond these circumstances, which we illustrate in simulation studies. Secondary tasks involving the Hessian (confidence interval coverage & model selection) require partial approximations to achieve close to nominal performance. We provide Hidden Markov and Tweedie model examples, for which the qEFS method is substantially easier to implement than alternative methods.2026-06-25T09:42:01Zv2: references updated, clarifications on applicability of Eq. 11Joshua KrauseJelmer P. BorstJacolien van Rijhttp://arxiv.org/abs/2608.19767v1skchange: Fast and Flexible Algorithms for Changepoint Detection2026-08-20T08:11:15ZSkchange is an open-source Python library for detecting structural changes in time series. It implements modern change detection algorithms within a unified and extensible framework. The algorithms are modular and composable, and they include changepoint search methods based on both cost minimisation and statistical tests. Key features include the detection of anomalous segments in addition to changepoints; theoretically well-founded fast and approximate search methods; theoretically well-founded algorithms for high-dimensional data, covering settings where either few or many features change simultaneously; utilities for automatic and data-driven penalty calibration, which balances false alarms against missed detections; and a large collection of built-in costs and statistical tests. The design follows established scikit-learn conventions to streamline both user and contributor experience, and Numba is used extensively to achieve high computational performance. Source code and documentation are available at https://github.com/NorskRegnesentral/skchange.2026-08-20T08:11:15Z6 pages, 3 figuresMartin TvetenJohannes Voll KolstøPer August Jarval Moenhttp://arxiv.org/abs/2608.19722v1Copula-Based Reconstruction and Clustering of Coccidioides Minimum Inhibitory Concentration Profiles2026-08-20T07:18:57ZCoccidioidomycosis is a fungal lung infection endemic to parts of the Pacific Northwest and southwestern United States, Mexico, Central America, and South America. A 2017 study by Thompson et al. reported that many analyzed Coccidioides isolates had elevated minimum inhibitory concentration values for fluconazole but comparatively low minimum inhibitory concentration values for other triazole drugs. We constructed 172 synthetic joint minimum inhibitory concentration profiles from the published drug-specific marginal frequency tables. Values sampled from each marginal distribution were paired across drugs using a Gaussian copula, and potential cross-drug patterns were examined using k-means clustering, hierarchical clustering, and Gaussian mixture models. On the selected reconstruction, a forced four-cluster partition consistently identified an upper caspofungin minimum-inhibitory-concentration tail. Across the evaluated dependence scenarios, the gap statistic favored a single cluster in 96% to 100% of nested reconstructions, providing no evidence that the reconstructed multivariate data supported a broader multicluster structure. A separate simulation study examined recovery of an upper-severity-score category defined from prespecified quantiles of a composite minimum-inhibitory-concentration score. The study included 4,000 replications across 20 prevalence and measurement-noise conditions. The multiclass adjusted Rand index ranged from moderate to high depending on the method and prevalence, whereas binary partition agreement for the upper-severity-score category approached zero at 1% prevalence even though recall remained near 1.0. Clustering conclusions therefore depended on the assumed cross-drug dependence structure and on the realized number of observations in the target category.2026-08-20T07:18:57ZMiranda WascoParomita Banerjeehttp://arxiv.org/abs/2608.19716v1A Bayesian Time-Varying SEIARD Model for State-Level COVID-19 Transmission and Mortality in the United States2026-08-20T07:14:48ZWe conduct a retrospective analysis of COVID-19 transmission dynamics across U.S. states using a modified population-based Susceptible-Exposed-Infectious-Asymptomatic-Recovered-Deceased (SEIARD) compartmental model. The proposed framework introduces time-varying transmission, reporting, and mortality rates to capture temporal variations in public behavior and policy interventions during the pandemic. In particular, the transmission rate is modeled as a function of population mobility (derived from Google Mobility Reports), with a residual time-decay term capturing the net effect of unobserved factors such as behavioral adaptation and control measures, while reporting is linked to nationwide testing strategies. We employ a Bayesian approach to integrate multiple data sources and quantify uncertainties in model parameters. The model explicitly distinguishes between symptomatic and asymptomatic infectious individuals and links the latent epidemic states to observable quantities, including reported cases and deaths, through a dynamic reporting function. This retrospective modeling framework provides insights into state-level epidemic trajectories and supports data-driven decision-making for optimal allocation of healthcare resources and evaluation of public health interventions during future pandemics. We further apply a clustering analysis to the posterior parameter estimates to identify groups of U.S. states exhibiting similar epidemiological characteristics, revealing substantial regional heterogeneity in transmission intensity, reproduction dynamics, and mortality burden.2026-08-20T07:14:48ZParomita Banerjeehttp://arxiv.org/abs/2608.19631v1Variational Goal-Oriented Optimal Experimental Design for Mixed-Distribution Quantities of Interest: Application to Ship Roll Safety2026-08-20T04:52:02ZGoal-oriented optimal experimental design (GO-OED) selects experiments according to the expected information gain (EIG) about a quantity of interest (QoI) rather than the full parameter vector. This work develops a variational GO-OED formulation for mixed discrete-continuous QoI laws arising in probabilistic mechanics when thresholding or event-based transformations map a positive-probability set of uncertain inputs to a common value while other inputs produce continuously varying responses. The motivating application is ship roll safety assessment in random waves, where the QoI is the temporal exceedance probability above a prescribed roll-angle threshold. This quantity is zero when no exceedance occurs and varies continuously over positive values otherwise. A purely continuous variational approximation does not dominate a posterior QoI law containing an atom, yielding an infinite Kullback-Leibler divergence and a trivial Barber-Agakov lower bound of $-\infty$. Scoring atom samples using continuous density values instead changes the objective and does not produce a valid lower-bound estimator. We introduce a mixed variational approximation that models the conditional atom probability and continuous component separately, with a normalizing flow used for the latter. An analytical example recovers the correct EIG landscape, while the ship roll application provides stable EIG lower-bound estimates and identifies informative wave conditions for temporal-exceedance-probability inference.2026-08-20T04:52:02ZChen ChengXun HuanYulin Panhttp://arxiv.org/abs/2608.19599v1Efficient Poisson Subsampling for the Partially Linear Additive Cox Model2026-08-20T03:37:45ZTo address the computational and storage challenges often encountered in large-scale survival data analysis, we propose an efficient Poisson subsampling method for the partially linear additive Cox model. This model provides a flexible yet interpretable framework by incorporating linear covariate effects, additive nonparametric components for nonlinear covariates, and a nonparametric baseline hazard function. The proposed method adopts B-spline basis functions to approximate the nonparametric components and employs the decorrelated score technique to construct a Poisson subsampling-based estimation equation, based on which we establish the asymptotic normality of the resulting estimator and derive the optimal subsampling probabilities according to the L-optimality criterion. Furthermore, we design a two-step adaptive algorithm for practical implementation. The proposed approach enables computationally efficient statistical inference for large-scale survival analysis without processing the full dataset. We validate the performance of the proposed method through extensive simulation studies and a real-world application to a lymphoma cancer dataset, demonstrating its efficiency and accuracy in large-scale settings.2026-08-20T03:37:45ZDongxiao HanLiuquan SunChunjie WangDehui WangHaiYing WangHaixiang Zhang