https://arxiv.org/api/dDwHt5amCqnaifb//TW4dh1s71o2026-09-11T18:50:11Z105721515http://arxiv.org/abs/2609.08494v1Nonparametric framework for the definition, adaptive detection and probabilistic interpretation of outliers2026-09-08T09:35:35ZOutlier detection is a fundamental challenge in data processing, with critical implications for robustness across statistical modeling, machine learning and exploratory data analysis. However, existing proposals rarely offer a universal, domain-agnostic definition of an outlier, often relying on heuristic trimming quotas that lack a statistical interpretation. To address this, we propose a nonparametric framework built on a pseudo-isolation outlier score. This score enables a formal, probabilistic definition of an anomaly tied to a false-alarm rate $α$, which extends into a rigorous geometric classification of internal and external outliers. We show that this mechanism seamlessly embeds into any objective-based clustering framework to identify cluster-specific outliers. Here, we integrate it into $K$-means to create ODK-means. The inferential capabilities and topological properties of this framework are explored both theoretically through formal propositions and empirically through extensive simulations and methodological tutorials, highlighting the practical actionability and intuitive appeal of the proposed outlier detection logic.2026-09-08T09:35:35ZTiziano IannaccioMaurizio Vichihttp://arxiv.org/abs/2609.08439v1Bayesian Group Testing Regression with a Shape-Free Dilution Curve2026-09-08T08:42:40ZGroup testing pools specimens to cut the cost of screening, but pooling positive specimens with negative ones lowers assay sensitivity, so that the sensitivity of a pooled test depends on how many of its members are positive. Regression models for group testing accommodate this dilution effect through submodels that fix a one-parameter shape for that dependence. We propose BADGER (Bayesian Analysis of Dilution in Group tEsting Regression), a regression model incorporating a shape-free dilution curve. The pooled sensitivity is modeled as a nondecreasing function represented by nonnegative increments with a Dirichlet prior, so that the parametric submodels become special cases. By an appropriate augmentation, every full conditional of BADGER is in closed form and inference is carried out by an exact Gibbs sampler. The evidence for dilution is measured by a Bayes factor computed from the posterior draws. We validate the BADGER model through simulation studies across different dilution shapes, prevalences and pool sizes. We further illustrate the method on hepatitis E serology from a national health survey.2026-09-08T08:42:40Z23 pages, 3 figures, 14 tables, submitted for journal publicationChun-Hao YangWei-Yan Honghttp://arxiv.org/abs/2609.03523v2Random mixtures in Bayes Hilbert spaces2026-09-08T08:35:19ZWe present a framework for the analysis and unmixing of random density mixtures in the Bayes Hilbert space. General identifiability results for mixtures in Hilbert spaces are established and applied to the Bayes Hilbert space setting. Building on these results, we propose a penalised maximum likelihood approach for the unmixing of Bayes Hilbert mixtures aimed at recovering the statistically space-efficient representation, together with a computationally efficient coordinate-wise maximisation algorithm for its implementation. The methodology is illustrated through a hyperspectral data application, where observations can be naturally embedded in the Bayes Hilbert space and analyzed in terms of distributional shape rather than amplitude. A complementary simulation study demonstrates the interpretability and practical performance of the proposed approach.2026-09-03T08:22:03ZGiulia PatanèSonja GrevenAlessandra Menafogliohttp://arxiv.org/abs/2603.02593v2Composite Wavelet Matrix-Based Transforms and Applications2026-09-08T04:13:17ZOrthogonal wavelet transforms are a cornerstone of modern signal and image denoising because they combine multiscale representation, energy preservation, and perfect reconstruction. In this paper, we show that these advantages can be retained and enhanced by moving beyond classical single-basis wavelet filterbanks to a broader class of composite wavelet-like matrices. By combining orthogonal wavelet matrices through products, Kronecker products, and block-diagonal constructions, we obtain new unitary transforms that generally fall outside the strict wavelet filterbank class, yet remain fully invertible and numerically stable.
The central finding is that such composite transforms often induce stronger concentration of signal energy into fewer coefficients than conventional wavelets. This increased sparsity, quantified using Lorenz curve diagnostics, directly translates into improved denoising under identical thresholding rules. Extensive simulations on Donoho-Johnstone benchmark signals, complex-valued unitary examples, and adaptive block constructions demonstrate consistent reductions in mean-squared error relative to single-basis transforms. Applications to atmospheric turbulence measurements and image denoising of the Barbara benchmark further confirm that composite transforms better preserve salient structures while suppressing noise.2026-03-03T04:39:20Z43 pages, 11 figures, 8 tablesRadhika KulkarniBrani Vidakovichttp://arxiv.org/abs/2609.08172v1Optimal Slice-Adaptive Tuning of Hybrid Slice Sampling2026-09-08T03:01:25ZSlice sampling is a Markov chain Monte Carlo algorithm that draws its next state uniformly from a "slice"---a super-level set of the target density function---at each iteration, thereby providing automatic local adaptivity to the scale of the target. In practice the exact slice is not known, so general-purpose implementations use an approximate slice that is grown from a starting interval of length $w>0$, with a computational cost that depends on $w$. This work presents an analysis of the average per-iteration number of target density evaluations, as a function of $w$, of hybrid slice sampling with various slice-finding schemes for targets with contiguous slices. The paper uses the results of the analysis to develop automated, slice-adaptive tuning schemes along with suboptimality bounds and asymptotic convergence guarantees. Simulations demonstrate that the tuning schemes reliably yield near-optimal slice-adaptive tuning with essentially no dependence on the initial setting of $w$.2026-09-08T03:01:25Z44 pages, 8 figuresTrevor Campbellhttp://arxiv.org/abs/2609.09236v1Identifiability of Latent Space Network Models on Anisotropic Thurston Geometries2026-09-08T01:27:17ZA latent space network model places the nodes in a metric space and lets the probability of a tie decrease with distance. In a space of constant curvature, pairwise distances determine the positions up to an isometry. In the products and in the three remaining three-dimensional model geometries they do not. We study the three geometries that are neither of constant curvature nor products: the Heisenberg group, the solvable group and the universal cover of the unit tangent bundle of the hyperbolic plane. We ask what one network identifies about positions in them and when their geometry is detectable. Two anchors remove the isometry ambiguity. Small configurations are not determined by their distances, and generic local identification holds beyond a finite threshold, certified in the Heisenberg group. We derive the posterior on the quotient by the isometry group. For small configurations, the divergence to the nearest product or constant-curvature competitor is the stress component of the curvature difference and vanishes at high order in the scale; undirected ties therefore detect the geometry only in large networks with large enclosed areas. Directed ties expose it at first order: in the two twisted geometries, asymmetric preferences circulate around triangles in proportion to enclosed area, which no additive ranking produces. On dense competitive-game counter networks, the coupled model beats rankings on every geometry, degree-corrected rankings and free antisymmetric terms. A Euclidean model with the same coupled term matches it, so the gain is the coupling of similarity and circulation through shared coordinates. An additive-and-multiplicative-effects model predicts better still.2026-09-08T01:27:17ZPreprintMarios PapamichalisRegina Ruanehttp://arxiv.org/abs/2609.08023v1Risk-Aware Goal-Oriented Bayesian Optimal Experimental Design2026-09-07T22:14:11ZTraditional Bayesian optimal experimental design (OED) selects measurements that best inform a model's parameters. However, such measurements can be suboptimal for downstream predictions. Goal-oriented OED targets the prediction directly. However, the existing goal-oriented criteria value all reductions in predictive uncertainty equally, with no way to prioritize rare, high-consequence outcomes. In this article, we develop a risk-aware framework that composes risk at three levels, each generalizing an ingredient of classical $I$- and $G$-optimal design: a deviation measure of the posterior predictive uncertainty (generalizing the predictive variance), a risk measure across the prediction domain (interpolating $I$-optimal averaging and $G$-optimal worst-case selection), and a risk measure over datasets (generalizing the expectation). We generate each level from a regret function in the risk quadrangle, so that one triple specifies a practitioner's risk preference. We relax the design to continuous weights on the unit simplex and construct a nested-quadrature estimator that is differentiable in the design variable. This enables solving the optimal design problem with gradient-based methods, avoiding a combinatorial search over candidate designs. For a linear-Gaussian lognormal model and a nonlinear extension, we derive closed-form objectives. These give exact references against which we verify that the estimator converges. We demonstrate this framework for finding optimal sensor placements in an inverse problem governed by an advection-diffusion equation. We find that the risk-aware designs substantially outperform the expected-information-gain baseline, which is statistically indistinguishable from a random allocation.2026-09-07T22:14:11ZJohn D. JakemanRebekah WhiteBart van Bloemen WaandersDrew P. KouriAlen Alexanderianhttp://arxiv.org/abs/2603.16014v2Fast multitask Gaussian processes, with application to surrogate modeling of the quark-gluon plasma2026-09-07T21:16:48ZGaussian processes (GPs) are broadly used for the surrogate modeling of computer experiments with reliable uncertainty quantification. Our motivating application comes from the study of the quark-gluon plasma (QGP), an extreme state of nuclear matter that filled the universe shortly after the Big Bang. To reliably infer properties of the QGP, multiple surrogate models need to be constructed for related particle collision simulation systems (i.e., multiple "tasks"). While there is a body of literature on multitask GPs (MTGPs), such models can be computationally expensive with large datasets: they require $\mathcal{O}(N^3)$ work and $\mathcal{O}(N^2)$ storage for model training, where $N$ is the total number of samples over all tasks. To address this, we propose a new fast MTGP approach, which pairs structured low-discrepancy design points, such as Sobol' points, with special kernel forms for efficient and exact model fitting. This pairing of a kernel and potentially different design points of different sizes across tasks provides a structured Gram matrix, e.g., a circulant block matrix, which we exploit via a novel algorithm for efficient Gram matrix inversion and determinant computation. Our algorithm reduces training costs to $\mathcal{O}(N \log N)$ work and $\mathcal{O}(N)$ storage in the case of equal sample sizes for each tasks. In the worst case of severely unbalanced sample sizes across tasks, our algorithm may require up to $\mathcal{O}(N^2)$ work and storage, with continuous interpolation between these best and worst cases depending on the balance of sample sizes across tasks. An open-source Python implementation is made available in the FastGPs package (https://alegresor.github.io/fastgps/). We demonstrate the effectiveness of our fast MTGP approach on a range of simulation experiments and on our motivating QGP application.2026-03-16T23:46:41ZAleksei G. SorokinPieterjan RobbeYen-Chun LiuSimon MakFred J. Hickernellhttp://arxiv.org/abs/2609.04126v2Fast Computation of Nested Cross-Validation for Penalized Regression2026-09-07T18:58:07ZCross-validation is a resampling procedure that provides a point estimate of generalization error for any predictive model. Cross-validation is widely used for model selection and evaluation. Uncertainty in the cross-validation estimate is challenging to quantify, and estimation of its variance is known to require multiple runs of the entire resampling procedure. Nested cross-validation computes a prediction interval for the generalization error for a given model and training dataset by resampling the entire cross-validation procedure but incurs extraordinary computational cost. We provide an efficient method for computing the nested cross-validation prediction interval using only a single model fit for some penalized regression models including ridge regression, spline smoothing, and some functional regression models. We characterize when our proposed method should be expected to out-perform resampling-based nested cross-validation in various scaling regimes as well as in finite samples. Experiments for functional principle components regression demonstrate non-trivial cases in which our proposed method improves run times substantially.2026-09-03T17:25:17Z17 pages, 5 figuresShuyang CaoAlex Stringerhttp://arxiv.org/abs/2609.07819v1Markov Chain CLTs: Resolving Open Problems2026-09-07T17:55:22ZMarkov chain central limit theorems (CLTs) and their associated variances are very important for implementing Markov chain Monte Carlo algorithms among other applications. Häggström and Rosenthal (2007) presented various results regarding the equality of different formulae for this variance and also posed seven open problems. We resolve all seven in this paper. For stationary, ergodic and reversible chains, we prove that whenever the normalized partial sums satisfy a $\sqrt{n}$-CLT, the variance limit is finite if the function is square-integrable, otherwise undefined. Moreover, failure of the $\sqrt{n}$-CLT forces the normalized partial sums to be non-tight. We also show that Roberts' holding-probability condition precludes a CLT even without assuming reversibility or square-integrability. Finally, we develop a general principle that expresses Fourier coefficients as the autocovariances of an ergodic nonreversible Markov chain.2026-09-07T17:55:22Z44 pages, 2 figures, 1 tableAustin BrownJeffrey S. RosenthalQuan Zhouhttp://arxiv.org/abs/2609.07660v1Thermodynamic Cyclic Processes with Markov Samplers in Bayesian Inference2026-09-07T15:46:03ZThe concept of Markov chain Monte Carlo (MCMC) cycles, an analogy to cyclic processes in heat engines, is presented in order to examine Bayesian inference problems. In this effort, we develop adaptive ensemble schedulers that allow the tuning of external parameters of a Bayesian canonical ensemble during an MCMC run, realising the MCMC cycles in practice. We run these cycles on different statistical models. As a fundamental insight, we find (both theoretically and in practice) that such systems can produce a non-zero net work output if and only if the considered model is non-Gaussian. As such, they may serve as a measure of non-Gaussianity in Bayesian inference, which we test on an example from supernova cosmology.2026-09-07T15:46:03ZHeinrich von CampeBjoern Malte Schaeferhttp://arxiv.org/abs/2609.07283v1Efficient model exploration with the integrated nested Laplace approximation2026-09-07T09:49:33ZModel and variable selection are important topics in Bayesian inference. In particular, the selection of different fixed, random effects and hyperparameters in hierarchical models can be difficult because of their complex structure. In this paper, we introduce the use of approximate Bayesian inference for model and variable selection for hierarchical Bayesian models.
The approach is based on the application of Markov chain Monte Carlo methods on the model space. In particular, the Metropolis-Hastings algorithm is used to sample from the set of model indices. In this way, the marginal likelihood is used to compute the acceptance probability so that it is not required to estimate the parameter models directly. For this, the integrated nested Laplace approximation (INLA) is used because it provides accurate estimates of the marginal likelihood and it can also provide estimates of the posterior marginal of the model parameters.
This method can be applied not only to variable selection but to a wide range of problems subject to model uncertainty. To illustrate the potential of this approach, examples on variable selection, changepoint models and log-Gaussian Cox processes are developed.2026-09-07T09:49:33ZHéctor López-GómezVirgilio Gómez-RubioGonzalo García-Donatohttp://arxiv.org/abs/2609.07257v1Bayesian Generalized Network Autoregressive Model with Structured Shrinkage and Persistence Priors2026-09-07T09:13:42ZWe propose a Bayesian generalized network autoregressive (BGNAR) model for multivariate time series whose component series are associated with the nodes of a known network. The proposed framework combines the parsimonious network structure of the generalized network autoregressive (GNAR) model with structured shrinkage and persistence priors adapted from Bayesian vector autoregressive (BVAR) modeling. We adapt Minnesota-type shrinkage to both own-lag and network-lag coefficients, with prior variances decreasing over temporal lags and, for network effects, neighborhood orders. A hierarchical prior on the own-lag coefficients allows information to be shared across nodes while retaining node-specific heterogeneity. We further adapt the sum-of-coefficients and dummy-initial-observation priors to the GNAR parameterization. Posterior inference is performed using a Gibbs sampler. Simulation studies show that BGNAR can use the same deliberately over-specified temporal and neighborhood structure across datasets without dataset-specific BIC order selection, while maintaining forecasting accuracy comparable to BIC-selected GNAR and outperforming the unrestricted BVAR benchmark in the settings considered. The structured prior regularizes weakly supported coefficients toward zero within this fixed model. Posterior distributions for the dynamic coefficients and posterior predictive distributions for future observations provide direct quantification of parameter and predictive uncertainty. An application to a wind-speed network demonstrates that BGNAR achieves point-forecast performance comparable to GNAR while additionally providing posterior inference on own-lag and network-lag effects and posterior predictive uncertainty.2026-09-07T09:13:42ZSeongmin KimKyusoon Kimhttp://arxiv.org/abs/2507.00923v3ForLion: An R Package for Finding Optimal Experimental Designs with Mixed Factors2026-09-07T03:35:52ZOptimal design is crucial for experimenters to maximize the information collected from experiments and estimate the model parameters most accurately. ForLion algorithms have been proposed to find D-optimal designs for experiments with mixed types of factors. In this paper, we introduce the ForLion package which implements the ForLion algorithm to construct locally D-optimal designs and the Expected Weighted (EW) ForLion algorithm to generate robust EW D-optimal designs, which maximize the determinant of the expected Fisher information matrix under parameter uncertainty. The package supports experiments under linear models (LM), generalized linear models (GLM), and multinomial logistic models (MLM) with continuous, discrete, or mixed-type factors. It provides both optimal approximate designs and an efficient function converting approximate designs into exact designs with integer-valued allocations of experimental units. Tutorials are included to show the package's usage across different scenarios.2025-07-01T16:28:37Z34 pages, 5 figures, 5 tablesSiting LinYifei HuangJie Yanghttp://arxiv.org/abs/2404.15649v3Exact Sampling of Gibbs Measures with Estimated Losses2026-09-06T22:47:49ZA popular strategy for ameliorating some of the shortcomings of Bayesian posterior inference is to instead target a Gibbs measure based on losses that connect a parameter of interest to observed data, and which are known to be robust to misspecification. Existing theory for these procedures treats these losses as being analytically available, but in many situations these losses must be stochastically estimated using pseudo-observations. In these settings, and even under strong assumptions, we show that posteriors based on standard Markov chain Monte Carlo (MCMC) algorithms exhibit strong dependence on the number of these pseudo-observations, and require utilizing a diverging number of pseudo-observations to ensure posterior concentration. To remedy this issue, we introduce a modified piecewise deterministic Markov process (PDMP) sampler, and formally show that its posterior draws have no dependence on the number of pseudo-observations used to estimate the loss within a Gibbs measure. We verify the practical utility of this approach on five examples spanning intractable likelihoods and intractable losses.2024-04-24T05:08:11ZThis is a substantially updated and expanded version of the previous submission, "The Impact of Loss Estimation on Gibbs Measures". This version contains a new algorithm that delivers unbiased samples from the intractable Gibbs measure and was not considered in the previous versionDavid T. FrazierJeremias KnoblauchJack JewsonChristopher Drovandi