https://arxiv.org/api/6ZAvt/zPeLbWZjicaMmNgHW3AAY2026-10-02T17:26:10Z1067519515http://arxiv.org/abs/2609.01133v1Scalable Inversion of Contests with Correlated Performances, Including Softmax and Multinomial Probit2026-09-01T12:09:42ZMultinomial probit choice probabilities over n alternatives are Gaussian orthant integrals, computed by simulation for thirty years, one expensive integral per alternative. Inversion, which is to say determining item attractiveness consistent with a prescribed choice probability vector, is even more difficult and has been considered impractical for correlated contests when n is large. Yet here, for families lying within a grammar including factor, block and hierarchical covariance structures, we exhibit a calibration tested at n = 1,000,000 reproducing probabilities to very high accuracy, even in the extreme tail. We must return to much smaller problems for any performance comparison to be possible due to limitations of the prior art. The Geweke-Hajivassiliou-Keane simulator is the standard (and still appropriate for high rank) but is two hundred times slower already at n = 200, and its measured cost grows roughly as n^2.8 while ours is linear. Furthermore our approach applies to any continuous performance distributions within reason: the Thurstone-Mosteller model families thereby become a practical alternative to logit at modern scale.2026-09-01T12:09:42ZPeter Cottonhttp://arxiv.org/abs/2510.01941v2A debiased Bernoulli factory and unbiased estimation of a probability2026-09-01T09:49:31ZGiven a known function $f : [0, 1] \to (0, 1)$ and a random but almost surely finite number of independent, Ber$(x)$-distributed random variables with unknown $x \in [0, 1]$, we prove the existence of an unbiased, $[0, 1]$-valued estimator of the probability $f(x) \in (0, 1)$. Our estimator is based on so-called debiasing, or randomly truncating a telescopic series of consistent estimators. Debiased estimators of a probability are not typically constrained to $[0, 1]$, or even bounded, even when all consistent estimators used as inputs are. We show that constructing the series of consistent estimators from the coefficients of a particular Bernoulli factory yields provable boundedness provided $f \in C^ρ[0, 1]$ for $ρ> 3$. Our result can be thought of as a novel Bernoulli factory with the appealing property that the required number of Ber$(x)$-distributed random variates is independent of their outcomes.2025-10-02T12:04:19Z10 pagesJere KoskelaToni KarvonenKrzysztof ŁatuszyńskiDario Spanòhttp://arxiv.org/abs/2609.00905v1When Metropolis and Hastings Meet Bradley and Terry: Exact MCMC From Preference Voting2026-09-01T08:36:43ZSampling from distributions conditioned on desired semantic properties is an emerging challenge in modern generative modeling. Metropolis-Hastings (MH) provides a principled route to conditional sampling, but requires access to exact pointwise target-density evaluations, which are not available in generative settings. Meanwhile, pairwise comparisons by humans or model "judge" are highly accessible and have proved valuable across diverse applications. We introduce Pref-MH, a general exact MH sampler for judge-induced conditional distributions using only stochastic binary pairwise comparisons. Our key observation is that the MH unnormalized density ratio matches the preference odds of the Bradley-Terry (BT) choice model. The central challenge is that while MH requires precise ratio computation, BT judges provide only sampled binary feedback. To this end, we develop a valid accept/reject rule whose resulting Markov chain provably converges to the target distribution. We further show that, for a fixed proposal kernel and budget, Pref-MH is optimal in the Peskun-Tierney sense among this class of exact reversible acceptance rules. Experiments on text generation and molecular design with LLM judges, as well as image generation with VLM judges, demonstrate that Pref-MH provides a practical and flexible approach to conditional sampling when comparative feedback is relatively easy to obtain.2026-09-01T08:36:43ZAriel SmogorghevskiNir RosenfeldYaniv Romanohttp://arxiv.org/abs/2607.25250v2A Copula-Based Regression Framework for Enhanced Prediction under Heteroscedasticity2026-09-01T08:22:58ZClassical regression approaches, including ordinary least squares, rely on strong assumptions such as constant variance and normality of residuals, which are often violated in real-world data. Although log-transformation is commonly used to stabilise variance, it may introduce re-transformation bias and fail to address heteroscedasticity and asymmetric dependence structures adequately. To overcome these limitations, this study proposes a copula-based regression framework for modelling data in the presence of heteroscedastic error structures. The proposed copula-based regression framework separates marginal distributions of the response and explanatory variables from their dependence structure, allowing flexible modelling of different tail-dependent relationships. The proposed approach explicitly accounts for heteroscedasticity without requiring restrictive distributional assumptions. A comprehensive simulation study and two real-world applications were considered under heteroscedastic scenarios to compare the performance of the proposed method with existing methods. The simulation results demonstrated that the proposed copula-based model consistently outperformed conventional approaches, achieving an average mean absolute percentage error of 0.21, compared with 0.27 and 0.36 for the linear and log-linear models, respectively. In the first application, which exhibited clear heteroscedasticity, the copula-based model achieved the lowest MAPE, although the overall differences between the copula-based and GAMLSS models were not substantial. In Application 2, all models achieved low predictive performance, as there was a moderate linear relationship between the response and predictor variables. Overall, the findings indicate that no single model consistently dominates across all settings, while copula-based regression provides a flexible and competitive alternative for heteroscedastic data.2026-07-28T03:52:19ZDeepani HemachandraJagath SenarathneMahasen Dehideniyahttp://arxiv.org/abs/2609.00773v1Deep Skew-t Mixture Models2026-09-01T06:02:11ZHigh-dimensional clustering is challenging when component distributions are both heavy-tailed and directionally asymmetric. We propose a deep skew-$t$ mixture model (DStMM), a hierarchical factor-analytic mixture based on the generalised-hyperbolic skew-$t$ normal mean--variance representation. A shared inverse-gamma mixing variable is propagated along each complete latent pathway, allowing heavy tails and directional asymmetry to be modelled jointly while preserving conditional Gaussianity. Each complete pathway therefore admits an exact GHST marginal representation. We formalise the reductions to symmetric deep $t$, Gaussian deep-mixture, and single-layer GHST factor-analytic models, discuss local non-identifiability and the implementation-level parameter-counting convention, and derive the conditional generalised inverse Gaussian law used for estimation. Estimation is carried out by a stochastic/Monte Carlo EM algorithm, with an explicit implementation-based parameter count for BIC architecture comparison. Simulation studies show that DStMM performs similarly to the symmetric robust model when skewness is absent but provides increasing gains as directional asymmetry becomes stronger, particularly under heavier tails; the same qualitative behaviour persists under smaller samples and unequal mixture proportions. Two real-data applications provide complementary evidence. On the UCI handwritten-digit benchmark, DStMM gives the strongest clustering performance under a common deep architecture, while on the Gas Sensor Array Drift data, DStMM improves on both deep Gaussian and deep $t$ alternatives and, under the implemented BIC criterion, selects a non-trivial second mixture layer. Together, these results support the value of propagating skewness and heavy-tail variation through a deep latent mixture while retaining an exact pathway-level likelihood.2026-09-01T06:02:11ZJinran WuYou-Gan WangGeoffrey J. McLachlanhttp://arxiv.org/abs/2603.07701v3Fractional Topological Phases, Flat Bands, and Protected Zero-Energy Edge States in Small Cyclic Quantum Walks2026-09-01T04:15:34ZWe report the first realization of a fractional topological phase in a fully unitary, noninteracting discrete-time quantum walk implemented on finite cyclic graphs. Using a single-coin split-step cyclic quantum walk (SCSS-CQW), we uncover topological phenomena that are inaccessible within conventional cyclic quantum-walk dynamics. The protocol enables controlled engineering of quasienergy spectra, flat bands, and topological phase transitions through the step-dependency parameter and coin-rotation angle. We show that cyclic graphs with even and odd numbers of sites exhibit qualitatively different band structures, while rotational flat bands arise exclusively in $4n$-site cycles; a general analytic condition for their emergence is derived. The SCSS-CQW produces fractional winding numbers $\pm \frac{1}{2}$ (Zak phases $\pm \fracπ{2}$), in sharp contrast with the integer invariants in standard cyclic quantum walks. These fractional invariants lead to an unconventional bulk-boundary correspondence and support protected {zero energy} edge states beyond the usual integer topological classification. In the step-dependent protocol, transitions between distinct fractional winding sectors generate robust edge modes pinned at zero quasienergy. Numerical simulations show that these states remain stable in the presence of both dynamic and static coin disorder as well as phase-preserving perturbations, while survival-probability analysis demonstrates their long-time persistence. Requiring only a constant number of detectors independent of the evolution time and fewer operators, the proposed scheme offers a minimal-resource and experimentally accessible platform for realizing fractional topology, flat bands, and protected zero energy edge states in small-scale synthetic quantum systems.2026-03-08T15:59:21Z30 pages, 27 figures, 3 tablesDinesh Kumar PandaAditi RathColin Benjaminhttp://arxiv.org/abs/2609.00616v1An efficient EM algorithm for both element-wise and structural missingness in matrix-variate normal mixture models2026-09-01T03:01:09ZMatrix-variate data with missing entries arise frequently in applications where observations are naturally organized as two-dimensional arrays. Although the matrix normal distribution provides a parsimonious model through its Kronecker covariance structure, standard EM estimation can be computationally expensive because arbitrary missingness patterns typically destroy this separability in the E-step. In this paper, we propose an efficient partial EM algorithm for matrix-variate normal data with missing entries. The proposed method updates the conditional mean and covariance of the missing component through coordinate-wise approximations, avoiding repeated inversion of pattern-specific covariance matrices and avoiding construction of the full vectorized covariance matrix. We further develop a specialized update for submatrix missingness, where the missing-block precision retains a Kronecker product structure, and the covariance update can be carried out independently in the row and column directions. Simulation studies show that the proposed methods substantially reduce computation time compared with exact EM while preserving nearly identical observed-data likelihood across a range of dimensions and missing proportions. A real-data application to hyperspectral image patches demonstrates that the proposed imputation strategy can be embedded within a matrix-variate mixture model for simultaneous imputation and clustering.2026-09-01T03:01:09ZHanzhang LuJeffrey L. AndrewsRyan P. Brownehttp://arxiv.org/abs/2512.13067v2Group-averaged Markov chains II: tuning of group action in finite state space2026-09-01T01:01:35ZWe study group-averaged Markov chains obtained by augmenting a $π$-stationary kernel $P$ with orbit kernels induced by a group action. We analyse the Gibbs ($G$), Metropolis--Hastings ($M$), and Barker ($B$) kernels, their sandwiches $QPQ$, and mixtures $\tfrac{1}{2}(P+Q)$, where $Q\in\{G,M,B\}$. Under suitable conditions, $M^t$ and $B^t$ converge blockwise to $G$. The projection chains of $GPG$ and $P$ coincide, while every sandwich $QPQ$ has absolute spectral gap no smaller than that of reversible $P$. For $GPG$, we derive an additive asymptotic-variance bound, prove monotonicity for $G$-invariant observables, and identify it as the Kullback--Leibler (KL) information projection of $P$ onto the $G$-invariant kernels. For a fixed orbit partition, the spectral and KL properties of $GPG$ reduce to those of a lower-dimensional orbit-space chain. Among Gibbs projections with a prescribed number of orbits, we identify the partition minimizing KL divergence to stationarity and characterize exact stationarity. Finally, alternating group projections converge at a rate determined by singular values of an overlap matrix and, in structured cases, can yield exact sampling with logarithmically many group actions. These results motivate tuning heuristics and yield polynomial mixing for a Curie--Weiss example in a regime where Glauber dynamics is exponentially slow.2025-12-15T08:05:24Z57 pages, 3 figuresMichael C. H. ChoiRyan J. Y. LimYoujia Wanghttp://arxiv.org/abs/2609.00408v1Non-Uniform Random Scans in Gibbs Sampling and CAVI2026-08-31T21:40:34ZGibbs sampling and coordinate ascent variational inference (CAVI) are two basic coordinate-wise methods for statistical computation. Recent analyses under strong log-concavity establish convergence rates for versions of these algorithms that update one uniformly selected block at each step. We extend both results to arbitrary fixed, strictly positive selection probabilities. The rates are governed by a selection-adapted convexity constant $λ^{\star}_θ$, defined using the block-smoothness constants and the selection probabilities $θ$. The same constant yields a contraction of relative entropy for the Gibbs sampler and a contraction of the mean-field objective gap for random-scan CAVI. The new bounds recover the uniform-scan results, and are never weaker than the naive comparison based on the smallest selection probability. They provide a principled way to adapt the scan to heterogeneous block geometry using curvature information.2026-08-31T21:40:34Z13 pages, no figuresSam Powerhttp://arxiv.org/abs/2609.00307v1Generalized Bayesian Clustering with Regression for Unaligned Longitudinal Binary Data2026-08-31T19:58:57ZWe propose a generalized Bayesian clustering with regression model for unaligned longitudinal binary outcomes, motivated by seizure diary data from the Human Epilepsy Project. Seizure diaries are sparse, irregularly observed, and vary enormously across patients. A single fully-specified generative model tends to be either misspecified or computationally inefficient. We address the challenge by two strategies. We set up a regression by way of clustering as model-based clustering using a mixture model. For the latter, we take a generalized Bayesian perspective which replaces the full likelihood with a loss-based update using a generalized likelihood. We combine a trajectory similarity loss and a regression loss, so that clustering is informed by both trajectory similarity and the prediction of outcomes The trajectory similarity loss is constructed by representing each trajectory as an (empirical) distribution of subsequences, called reads, and then is defined based on the sliced Wasserstein distance between these empirical distributions. This loss allows alignment-free comparison of sequences that are irregularly observed or temporally misaligned, and it scales quasi-linearly in trajectory length. The regression loss is the negative log-likelihood of a probit regression. A prior on the cluster-specific parameters is defined by way of a Dirichlet process prior on the mixing measure.2026-08-31T19:58:57Z29 pages, 6 figuresKhai NguyenElizabeth Juarez-ColungaPeter Muellerhttp://arxiv.org/abs/2504.00049v3Scalable Durational Event Models: Application to Physical and Digital Interactions2026-08-31T15:11:05ZDurable interactions are increasingly observed in social network analysis with precise timestamps. Phone and video calls, for instance, are events to which a specific duration can be assigned. We term data encoding interactions with start and end times ``durational event data''. Recent advances in data collection have enabled the observation of such data over extended periods and across large populations of actors. Methodologically, we propose the Durational Event Model, an extension of Relational Event Models that decouples the modeling of event incidence from event duration. Computationally, we derive a fast, memory-efficient, and exact block-coordinate ascent algorithm to facilitate large-scale inference. Theoretical complexity analysis and numerical simulations demonstrate the computational superiority of this approach over state-of-the-art methods. We apply the model implemented in the R package redeem to physical and digital interactions among college students in Copenhagen. Our empirical findings reveal that past interactions drive physical interactions, whereas digital interactions are influenced predominantly by friendship ties and prior dyadic contact.2025-03-31T00:03:11ZCornelius FritzRiccardo RastelliMichael FopAlberto Caimohttp://arxiv.org/abs/2608.30845v1Scalable Statistical Inference in Stochastic Gradient Descent2026-08-31T14:13:55ZConstructing confidence regions for stochastic gradient descent (SGD) ideally requires estimating the asymptotic covariance matrix, a severe computational bottleneck in high dimensions. Traditional cancellation-based batch means methods bypass this estimation but require inverting a sample batch covariance matrix. This introduces strict mathematical degeneracy when the parameter dimension exceeds the number of batches. To address this problem, we utilize equal batch size batch means method and propose a simultaneous, marginal-friendly framework. The proposed marginal statistics has a asymptotic Student's $t$-distribution, and eliminates the matrix inversion step, entirely circumventing high-dimensional degeneracy. To achieve valid simultaneous coverage, we present an algorithm utilizing wild bootstrap samples drawn from a statistic as a function of only the diagonals of the variance-covariance estimator, and to further incorporate the contribution of cross-dependencies, we introduce an efficient Quasi-Monte Carlo procedure utilizing a $t$-copula approximation. Additionally, we integrate a Lugsail variance estimator to aggressively correct finite-sample bias and under-coverage. The proposed methodology delivers interpretable, simultaneous hyper-rectangular confidence regions that are statistically robust, memory-efficient, and strictly scalable for high-dimensional inference. The theoretical results are supported by extensive numerical simulation analysis through various aspects of dimension, number of batches and error structure.2026-08-31T14:13:55Z18 pages, 9 FiguresRahul SinghAbhinek Shuklahttp://arxiv.org/abs/2608.30239v1GPU-Parallelization of Markov Chain Pool Decoding with Unbiased MCMC2026-08-31T04:46:43ZMarkov chain pool decoding (MCPD) devised by Knill et al. (1996) identifies likely positive clones from noisy pooled-test results. The standard MCPD estimates clone-wise posterior probabilities using Gibbs sampling, but it may allocate excessive computational effort to low-scoring clones. This paper focuses on parallelizing MCPD on GPU architectures. Whereas the standard MCPD employs systematic-scan updates, we propose a score-weighted update scheme that updates high-scoring clones more frequently. We prove that the stationary distribution of the proposed Markov chain coincides with the target posterior distribution. To enable efficient GPU parallelization, we further incorporate the unbiased MCMC framework of Jacob et al. (2020) and employ a slot-refilling technique based on the arguments by Glynn and Heidelberger (1991) about the coupling of Markov chains. Experiments involving 1,298 clones, 97 pools, and three true positives demonstrate improved recovery compared with uniform decoders, while maintaining high overlap under high-noise conditions.2026-08-31T04:46:43ZTakato UenoShuji Kijimahttp://arxiv.org/abs/2608.30096v1Multifidelity Computer Model Emulation Via Diffusion Model Steering and Targeted Maximum Likelihood2026-08-31T00:00:41ZWe develop a multifidelity method for fusing low-resolution simulations with computationally expensive high-resolution simulations, which are run infrequently and are therefore prone to bias. We formulate this fusion as a constrained optimization under missing-not-at-random (MNAR) selection bias. This formulation searches for the exponentially tilted high-resolution distribution that minimizes KL divergence from the biased baseline, subject to moment constraints derived from low-resolution simulations. This optimization requires first estimating the biased baseline conditional density $f$ as a nuisance parameter. We estimate $f$ using a score-based diffusion model. To eliminate the generative model's regularization bias that harms the downstream task, we apply targeted maximum likelihood estimation (TMLE). TMLE debiases $\hat{f}$ via a targeted exponential tilting, rendering the target parameters insensitive to first-order nuisance estimation errors. To execute this computationally, we adapt generative model steering, a technique originally developed for human-preference alignment. Using Feynman-Kac steering with a reward function based on our formulation, we simultaneously execute the exponential tilts for MNAR and TMLE at inference time, avoiding expensive retraining costs. Code available [here](https://github.com/Jong-Min-Moon/multifidel_emul_by_FK).2026-08-31T00:00:41Z27 pages, 1 table, 4 algorithmsJongmin Munhttp://arxiv.org/abs/2608.30093v1Robust K-means Clustering using the Density Power Divergence Measure2026-08-30T23:41:19ZWe introduce a robust clustering method, MK-means DPD, that estimates cluster centers and covariance matrices using density power divergence (DPD) measures combined with Mahalanobis distance, making it resistant to outliers and adaptable to heterogeneous, elliptical clusters, unlike the classical K-means algorithm. Since Mahalanobis distance-based K-means lacks a general convergence guarantee, we further introduce a convergent variant, Density-Consistent MK-means DPD (DC-MK-means DPD), which redefines the cluster assignment step in terms of a pointwise DPD loss. We prove a formal theorem establishing that the resulting algorithm converges in a finite number of steps. We also propose two new robust internal evaluation indices, a Median Davies-Bouldin Index and a Trimmed Calinski-Harabasz Index, to ensure that performance comparisons are not themselves distorted by outliers. The efficacy of the proposed methods is demonstrated on simulated data, showing superiority over existing methods, and on two real datasets: Iris data, to identify similar species, and COVID-19 case fatality rate and infection rate data for countries worldwide, examining the resulting clusters' geographic and socio-economic patterns.2026-08-30T23:41:19ZAnirban MondalParomita BanerjeeAbhijit Mandal