https://arxiv.org/api/6m8TVDC+jF1ZNYQf7COQEBMdwEM 2026-10-02T20:20:35Z 10675 255 15 http://arxiv.org/abs/2608.23625v1 Separating Time-Varying Network Composition from Predictive Dependence under Noisy Network Measurement 2026-08-22T23:01:02Z A common question about networked time series is whether outcomes changed because shocks transmit more strongly or because the pattern of connections changed. Standard practice inserts a recorded network into an outcome regression and reads movements of the fitted coefficient as changes in transmission strength. When the network is latent, time varying, and measured with error, this reading fails: changes in strength and changes in composition can produce the same outcome distribution at a single date, and the population coefficient moves under composition changes alone. The question becomes answerable when outcomes are analyzed jointly with repeated noisy measurements of the network, such as paired reports of bilateral trade flows. For the joint model we establish necessary and sufficient conditions for local identification, estimators of the strength and composition paths, a simultaneous confidence band for the strength path, confidence sets that remain exact under weak identification, breakdown bounds under common reporting bias, and an exactly sized test that detects changes on the observed path and attributes them to strength or to composition. Simulations assess each procedure at its stated boundary. On a mirror-reported trade panel of eighteen economies over 1995 to 2020, the diagnostics flag exactly the crisis years and the composition coordinate attached to European Union membership declines by roughly two thirds. The estimand is predictive dependence, not a causal effect. 2026-08-22T23:01:02Z Marios Papamichalis Regina Ruane Theofanis Papamichalis http://arxiv.org/abs/2509.03945v3 Prob-GParareal: A Probabilistic Numerical Parallel-in-Time Solver for Differential Equations 2026-08-22T11:55:43Z We introduce Prob-GParareal, a probabilistic extension of the GParareal algorithm designed to provide uncertainty quantification for the Parallel-in-Time (PinT) solution of (ordinary and partial) differential equations (ODEs, PDEs). The method employs Gaussian processes (GPs) to model the Parareal correction function, in line with GParareal, further enabling the propagation of numerical uncertainty across time and yielding probabilistic forecasts of the system's evolution. Furthermore, Prob-GParareal accommodates probabilistic initial conditions and maintains compatibility with classical numerical solvers, ensuring its straightforward integration into existing Parareal frameworks. Here, we first conduct a theoretical analysis of the computational complexity and derive error bounds of Prob-GParareal. Then, we numerically demonstrate the accuracy and robustness of the proposed algorithm on five benchmark ODE systems, including chaotic, stiff, and bifurcation problems. To showcase the flexibility and potential scalability of the proposed algorithm, we also consider Prob-nnGParareal, a variant obtained by replacing the GPs in Parareal with the nearest-neighbors GPs, illustrating its improved computational performance on an additional PDE example. This work bridges a critical gap in the development of probabilistic counterparts to established PinT methods. 2025-09-04T07:09:59Z Guglielmo Gattiglio Lyudmila Grigoryeva Massimiliano Tamborrino http://arxiv.org/abs/2607.23280v2 Uniqueness in multivariate curve resolution, re-tuned 2026-08-22T05:29:52Z There are many misunderstandings about the term and interpretation of uniqueness in multivariate curve resolution tasks. In CAC2026 Tarragona, it turned out that even mathematicians do not properly construe Manne's theorems. So we decided to provide this re-tuned summary, using several original quotes, their explanations, and newly created illustrative figures to get the points across. This contribution then explicitly offers the following: (i) a unified interpretation of Manne's conditions, DBU, GRU, minimal constrained duality, and particular-solution-based criteria; (ii) newly constructed visual illustrations, including a counterexample; and (iii) a practical workflow for diagnosing uniqueness in nonnegative bilinear MCR. 2026-07-25T16:35:13Z Robert Rajko Rabeea Razaq http://arxiv.org/abs/2112.03738v3 A more efficient algorithm to compute the Rand Index for change-point problems 2026-08-22T03:45:38Z We provide a more efficient algorithm for computing the Rand Index when the data clusters come from a change-point detection problem. Given the number of data points $N$ and two change-point sets of size $r$ and $s$, the algorithm runs on $O(r+s)$ time complexity and $O(1)$ memory complexity. The Rand Index computation for the general clustering problem, in contrast, requires the $N$ cluster memberships and has a $O(N)$ complexity in both time and memory. 2021-12-07T14:52:04Z Lucas de Oliveira Prates http://arxiv.org/abs/2609.27924v1 A Prediction--Correction Analysis of Two-Way Block Splitting in Distributed Learning 2026-08-22T03:05:36Z This note revisits the convergence of the Parikh--Boyd two-way block-splitting algorithm for large-scale distributed learning through the He--Yuan prediction--correction framework. Simultaneous row--column partitioning is also relevant to hybrid federated learning, where data may be heterogeneous in both samples and features. We lift the reduced iteration to an equal-dimensional product space and reconstruct the primal and dual coordinates omitted by its implementation. The induced orthogonal-complement structure establishes exact iteration-by-iteration equivalence with the published updates. A mixed variational-inequality representation then yields a fundamental descent inequality, global convergence, an ergodic complexity bound, and current-iterate residual estimates under standard convexity, solvability, exact-subproblem, and invariant-initialization assumptions. The analysis also shows that the reduced state recursion is a metric proximal point iteration. No strong convexity, differentiability, or full-rank condition is imposed. The derivation clarifies which algebraic initialization conditions allow the reduced implementation to inherit the full-space convergence and complexity guarantees without modifying its local updates. 2026-08-22T03:05:36Z Xiaofei Wu http://arxiv.org/abs/2608.21654v1 $K$-functions for point processes on complex surfaces 2026-08-21T21:50:51Z The $K$-function is a fundamental summary statistic for assessing clustering or regularity of point processes in two or three dimensional Euclidean space. In practice, however, many planar point patterns arise from projecting locations of objects on a surface in three dimensional space to two dimensional space. For example, when events or objects occur in a landscape their elevation is often ignored. This can lead to erroneous conclusions regarding properties of the point process generating the point pattern. There is not a unique way to extend the classical $K$-function to point patterns on a complex surface. In this paper we propose, explore, and discuss several approaches in terms of their theoretical and computational properties. The best performing approach, coined the surface area $K$-function, can be viewed is an analogue of the classical $K$-function replacing counts of points in Euclidean balls with counts of points in surface geodesic balls. However, an important distinction is that the argument of our surface area $K$-function is area instead of radius of geodesic balls. The performances of the various surface $K$-functions are compared in applications to simulated and real data. 2026-08-21T21:50:51Z Francisco Cuevas-Pacheco Scott Ward Ed Cohen Rasmus Waagepetersen http://arxiv.org/abs/2608.21551v1 Sparse Separable Factor Analysis in the Complex Domain with an Application to Local Field Potential Data 2026-08-21T18:39:10Z Complex-valued arrays arise in signal processing, where scientific interpretation depends on retaining amplitude and phase information. Existing covariance estimation methods either ignore the multiway organization of such data or rely on real-domain embeddings that do not directly exploit their complex structure. We develop sparse separable factor analysis (SSFA), a latent factor model for complex-valued arrays with a separable covariance structure across modes. Each mode-specific covariance matrix is modeled through a low-rank Hermitian factor structure and a diagonal residual covariance matrix. To obtain interpretable estimates, we impose elementwise lasso penalties on the complex loading matrices and estimate the SSFA parameters using a mode-wise parameter-expanded expectation-maximization procedure. The resulting loading updates admit closed-form complex soft-thresholding solutions, which shrink the modulus of each loading while preserving its phase. A separate balancing step resolves the scale nonidentifiability of the separable covariance structure. Simulation studies show that SSFA improves covariance estimation relative to vectorization-based methods, including complex principal component analysis. We apply SSFA to local field potential recordings from mice, where we compare separability structures induced by different groupings of brain region, frequency, and time and perform model-based imputation of recordings missing because of electrode misplacement. 2026-08-21T18:39:10Z 50 pages, 9 figures, and 7 tables Ian Hultman Kirtikanth Kalapatapu Yassine Filali Rainbo Hultman Sanvesh Srivastava http://arxiv.org/abs/2608.21287v1 Comprehensive Regression and Diagnostics for Non-Negative Data Using the BCSreg Package 2026-08-21T16:43:06Z Continuous positive data characterized by high skewness and heavy tails frequently arise in applied statistics. In other applications, these characteristics are accompanied by a point mass at zero, resulting in a non-negative response with a mixed discrete-continuous distribution. Standard regression models often fail to capture these complex features adequately, requiring more flexible approaches. In this paper, we introduce the BCSreg package for R, which provides a comprehensive and unified computational framework for fitting Box-Cox symmetric and log-symmetric regression models for positive continuous data and their zero-adjusted extensions for mixed non-negative data. These broad classes of models accommodate varying degrees of skewness and tail-heaviness while allowing the parameters to be interpreted directly on the original scale of the data. Through a user-friendly multi-part formula interface, the BCSreg package allows practitioners to simultaneously specify regression structures for the scale parameter (which is proportional to the quantiles of the response), the relative dispersion, and, when appropriate, the probability of zero occurrences. Furthermore, the package provides a complete suite of diagnostic tools specifically tailored to these classes of models, including randomized quantile residuals, simulated envelopes, and influence diagnostics. The package's features and capabilities are illustrated through applications to real data. 2026-08-21T16:43:06Z Francisco F. Queiroz Rodrigo M. R. de Medeiros http://arxiv.org/abs/2608.21122v1 LABS: Extending the scope of binary segmentation via a look-ahead device 2026-08-21T14:01:01Z Binary segmentation is widely used for multiple change-point detection because it is fast, simple to describe, and simple to implement. Its validity rests on the requirement that, at each recursive stage, the procedure identifies one of the true change-points when several are present in the current interval. This holds for detecting changes in mean using the CUSUM statistic, but fails in some other settings, in particular in slope change detection for continuous piecewise-linear signals. We propose Look-Ahead Binary Segmentation (LABS), a modification in which the change-points returned by the two child recursions define a narrower interval on which the parent estimate is re-evaluated. LABS inherits the computational speed of standard binary segmentation, but achieves the near-optimal consistency rate of $O\{(n\log n)^{1/2}\}$ in the slope-change signal setting when the LABS model is chosen via either thresholding or a Schwarz-like information criterion. Simulations show that LABS is fast and achieves state-of-the-art performance. 2026-08-21T14:01:01Z Piotr Fryzlewicz http://arxiv.org/abs/2607.09431v2 comprisk: A scikit-learn-compatible Python toolkit for competing-risks survival analysis 2026-08-21T12:18:01Z Medical time-to-event data are frequently subject to competing risks, where the occurrence of one terminal event precludes the others and standard survival methods that treat competing events as censoring yield biased absolute-risk estimates. Valid analysis instead targets the cause-specific cumulative incidence function (CIF). This methodology has been available to applied researchers almost exclusively through R packages, forcing Python-based machine-learning workflows into a Python-to-R round trip. We present comprisk, a scikit-learn-compatible Python toolkit that puts the canonical competing-risks methods behind one API: a scalable competing-risks random survival forest, Fine-Gray subdistribution-hazard regression and a penalized variant, cause-specific Cox regression, the Aalen-Johansen CIF estimator, and Gray's K-sample test, together with competing-risks-aware model evaluation. Every estimator is validated numerically against its R reference implementation. The forest uses a histogram-based, numba-compiled split kernel that fits 10-22x faster than randomForestSRC at comparable discrimination on real clinical cohorts and scales to n = 10^6 on a consumer CPU. comprisk is distributed on PyPI and lets applied researchers run correct, scalable competing-risks analysis without leaving the Python scientific stack. 2026-07-10T13:57:13Z 5 pages. Software paper; package available on PyPI (pip install comprisk). Code archived on Zenodo: 10.5281/zenodo.19876282 Sunny Yang Weiyan Zhao Wanqi Zhao http://arxiv.org/abs/2609.27885v1 A Uniformly Efficient Rejection Sampler for the Multivariate Expectile-Based Distribution 2026-08-21T08:25:11Z The multivariate expectile-based distribution is a `spiked' perturbation of a multivariate Gaussian distribution, proposed by Arbel et al. in 2023. The rejection sampler proposed for this distribution in the initial work is valid, but its acceptance probability can degenerate badly both in high dimension and in the strongly asymmetric limit. We give a simple alternative. After affine whitening and a polar decomposition, the sampling problem reduces exactly to a univariate distribution, and an additional hyperbolic change of variables exhibits this distribution as log-concave, so that a universal and uniformly efficient construction of Devroye applies. The resulting exact sampler therefore enjoys an acceptance probability of at least $1 - \exp \left( -1 \right) = 0.632120 \ldots$, uniformly over the dimension and all admissible asymmetry parameters. 2026-08-21T08:25:11Z 5 pages, no figures Sam Power http://arxiv.org/abs/2608.21466v1 Spectral partitioning for $k$-block averaging kernels of finite Markov chains 2026-08-21T02:53:38Z We develop spectral algorithms for selecting state-space partitions that define averaging kernels for finite, ergodic and reversible Markov chains. For a partition $\mathcal O$, the Gibbs kernel $G_{\mathcal O}$ resamples within the current block from the stationary conditional distribution; when this update is tractable, composing or mixing it with a baseline kernel $P$ can accelerate convergence. We select $\mathcal O$ by rounding the bottom nonconstant eigenfunctions of $P^2$, or the algebraically smallest eigenfunctions of $P$ for additive mixtures, using weighted $k$-means. For $F(\mathcal O)=\|G_{\mathcal O}P-Π\|_{F,π}^2$, we derive exact trace and normalized-cut representations and show that $F$ equals the Pearson $χ^2$-mutual information between the initial block label and the state after one transition, giving this matrix objective a natural probabilistic interpretation. In the two-block case, a threshold sweep exactly solves the associated one-dimensional weighted two-means rounding problem. For general $k \geq 2$, weighted $k$-means rounds the bottom $(k-1)$-dimensional embedding, after which candidates are rescored by $F$; the rounding distortion is a distance between subspaces that yields spectral approximation bounds. We extend the framework to additive mixtures, finite-horizon objectives, and discounted infinite-horizon objectives. In contrast to classical normalized spectral clustering, which uses top nonconstant modes to find low-flow persistent clusters, our method uses bottom modes to favor large normalized cross-block flow and rapid loss of block-label information. Experiments on a controlled-spectrum graph, a mean-field Ising model, and Bayesian variable selection show notable per-iteration improvements in convergence and statistical estimation. 2026-08-21T02:53:38Z 43 pages, 8 figures Michael C. H. Choi Youjia Wang http://arxiv.org/abs/2609.27880v1 Bayesian Tensor Regression for Neuroimaging Data 2026-08-21T01:45:40Z Multidimensional array data, or tensors, arise naturally in neuroimaging and other high-dimensional applications. We propose a parsimonious Bayesian tensor regression model for studies in which a brain image is the response and predictors are vector-valued covariates. The method extends Bayesian envelope dimension reduction to tensor responses, identifying material subspaces that contain regression information while removing variation that is immaterial to the predictors. This formulation leads naturally to a Tucker tensor decomposition and allows spatial dependence and multiple sources of uncertainty to be modeled jointly. We develop a computationally feasible Markov chain Monte Carlo algorithm based on Gibbs sampling and establish posterior consistency for the proposed model. Simulation studies demonstrate substantial gains in estimation accuracy and uncertainty quantification when meaningful dimension reduction is present. We apply the method to Human Connectome Project neuroimaging data to investigate associations between alcohol use and brain activity. The results illustrate the value of Bayesian tensor envelope regression for inference with high-dimensional, spatially dependent imaging responses. 2026-08-21T01:45:40Z 29 pages, 3 figures, 5 tables; includes theoretical results, simulation studies, and an application to Human Connectome Project neuroimaging data Zahra Naji Montserrat Fuentes Liangsuo Ma Hossein Moradi Rekabdarkolaee http://arxiv.org/abs/2608.20633v1 A new analysis of the randomly pivoted Cholesky algorithm 2026-08-21T00:15:20Z The randomly pivoted Cholesky algorithm is one of the leading methods for computing a low-rank approximation to a large positive-semidefinite matrix. However, while it consistently achieves accuracy comparable to or better than competing methods of its type in experiments, its theoretical analysis lags somewhat behind other methods. This paper closes this gap, proving that randomly pivoted Cholesky produces an approximation with expected error within a $1+\varepsilon$ factor of the optimal rank-$r$ approximation in $\mathcal{O}(r/\varepsilon + r\sqrt{\log r})$ steps. This result nearly matches the optimal complexity $Θ(r/\varepsilon)$ for any low-rank approximation method based on a partial Cholesky decomposition (also known as a column Nyström approximation). The paper also presents bounds on the randomly pivoted Cholesky trace and spectral-norm errors that hold with high probability. The mathematical argument is largely due to GPT 5.6-Sol (Pro), with some refinements by the author. 2026-08-21T00:15:20Z 12 pages, 1 figure Ethan N. W. Epperly http://arxiv.org/abs/2505.00877v3 Particle Filter for Bayesian Inference on Privatized Data 2026-08-20T23:00:26Z Differential Privacy (DP) is a probabilistic framework that protects privacy while preserving data utility. To protect the privacy of the individuals in the dataset, DP requires adding a precise amount of noise to a statistic of interest; however, this noise addition alters the resulting sampling distribution, making statistical inference challenging. One of the main DP goals in Bayesian analysis is to make statistical inference based on the private posterior distribution. While existing methods have strengths in specific conditions, they can be limited by poor mixing, strict assumptions, or low acceptance rates. We propose a novel particle filtering algorithm, which features (i) consistent estimates, (ii) Monte Carlo error estimates and asymptotic confidence intervals, (iii) computational efficiency, and (iv) accommodation to a wide variety of priors, models, and privacy mechanisms with minimal assumptions. We empirically evaluate our algorithm through a variety of simulation settings as well as an application to a 2021 Canadian census dataset, demonstrating the efficacy and adaptability of the proposed sampler. 2025-05-01T21:39:12Z Yu-Wei Chen Pranav Sanghi Jordan Awan