https://arxiv.org/api/6m8TVDC+jF1ZNYQf7COQEBMdwEM2026-10-02T20:20:35Z1067525515http://arxiv.org/abs/2608.23625v1Separating Time-Varying Network Composition from Predictive Dependence under Noisy Network Measurement2026-08-22T23:01:02ZA common question about networked time series is whether outcomes changed because shocks transmit more strongly or because the pattern of connections changed. Standard practice inserts a recorded network into an outcome regression and reads movements of the fitted coefficient as changes in transmission strength. When the network is latent, time varying, and measured with error, this reading fails: changes in strength and changes in composition can produce the same outcome distribution at a single date, and the population coefficient moves under composition changes alone. The question becomes answerable when outcomes are analyzed jointly with repeated noisy measurements of the network, such as paired reports of bilateral trade flows. For the joint model we establish necessary and sufficient conditions for local identification, estimators of the strength and composition paths, a simultaneous confidence band for the strength path, confidence sets that remain exact under weak identification, breakdown bounds under common reporting bias, and an exactly sized test that detects changes on the observed path and attributes them to strength or to composition. Simulations assess each procedure at its stated boundary. On a mirror-reported trade panel of eighteen economies over 1995 to 2020, the diagnostics flag exactly the crisis years and the composition coordinate attached to European Union membership declines by roughly two thirds. The estimand is predictive dependence, not a causal effect.2026-08-22T23:01:02ZMarios PapamichalisRegina RuaneTheofanis Papamichalishttp://arxiv.org/abs/2509.03945v3Prob-GParareal: A Probabilistic Numerical Parallel-in-Time Solver for Differential Equations2026-08-22T11:55:43ZWe introduce Prob-GParareal, a probabilistic extension of the GParareal algorithm designed to provide uncertainty quantification for the Parallel-in-Time (PinT) solution of (ordinary and partial) differential equations (ODEs, PDEs). The method employs Gaussian processes (GPs) to model the Parareal correction function, in line with GParareal, further enabling the propagation of numerical uncertainty across time and yielding probabilistic forecasts of the system's evolution. Furthermore, Prob-GParareal accommodates probabilistic initial conditions and maintains compatibility with classical numerical solvers, ensuring its straightforward integration into existing Parareal frameworks. Here, we first conduct a theoretical analysis of the computational complexity and derive error bounds of Prob-GParareal. Then, we numerically demonstrate the accuracy and robustness of the proposed algorithm on five benchmark ODE systems, including chaotic, stiff, and bifurcation problems. To showcase the flexibility and potential scalability of the proposed algorithm, we also consider Prob-nnGParareal, a variant obtained by replacing the GPs in Parareal with the nearest-neighbors GPs, illustrating its improved computational performance on an additional PDE example. This work bridges a critical gap in the development of probabilistic counterparts to established PinT methods.2025-09-04T07:09:59ZGuglielmo GattiglioLyudmila GrigoryevaMassimiliano Tamborrinohttp://arxiv.org/abs/2607.23280v2Uniqueness in multivariate curve resolution, re-tuned2026-08-22T05:29:52ZThere are many misunderstandings about the term and interpretation of uniqueness in multivariate curve resolution tasks. In CAC2026 Tarragona, it turned out that even mathematicians do not properly construe Manne's theorems. So we decided to provide this re-tuned summary, using several original quotes, their explanations, and newly created illustrative figures to get the points across.
This contribution then explicitly offers the following: (i) a unified interpretation of Manne's conditions, DBU, GRU, minimal constrained duality, and particular-solution-based criteria; (ii) newly constructed visual illustrations, including a counterexample; and (iii) a practical workflow for diagnosing uniqueness in nonnegative bilinear MCR.2026-07-25T16:35:13ZRobert RajkoRabeea Razaqhttp://arxiv.org/abs/2112.03738v3A more efficient algorithm to compute the Rand Index for change-point problems2026-08-22T03:45:38ZWe provide a more efficient algorithm for computing the Rand Index when the data clusters come from a change-point detection problem. Given the number of data points $N$ and two change-point sets of size $r$ and $s$, the algorithm runs on $O(r+s)$ time complexity and $O(1)$ memory complexity. The Rand Index computation for the general clustering problem, in contrast, requires the $N$ cluster memberships and has a $O(N)$ complexity in both time and memory.2021-12-07T14:52:04ZLucas de Oliveira Prateshttp://arxiv.org/abs/2609.27924v1A Prediction--Correction Analysis of Two-Way Block Splitting in Distributed Learning2026-08-22T03:05:36ZThis note revisits the convergence of the Parikh--Boyd two-way block-splitting algorithm for large-scale distributed learning through the He--Yuan prediction--correction framework. Simultaneous row--column partitioning is also relevant to hybrid federated learning, where data may be heterogeneous in both samples and features. We lift the reduced iteration to an equal-dimensional product space and reconstruct the primal and dual coordinates omitted by its implementation. The induced orthogonal-complement structure establishes exact iteration-by-iteration equivalence with the published updates. A mixed variational-inequality representation then yields a fundamental descent inequality, global convergence, an ergodic complexity bound, and current-iterate residual estimates under standard convexity, solvability, exact-subproblem, and invariant-initialization assumptions. The analysis also shows that the reduced state recursion is a metric proximal point iteration. No strong convexity, differentiability, or full-rank condition is imposed. The derivation clarifies which algebraic initialization conditions allow the reduced implementation to inherit the full-space convergence and complexity guarantees without modifying its local updates.2026-08-22T03:05:36ZXiaofei Wuhttp://arxiv.org/abs/2608.21654v1$K$-functions for point processes on complex surfaces2026-08-21T21:50:51ZThe $K$-function is a fundamental summary statistic for assessing clustering or regularity of point processes in two or three dimensional Euclidean space. In practice, however, many planar point patterns arise from projecting locations of objects on a surface in three dimensional space to two dimensional space. For example, when events or objects occur in a landscape their elevation is often ignored. This can lead to erroneous conclusions regarding properties of the point process generating the point pattern.
There is not a unique way to extend the classical $K$-function to point patterns on a complex surface. In this paper we propose, explore, and discuss several approaches in terms of their theoretical and computational properties. The best performing approach, coined the surface area $K$-function, can be viewed is an analogue of the classical $K$-function replacing counts of points in Euclidean balls with counts of points in surface geodesic balls. However, an important distinction is that the argument of our surface area $K$-function is area instead of radius of geodesic balls. The performances of the various surface $K$-functions are compared in applications to simulated and real data.2026-08-21T21:50:51ZFrancisco Cuevas-PachecoScott WardEd CohenRasmus Waagepetersenhttp://arxiv.org/abs/2608.21551v1Sparse Separable Factor Analysis in the Complex Domain with an Application to Local Field Potential Data2026-08-21T18:39:10ZComplex-valued arrays arise in signal processing, where scientific interpretation depends on retaining amplitude and phase information. Existing covariance estimation methods either ignore the multiway organization of such data or rely on real-domain embeddings that do not directly exploit their complex structure. We develop sparse separable factor analysis (SSFA), a latent factor model for complex-valued arrays with a separable covariance structure across modes. Each mode-specific covariance matrix is modeled through a low-rank Hermitian factor structure and a diagonal residual covariance matrix. To obtain interpretable estimates, we impose elementwise lasso penalties on the complex loading matrices and estimate the SSFA parameters using a mode-wise parameter-expanded expectation-maximization procedure. The resulting loading updates admit closed-form complex soft-thresholding solutions, which shrink the modulus of each loading while preserving its phase. A separate balancing step resolves the scale nonidentifiability of the separable covariance structure. Simulation studies show that SSFA improves covariance estimation relative to vectorization-based methods, including complex principal component analysis. We apply SSFA to local field potential recordings from mice, where we compare separability structures induced by different groupings of brain region, frequency, and time and perform model-based imputation of recordings missing because of electrode misplacement.2026-08-21T18:39:10Z50 pages, 9 figures, and 7 tablesIan HultmanKirtikanth KalapatapuYassine FilaliRainbo HultmanSanvesh Srivastavahttp://arxiv.org/abs/2608.21287v1Comprehensive Regression and Diagnostics for Non-Negative Data Using the BCSreg Package2026-08-21T16:43:06ZContinuous positive data characterized by high skewness and heavy tails frequently arise in applied statistics. In other applications, these characteristics are accompanied by a point mass at zero, resulting in a non-negative response with a mixed discrete-continuous distribution. Standard regression models often fail to capture these complex features adequately, requiring more flexible approaches. In this paper, we introduce the BCSreg package for R, which provides a comprehensive and unified computational framework for fitting Box-Cox symmetric and log-symmetric regression models for positive continuous data and their zero-adjusted extensions for mixed non-negative data. These broad classes of models accommodate varying degrees of skewness and tail-heaviness while allowing the parameters to be interpreted directly on the original scale of the data. Through a user-friendly multi-part formula interface, the BCSreg package allows practitioners to simultaneously specify regression structures for the scale parameter (which is proportional to the quantiles of the response), the relative dispersion, and, when appropriate, the probability of zero occurrences. Furthermore, the package provides a complete suite of diagnostic tools specifically tailored to these classes of models, including randomized quantile residuals, simulated envelopes, and influence diagnostics. The package's features and capabilities are illustrated through applications to real data.2026-08-21T16:43:06ZFrancisco F. QueirozRodrigo M. R. de Medeiroshttp://arxiv.org/abs/2608.21122v1LABS: Extending the scope of binary segmentation via a look-ahead device2026-08-21T14:01:01ZBinary segmentation is widely used for multiple change-point detection because it is fast, simple to describe, and simple to implement. Its validity rests on the requirement that, at each recursive stage, the procedure identifies one of the true change-points when several are present in the current interval. This holds for detecting changes in mean using the CUSUM statistic, but fails in some other settings, in particular in slope change detection for continuous piecewise-linear signals. We propose Look-Ahead Binary Segmentation (LABS), a modification in which the change-points returned by the two child recursions define a narrower interval on which the parent estimate is re-evaluated. LABS inherits the computational speed of standard binary segmentation, but achieves the near-optimal consistency rate of $O\{(n\log n)^{1/2}\}$ in the slope-change signal setting when the LABS model is chosen via either thresholding or a Schwarz-like information criterion. Simulations show that LABS is fast and achieves state-of-the-art performance.2026-08-21T14:01:01ZPiotr Fryzlewiczhttp://arxiv.org/abs/2607.09431v2comprisk: A scikit-learn-compatible Python toolkit for competing-risks survival analysis2026-08-21T12:18:01ZMedical time-to-event data are frequently subject to competing risks, where the occurrence of one terminal event precludes the others and standard survival methods that treat competing events as censoring yield biased absolute-risk estimates. Valid analysis instead targets the cause-specific cumulative incidence function (CIF). This methodology has been available to applied researchers almost exclusively through R packages, forcing Python-based machine-learning workflows into a Python-to-R round trip. We present comprisk, a scikit-learn-compatible Python toolkit that puts the canonical competing-risks methods behind one API: a scalable competing-risks random survival forest, Fine-Gray subdistribution-hazard regression and a penalized variant, cause-specific Cox regression, the Aalen-Johansen CIF estimator, and Gray's K-sample test, together with competing-risks-aware model evaluation. Every estimator is validated numerically against its R reference implementation. The forest uses a histogram-based, numba-compiled split kernel that fits 10-22x faster than randomForestSRC at comparable discrimination on real clinical cohorts and scales to n = 10^6 on a consumer CPU. comprisk is distributed on PyPI and lets applied researchers run correct, scalable competing-risks analysis without leaving the Python scientific stack.2026-07-10T13:57:13Z5 pages. Software paper; package available on PyPI (pip install comprisk). Code archived on Zenodo: 10.5281/zenodo.19876282Sunny YangWeiyan ZhaoWanqi Zhaohttp://arxiv.org/abs/2609.27885v1A Uniformly Efficient Rejection Sampler for the Multivariate Expectile-Based Distribution2026-08-21T08:25:11ZThe multivariate expectile-based distribution is a `spiked' perturbation of a multivariate Gaussian distribution, proposed by Arbel et al. in 2023. The rejection sampler proposed for this distribution in the initial work is valid, but its acceptance probability can degenerate badly both in high dimension and in the strongly asymmetric limit. We give a simple alternative. After affine whitening and a polar decomposition, the sampling problem reduces exactly to a univariate distribution, and an additional hyperbolic change of variables exhibits this distribution as log-concave, so that a universal and uniformly efficient construction of Devroye applies. The resulting exact sampler therefore enjoys an acceptance probability of at least $1 - \exp \left( -1 \right) = 0.632120 \ldots$, uniformly over the dimension and all admissible asymmetry parameters.2026-08-21T08:25:11Z5 pages, no figuresSam Powerhttp://arxiv.org/abs/2608.21466v1Spectral partitioning for $k$-block averaging kernels of finite Markov chains2026-08-21T02:53:38ZWe develop spectral algorithms for selecting state-space partitions that define averaging kernels for finite, ergodic and reversible Markov chains. For a partition $\mathcal O$, the Gibbs kernel $G_{\mathcal O}$ resamples within the current block from the stationary conditional distribution; when this update is tractable, composing or mixing it with a baseline kernel $P$ can accelerate convergence. We select $\mathcal O$ by rounding the bottom nonconstant eigenfunctions of $P^2$, or the algebraically smallest eigenfunctions of $P$ for additive mixtures, using weighted $k$-means. For $F(\mathcal O)=\|G_{\mathcal O}P-Π\|_{F,π}^2$, we derive exact trace and normalized-cut representations and show that $F$ equals the Pearson $χ^2$-mutual information between the initial block label and the state after one transition, giving this matrix objective a natural probabilistic interpretation. In the two-block case, a threshold sweep exactly solves the associated one-dimensional weighted two-means rounding problem. For general $k \geq 2$, weighted $k$-means rounds the bottom $(k-1)$-dimensional embedding, after which candidates are rescored by $F$; the rounding distortion is a distance between subspaces that yields spectral approximation bounds. We extend the framework to additive mixtures, finite-horizon objectives, and discounted infinite-horizon objectives. In contrast to classical normalized spectral clustering, which uses top nonconstant modes to find low-flow persistent clusters, our method uses bottom modes to favor large normalized cross-block flow and rapid loss of block-label information. Experiments on a controlled-spectrum graph, a mean-field Ising model, and Bayesian variable selection show notable per-iteration improvements in convergence and statistical estimation.2026-08-21T02:53:38Z43 pages, 8 figuresMichael C. H. ChoiYoujia Wanghttp://arxiv.org/abs/2609.27880v1Bayesian Tensor Regression for Neuroimaging Data2026-08-21T01:45:40ZMultidimensional array data, or tensors, arise naturally in neuroimaging and other high-dimensional applications. We propose a parsimonious Bayesian tensor regression model for studies in which a brain image is the response and predictors are vector-valued covariates. The method extends Bayesian envelope dimension reduction to tensor responses, identifying material subspaces that contain regression information while removing variation that is immaterial to the predictors. This formulation leads naturally to a Tucker tensor decomposition and allows spatial dependence and multiple sources of uncertainty to be modeled jointly. We develop a computationally feasible Markov chain Monte Carlo algorithm based on Gibbs sampling and establish posterior consistency for the proposed model. Simulation studies demonstrate substantial gains in estimation accuracy and uncertainty quantification when meaningful dimension reduction is present. We apply the method to Human Connectome Project neuroimaging data to investigate associations between alcohol use and brain activity. The results illustrate the value of Bayesian tensor envelope regression for inference with high-dimensional, spatially dependent imaging responses.2026-08-21T01:45:40Z29 pages, 3 figures, 5 tables; includes theoretical results, simulation studies, and an application to Human Connectome Project neuroimaging dataZahra NajiMontserrat FuentesLiangsuo MaHossein Moradi Rekabdarkolaeehttp://arxiv.org/abs/2608.20633v1A new analysis of the randomly pivoted Cholesky algorithm2026-08-21T00:15:20ZThe randomly pivoted Cholesky algorithm is one of the leading methods for computing a low-rank approximation to a large positive-semidefinite matrix. However, while it consistently achieves accuracy comparable to or better than competing methods of its type in experiments, its theoretical analysis lags somewhat behind other methods. This paper closes this gap, proving that randomly pivoted Cholesky produces an approximation with expected error within a $1+\varepsilon$ factor of the optimal rank-$r$ approximation in $\mathcal{O}(r/\varepsilon + r\sqrt{\log r})$ steps. This result nearly matches the optimal complexity $Θ(r/\varepsilon)$ for any low-rank approximation method based on a partial Cholesky decomposition (also known as a column Nyström approximation). The paper also presents bounds on the randomly pivoted Cholesky trace and spectral-norm errors that hold with high probability. The mathematical argument is largely due to GPT 5.6-Sol (Pro), with some refinements by the author.2026-08-21T00:15:20Z12 pages, 1 figureEthan N. W. Epperlyhttp://arxiv.org/abs/2505.00877v3Particle Filter for Bayesian Inference on Privatized Data2026-08-20T23:00:26ZDifferential Privacy (DP) is a probabilistic framework that protects privacy while preserving data utility. To protect the privacy of the individuals in the dataset, DP requires adding a precise amount of noise to a statistic of interest; however, this noise addition alters the resulting sampling distribution, making statistical inference challenging. One of the main DP goals in Bayesian analysis is to make statistical inference based on the private posterior distribution. While existing methods have strengths in specific conditions, they can be limited by poor mixing, strict assumptions, or low acceptance rates. We propose a novel particle filtering algorithm, which features (i) consistent estimates, (ii) Monte Carlo error estimates and asymptotic confidence intervals, (iii) computational efficiency, and (iv) accommodation to a wide variety of priors, models, and privacy mechanisms with minimal assumptions. We empirically evaluate our algorithm through a variety of simulation settings as well as an application to a 2021 Canadian census dataset, demonstrating the efficacy and adaptability of the proposed sampler.2025-05-01T21:39:12ZYu-Wei ChenPranav SanghiJordan Awan