https://arxiv.org/api/YKyXcefmptze+RztGix0dUngP4E2026-09-11T00:25:13Z7988012015http://arxiv.org/abs/2609.06382v1Recovering Weak Signals with Normalizing Flows2026-09-06T04:44:34ZIn many scientific disciplines, weak signals of interest are obscured by dominant nuisance signals that are several orders of magnitude stronger. Recovering these weak signals requires subtracting the dominant ones; however, this calibration process inherently distorts or partially suppresses the underlying signal of interest. To address this problem, we propose the use of normalizing flow models to reconstruct calibration-affected weak signals. By leveraging the statistical invariance of the target signals and assuming minimal initial suppression, our framework effectively recovers the lost signal components. We provide a comprehensive theoretical overview of this normalizing flow-based recovery method and demonstrate its efficacy using simulated data.2026-09-06T04:44:34ZSarod Yatawattahttp://arxiv.org/abs/2101.02307v4Directed mixed membership stochastic blockmodel2026-09-06T03:18:12ZMixed membership modeling for undirected networks has been extensively explored in network science over the past few years. Despite the substantial progress made for undirected cases, handling mixed membership structures in directed networks continues to pose substantial difficulties. To address this gap, we introduce the Directed Mixed Membership Stochastic Blockmodel (DiMMSB), a novel framework tailored for directed networks with overlapping communities. A key feature of DiMMSB is its ability to treat the row and column nodes of the adjacency matrix as distinct entities, each potentially following its own community organization. Building on this model, we develop DiSP, an efficient spectral procedure to estimate mixed memberships for both sets of nodes. Through delicate analysis, we derive node-specific error bounds of DiSP under mild sparsity conditions. Simulation results support the theoretical results, demonstrating that DiSP achieves lower error rates and faster computation than its competitor. Moreover, applications to real data highlight DiSP's effectiveness in uncovering asymmetric structural patterns.2021-01-07T00:21:50ZQing H, Wang J. Directed mixed membership stochastic blockmodel. Information Sciences. 2026 Apr 30:123577Huan QingJingli Wang10.1016/j.ins.2026.123577http://arxiv.org/abs/2310.18727v2Latent class analysis by regularized spectral clustering2026-09-06T03:04:30ZThe latent class model is a highly effective tool in the analysis of categorical data from social, psychological, and behavioral sciences, where populations often share hidden common characteristics. In this article, we introduce two new algorithms for estimating the parameters of a latent class model for ordered categorical data with polytomous responses. These algorithms are based on a newly defined regularized Laplacian matrix derived from the response matrix. We provide theoretical convergence rates for our algorithms by considering a sparsity parameter and demonstrate that under a mild condition on data's sparsity, our algorithms yield consistent latent class analysis. Furthermore, we introduce a metric to assess the strength of latent class analysis and develop procedures based on this metric to determine the optimal number of latent classes for real-world ordered categorical data. Extensive simulation experiments demonstrate the efficiency and accuracy of our algorithms, and we demonstrate their practical application to real-world ordered categorical data with promising results.2023-10-28T15:09:08ZComput Stat 41, 92 (2026)Huan Qing10.1007/s00180-026-01772-0http://arxiv.org/abs/2609.09211v1A Subsampled Davis-Kahan Bound for Large-Scale Eigenspace Estimation2026-09-06T02:31:45ZThe Davis-Kahan theorem is a fundamental tool in spectral analysis, providing quantitative control over the distance between the eigenspaces of a symmetric matrix and its perturbation. However, when the matrix dimension is large, computing leading eigenvectors is computationally expensive, limiting the practical use of spectral methods in modern large-scale applications. This paper addresses this problem by proposing an independent Bernoulli sampling scheme and proves that the leading left singular vectors of the subsampled matrix faithfully approximate the target subspace of a low-rank symmetric matrix. Our main result is a subsampled Davis-Kahan bound that gives an explicit error bound depending directly on the sampling probability. The bound reveals the trade-off: the computational cost scales linearly with the sampling probability, while the statistical error scales as the inverse square root of the sampling probability. Our result thus extends the Davis-Kahan theorem to the subsampled setting, enabling scalable spectral analysis of large-scale symmetric matrices.2026-09-06T02:31:45Z24 pagesHuan Qinghttp://arxiv.org/abs/2509.15611v2Interpretable Network-assisted Random Forest+2026-09-06T02:23:29ZMachine learning algorithms often assume that training samples are independent. When data points are connected by a network, the induced dependency between samples is both a challenge, reducing effective sample size, and an opportunity to improve prediction by leveraging information from network neighbors. Multiple methods taking advantage of this opportunity are available, but many, including graph neural networks, are not easily interpretable, limiting their usefulness for understanding how models make predictions. Others, such as network-assisted linear regression, are interpretable but often yield worse prediction performance. We bridge this gap by proposing a family of flexible network-assisted models built upon a generalization of random forests (RF+), which achieves highly-competitive prediction accuracy and can be understood through intrinsic interpretability measures, derived directly from the model parameters and structure. In particular, we develop a suite of interpretation tools that enable researchers to both identify important features that drive model predictions and quantify the importance of the network contribution to prediction. Importantly, we provide global and local feature importances as well as sample influence measures to assess the impact of individual observations. This suite of tools broadens the scope and applicability of network-assisted machine learning for high-impact problems where interpretability and transparency are essential.2025-09-19T05:22:06ZTiffany M. TangElizaveta LevinaJi Zhuhttp://arxiv.org/abs/2607.02212v3An Additive MLP-GNN Framework for Characterizing Chemical and Structural Contributions to Aqueous Solubility2026-09-06T01:17:24ZAqueous solubility is a key property in early-stage drug discovery, but most predictive models merge physicochemical descriptors and molecular graph information into a single representation, obscuring whether a prediction is driven by global chemistry, molecular structure, or both. We present an additive deep-learning framework that keeps these two sources of information separate throughout training: physicochemical descriptors are encoded by a multilayer perceptron (the chemical branch) and molecular graph topology by a graph neural network (the structural branch), with the two outputs combined only at the prediction stage through an additive model with an optional multiplicative interaction. This design provides a direct decomposition of chemical and structural components that can be examined separately after training. Furthermore, pretraining on the larger AqSolDB dataset and fine-tuning on the smaller BigSolDB2 dataset substantially improve accuracy and reduce run-to-run variations, indicating generalizability of the learned features from the data-rich settings. We further interpret the fitted model using best linear projections of the branch outputs, molecule-level embedding summaries across solubility classes, and atom-level GNNExplainer masks aggregated over functional groups. These analyses show that the chemical branch aligns with familiar physicochemical descriptors, while the structural branch captures graph-topological and functional-group patterns associated with solubility. Across both datasets, the framework attains competitive predictive performance while making the distinct roles of chemical and structural information more transparent.2026-07-02T14:21:44ZSampreeti BhattacharyaArkaprava Royhttp://arxiv.org/abs/2601.21873v2Low-Rank Plus Sparse Matrix Transfer Learning under Growing Representations and Ambient Dimensions2026-09-06T00:29:26ZLearning systems often expand their ambient features or latent representations over time, embedding earlier representations into larger spaces with limited new latent structure. We study transfer learning for structured matrix estimation under simultaneous growth of the ambient dimension and the intrinsic representation, where a well-estimated source task is embedded as a subspace of a higher-dimensional target task.
We propose a general transfer framework in which the target parameter decomposes into an embedded source component, low-dimensional low-rank innovations, and sparse edits, and develop an anchored alternating projection estimator that preserves transferred subspaces while estimating only low-dimensional innovations and sparse modifications. We establish deterministic error bounds that separate target noise, representation growth, and source estimation error, yielding strictly improved rates when rank and sparsity increments are small.
We demonstrate the generality of the framework by applying it to two canonical problems. For Markov transition matrix estimation from a single trajectory, we derive end-to-end theoretical guarantees under dependent noise. For structured covariance estimation under enlarged dimensions, we provide complementary theoretical analysis in the appendix and empirically validate consistent transfer gains.2026-01-29T15:40:05ZThe authors have withdrawn this manuscript because substantial changes to its scope, formulation, and analysis are needed. The current version should not be citedJinhang ChaiXuyuan LiuElynn ChenYujun Yanhttp://arxiv.org/abs/2505.24281v2Multi-Task Learning with Covariate-Overlap Regularization2026-09-05T22:48:40ZMulti-task learning improves data efficiency by sharing information across related tasks, but indiscriminate sharing can be harmful when their covariate distributions and response relationships differ. We propose COVariate-ovERlap regularized multi-task learning (COVER) to address covariate and posterior heterogeneity. The model combines a common component function with a shared neural representation and low-dimensional task-specific coefficients. Taskwise second-moment matrices summarize covariate heterogeneity and determine the strength of coefficient integration in each representation direction. We derive a covariate-overlap penalty by minimizing the total squared change in two task predictors when their coefficients are replaced by one auxiliary coefficient. An equivalent auxiliary formulation supports end-to-end training without matrix inversion. An exact fixed-representation bias--variance decomposition quantifies how covariate overlap controls variance reduction and how posterior heterogeneity determines shrinkage bias. Global and localized end-to-end oracle inequalities account for jointly learning the neural functions and estimating the overlap matrices from the same observations. We give explicit neural-network rates and sharpen the stochastic prediction term when the regularized oracle risk and overlap-estimation error are small. Simulations across diverse heterogeneity settings show competitive performance against deep-learning and statistical data-integration methods, with the largest gains under joint heterogeneity. In a GTEx central-nervous-system analysis, COVER achieves the lowest response-averaged prediction error among the compared methods and reveals tissue-pair integration patterns.2025-05-30T06:58:42ZYang SuiQi XuYang BaiAnnie Quhttp://arxiv.org/abs/2609.06284v1Robust conditional dimension reduction for dissimilarity data2026-09-05T22:36:29ZConditional dimension reduction (cDR) learns low-dimensional latent coordinates while accounting for observed covariates that represent known sources of variation in the data. Conditional Multidimensional Scaling (cMDS) is a cDR technique that works directly with dissimilarity data. Its standard squared-stress formulation, however, is sensitive to contaminated dissimilarity, since outliers can dominate the objective and distort the learned configuration. We proposed Robust Conditional Multidimensional Scaling (rcMDS) by replacing the squared-stress criterion with a Fair M-estimation objective. We developed a reweighted conditional SMACOF algorithm to optimize this objective. The proposed algorithm admits computationally tractable updates, and its stabilized objective values decrease monotonically and converge to a finite limit. Experiments on synthetic and real data show that the pro2026-09-05T22:36:29ZXiao LingAnh Buihttp://arxiv.org/abs/2512.12749v3Residual-augmented flow matching operators for probabilistic partial differential equations2026-09-05T20:12:47ZLearning surrogate models for physical systems with latent uncertainty remains challenging in data-scarce regimes: deterministic neural operators fail to characterize uncertainty, while generative approaches require large ensembles of high-fidelity solution operator simulations and often sacrifice resolution generalizability. In this work, we propose a residual-augmented probabilistic operator learning framework that casts flow-matching-based generative modeling in infinite-dimensional function spaces while leveraging inexpensive low-fidelity solution operators as an inductive bias. Rather than learning the full high-fidelity stochastic solution operator directly, the proposed framework learns probabilistic residual operators that characterize the discrepancy between low- and high-fidelity solutions. By parameterizing the vector field in flow matching using neural operators conditioned on both the known system input and low-fidelity solution, the framework amortizes probabilistic inference across input conditions while enabling uncertainty-aware and resolution-generalizable predictions across spatial discretizations. Numerical experiments on stochastic advection, Burgers', and Darcy flow systems demonstrate that the residual-augmented formulation improves predictive accuracy under the same high-fidelity data budget, while the probabilistic operator learning formulation enables accurate characterization of uncertainty in low-data regimes compared to learning high-fidelity stochastic operators directly from data.2025-12-14T16:06:10ZSahil BholaKarthik Duraisamyhttp://arxiv.org/abs/2411.15876v3DUA-D2C: Dynamic Uncertainty Aware Method for Overfitting Remediation in Deep Learning2026-09-05T18:47:48ZOverfitting remains a significant challenge in deep learning, often arising from data outliers, noise, and limited training data. To address this, we previously proposed the Divide2Conquer (D2C) method, which partitions training data into multiple subsets and trains identical models independently on each. This strategy enables learning more consistent patterns while minimizing the influence of individual outliers and noise. D2C's standard aggregation typically treats all subset models equally or based on fixed heuristics (like data size), potentially underutilizing information about their varying generalization capabilities. Building upon this foundation, we introduce Dynamic Uncertainty-Aware Divide2Conquer (DUA-D2C), an advanced technique that refines the aggregation process. DUA-D2C dynamically weights the contributions of subset models based on their performance on a shared validation set, employing a novel composite score of accuracy and normalized prediction entropy. This intelligent aggregation allows the central model to preferentially learn from subsets yielding more generalizable and confident edge models, thereby more effectively combating overfitting. In this work, we provide a rigorous theoretical justification for this approach, analytically demonstrating how dynamic parameter fusion reduces model variance. Empirical evaluations on benchmark datasets spanning image, audio, and text domains demonstrate that DUA-D2C significantly improves generalization. Our analysis includes evaluations of decision boundaries, loss curves, and ablation studies, highlighting that DUA-D2C provides additive performance gains even when applied on top of standard regularizers like Dropout. This study establishes it as a theoretically grounded and effective approach to combating overfitting in deep learning. Our code is publicly available at: https://github.com/Saiful185/DUAD2C.2024-11-24T15:31:22ZThis version (v2) extends our previous work (arXiv:2411.15876v1) on Divide2Conquer (D2C) by introducing Dynamic Uncertainty-Aware Divide2Conquer (DUA-D2C). The manuscript has been published at Complex and Intelligent Systems. Find the published version at https://doi.org/10.1007/s40747-026-02251-1Complex Intell. Syst. 12, 149 (2026)Md. Saiful Bari SiddiquiMd Mohaiminul IslamMd. Golam Rabiul Alam10.1007/s40747-026-02251-1http://arxiv.org/abs/2307.02719v5Understanding Uncertainty Sampling via Equivalent Loss2026-09-05T18:14:40ZUncertainty sampling is a classical active-learning strategy, yet the statistical objective induced by its query rule is often implicit. We introduce the equivalent loss, whose gradient is the original loss gradient multiplied by the query probability. This construction places probabilistic, margin-based, and threshold-based uncertainty rules within a common framework. For binary classification, concrete equivalent losses and surrogate link functions show how uncertainty weighting preserves calibration while reshaping optimization geometry. When the equivalent loss is convex, we derive a finite-sample excess-risk bound with a fixed learning rate and an explicit constant controlling the tradeoff between initial error and query-weighted gradient variance. A feasible choice based only on maximal uncertainty yields a leading classification-risk upper bound no larger than the corresponding passive-learning bound at the same expected label budget; knowledge of the average query rate sharpens this comparison. We also analyze pool-based sampling, characterize the integrability obstruction beyond scalar prediction, and examine the query-clock dynamics of momentum methods. Together, these results provide a reusable route from a classical acquisition rule to its induced objective, statistical guarantees, and optimization behavior.2023-07-06T01:57:37ZAn updated version of the previous paper titled "Understanding Uncertainty Sampling." Corrected some gaps and typosShang LiuXiaocheng Lihttp://arxiv.org/abs/2609.06182v1Recovering linear images of sparse signals from indirect observations2026-09-05T16:57:31ZIn this paper, we develop and analyze techniques for recovering a linear image $Bx$ of an unknown signal $x$ from indirect noisy observation $ω=Ax+ξ$. It is {\em a priori} known that $x\in \cX$, a given convex compact set, and that $x$ is $s$-sparse---has at most $s$ nonvanishing entries. The proposed estimates belong to a large family of recovery routines by $\ell_1$-minimization. However, unlike the classical result describing performance of such estimates, we do not make any special (and hard to check) assumptions about the sensing matrix $A$ such as nullspace or Restricted Isometry condition and the like. As a consequence, parameters of the estimates and the upper bounds on their risks are not available in a closed analytic form, but are delivered instead by efficient computation as solutions to explicit convex optimization problems.2026-09-05T16:57:31ZAnatoli JuditskyArkadi Nemirovskihttp://arxiv.org/abs/2609.06098v1Causal DAG Identification for Count Data via Poisson Thinning Structural Equation Models2026-09-05T13:50:11ZCount-valued variables arise in many scientific and applied settings, yet explicit structural models that allow full identification of causal DAGs from observational data remain limited. The Poisson branching structural causal model (PB-SCM) provides a count-valued analogue of linear structural equation models using binomial thinning and independent Poisson exogenous variables, but its causal DAG is generally only partially identifiable.
Building on this framework, we propose the Poisson thinning structural equation model (PT-SEM), which replaces binomial thinning in PB-SCM with Poisson thinning and allows node-wise exogenous distributions from diverse count-distribution families. Under node-wise regularity conditions, we establish identifiability of the causal DAG, the thinning coefficients, and the node-wise exogenous distributions. The same identification analysis extends to binomial thinning, yielding full identifiability whenever every nonsink has non-Poisson exogenous noise. We further develop a structure learning algorithm that optimizes, via dynamic programming, a BIC score based on local likelihoods evaluated at plug-in moment estimates, and establish its consistency for DAG selection.
Simulations demonstrate favorable performance in DAG recovery and thinning-coefficient estimation, and a real-data application illustrates the practical utility of PT-SEM.2026-09-05T13:50:11Z35 pages, 6 figuresPenggang GaoMing CaiHisayuki Harahttp://arxiv.org/abs/2609.06064v1The Role of Gradient Modification in Heavy-Tailed Nonconvex Stochastic Min-Max Optimization2026-09-05T12:49:28ZStochastic min-max optimization has attracted increasing attention due to its applications in modern machine learning, while existing theoretical studies mainly rely on the bounded variance assumption for stochastic gradients. Under heavy-tailed noise, where stochastic gradients only possess a finite $p$-th moment for $p\in(1,2]$, gradient clipping or normalization is commonly believed to be necessary to guarantee convergence. In this work, we revisit stochastic min-max optimization under heavy-tailed noise and provide a comprehensive theoretical study of stochastic gradient descent ascent (SGDA). We first show that vanilla SGDA, without any modification to its update rule, can converge under heavy-tailed noise in both nonconvex-strongly-concave (NC-SC) and nonconvex-concave (NC-C) settings, establishing the first convergence guarantees for SGDA in these regimes. Beyond unregularized problems, we further investigate regularized stochastic min-max optimization, where directly incorporating gradient normalization into proximal updates is nontrivial due to the incompatibility between normalization and proximal structures. We overcome this difficulty by developing new clipping-free algorithms, i.e., Stoc-TRGDAM and Stoc-TRGDmax, and they both can achieve the optimal dependence on the target accuracy without using gradient clipping.2026-09-05T12:49:28ZTianxi ZhuYi XuXiangyang Ji