https://arxiv.org/api/f6UzKilqBCWAZB1PdBCAhaRGqRY2026-09-10T20:14:17Z798806015http://arxiv.org/abs/2609.08581v1AlphaRJM: Reward-Jump Memory for Stochastic Return-Guided Alpha Discovery2026-09-08T11:17:07ZFormulaic alpha discovery is a pool-dependent symbolic search problem in which informative feedback is observed primarily when a complete expression is evaluated. This delayed feedback creates two coupled difficulties: the retained alpha pool does not preserve the full history of realized evaluation feedback, and the value of an intermediate construction action is uncertain because its consequence depends on the formula eventually completed. We introduce AlphaRJM, which addresses these difficulties through Reward-Jump Memory, an event-driven latent state that remains fixed during token construction and updates only at terminal evaluation events using the realized pool reward and evaluation outcome, and an action-conditioned SDE return critic that represents future discounted discovery returns with stochastic particles. The particles guide action selection through their mean and uncertainty and are learned using a distributional Bellman objective combining energy-distance matching, mean calibration, and jump regularization. Empirically, AlphaRJM delivers strong and stable gains across multiple equity universes, forecasting horizons, and random seeds, while ablations confirm the complementary roles of persistent evaluation history, stochastic return modeling, and distributional supervision.2026-09-08T11:17:07Z26 pages,5 Tables, 4 FiguresSayan DhanSelvaraju Natarajanhttp://arxiv.org/abs/2505.15215v4Clustering and Pruning in Causal Data Fusion2026-09-08T11:16:22ZData fusion, the process of combining observational and experimental data, can enable the identification of causal effects that would otherwise remain non-identifiable. Although identification algorithms have been developed for specific scenarios, do-calculus remains the only general-purpose tool for causal data fusion, particularly when variables are present in some data sources but not others. However, approaches based on do-calculus may encounter computational challenges as the number of variables increases and the causal graph grows in complexity. Consequently, there exists a need to reduce the size of such models while preserving the essential features. For this purpose, we propose pruning (removing unnecessary variables) and clustering (combining variables) as preprocessing operations for causal data fusion. We generalize earlier results on a single data source and derive conditions for applying pruning and clustering in the case of multiple data sources. We give sufficient conditions for inferring the identifiability or non-identifiability of a causal effect in a larger graph based on a smaller graph and show how to obtain the corresponding identifying functional for identifiable causal effects. Examples from epidemiology and social science demonstrate the use of the results.2025-05-21T07:44:39ZJournal of Machine Learning Research, 27(175):1-56, 2026Otto TabellSanttu TikkaJuha Karvanenhttp://arxiv.org/abs/2609.08577v1Instability in Patient Clustering: A Multiverse Analysis of Unsupervised Clustering in the CENTER-TBI cohort2026-09-08T11:12:59ZUnderstanding patient heterogeneity is key to improving prognostic modeling in traumatic brain injury (TBI). Unsupervised clustering is widely used to explore patterns in patient characteristics that may define subgroups. However, it involves a multitude of decisions, including the choice of algorithm, the distance metric, and the method used to determine the "optimal" number of clusters. The aim of this study is to investigate how these choices influence the resulting clustering solution.
We analyzed data from 4,509 patients enrolled in the Collaborative European NeuroTrauma Effectiveness Research in TBI (CENTER-TBI) study. K-medoids, agglomerative, and spectral clustering were applied in a complete 3 X 2 X 2 factorial design, in combination with Euclidean or Gower's distances, and silhouette score or gap statistic to choose the number of clusters. We investigated the agreement of clustering solutions with UpSet Plots and stability with the (adjusted) Rand index. Comparisons were made both across approaches using the original dataset and within approaches using bootstrap resampling.
Clustering results varied substantially depending on the analysis choices. The number of suggested clusters varied widely, from one to twenty-five. Adjusted Rand indices confirmed low concordance between methods. Moreover, none of the clustering solutions demonstrated discriminatory performance comparable to a supervised logistic regression model in classifying patient recovery illustrating the limited usefulness of clustering for this purpose. The high instability in clustering results compromises interpretability and underscores that such solutions should not be blindly interpreted as underlying structure.2026-09-08T11:12:59ZSean R E A BagcikAneeta Merlin ChackoEwout W SteyerbergMaarten van SmedenAndrew I R MaasErik van ZwetNicole S Erlerhttp://arxiv.org/abs/1601.08057v5On the Geometric Ergodicity of Hamiltonian Monte Carlo2026-09-08T11:06:59ZWe establish general conditions under which Markov chains produced by the Hamiltonian Monte Carlo method will and will not be geometrically ergodic. We consider implementations with both position-independent and position-dependent integration times. In the former case we find that the conditions for geometric ergodicity are essentially a gradient of the log-density which asymptotically points towards the centre of the space and grows no faster than linearly. In an idealised scenario in which the integration time is allowed to change in different regions of the space, we show that geometric ergodicity can be recovered for a much broader class of tail behaviours, leading to some guidelines for the choice of this free parameter in practice.2016-01-29T11:13:46Z29 pages + supplement (included in arXival as Appendix), 1 figure. Corrigendum related to Theorem 5.14 and Corollary 2.3 included as Appendix DBernoulli 25(4A) (2019), 3109-3138Samuel LivingstoneMichael BetancourtSimon ByrneMark Girolami10.3150/18-BEJ1083http://arxiv.org/abs/2503.24022v3Information Geometry for Wasserstein KL Divergence of Gaussian Measures on $\mathbb{R}^n$2026-09-08T10:59:17ZWe study the Wasserstein Kullback--Leibler divergence (WKL divergence) on the manifold of nondegenerate Gaussian measures over $\mathbb R^n$. In the canonical-divergence construction, the Fisher--Rao metric recovers forward KL along intrinsic mixture geodesics and reverse KL along geodesics of the conjugate exponential connection. Replacing the Fisher--Rao metric by the Otto metric and following the latter route produces the $e_1$-connection underlying WKL divergence. We establish its geodesic completeness, classify its forward limits, prove that every ordered pair is joined by a unique $e_1$-connector generated by a quadratic potential, and derive an explicit WKL divergence formula with separate mean and covariance contributions. WKL divergence is nonnegative and separating. At equal covariances WKL divergence equals one half of the squared Euclidean mean distance. Finally, WKL divergence extends finitely and continuously to singular targets from a nondegenerate source, but diverges when the target remains nondegenerate and the source covariance becomes singular. Path-dependent joint limits at Dirac pairs preclude a continuous extension to the full positive-semidefinite covariance product, although a lower-semicontinuous extended-real extension exists.2025-03-31T12:49:01ZAdwait DatarNihat Ayhttp://arxiv.org/abs/2609.08564v1Non-Adaptive 1-Bit Mean Estimation: Minimax Rates and the Sample-Interval Tradeoff2026-09-08T10:52:52ZWe study distributed one-dimensional mean estimation under a 1-bit communication constraint. Each agent observes one sample, drawn independently from an unknown distribution, and returns a single bit in response to a query $Q: \mathbb{R}\to\{0,1\}$ chosen by a central learner. The distribution has mean in $[-λ,λ]$ and $k$-th central moment at most $σ^k$, for a fixed $k>1$. The order-optimal two-stage protocol of Lau and Scarlett uses responses from the first batch to choose the second-batch queries, motivating the question of whether this single round of interaction is necessary. We answer this negatively: for every $k>1$, a non-adaptive protocol attains the adaptive 1-bit minimax rate (and concurrent works reached the same conclusion via different strategies). We further determine the minimax sample complexity among non-adaptive 1-bit estimators when every one-set $Q^{-1}(1)$ is restricted to a union of at most $s$ intervals. Relative to unrestricted non-adaptive 1-bit querying, this constraint adds a term of order $(λσ/(s\varepsilon^2))\log(1/δ)$, giving the full tradeoff between sample complexity and interval complexity to within $k$-dependent constant factors. As a corollary, we identify, order-wise, the minimum interval budget needed to retain the unrestricted 1-bit minimax sample rate.2026-09-08T10:52:52ZIvan LauJonathan Scarletthttp://arxiv.org/abs/2609.09245v1What Fixed-Rollout pass@k Evaluations Can Identify2026-09-08T10:27:52ZRepeated-sampling evaluations increasingly extrapolate pass@k far beyond the number n of samples collected per problem. We show that, in the pooled/random-task conditional-Binomial model, fixed-n success counts identify only the n free moments of the latent per-task success distribution. Consequently, direct pass@k is identified for k <= n, but generic extrapolated pass@k, tail exponents, and tail constants are not identified for k > n, even with arbitrarily many exchangeable tasks at the same rollout budget. This is stronger than the observation that the usual estimator is undefined beyond n: it characterizes the information missing from the fixed-depth count-law experiment. We give exact count-law-preserving constructions with incompatible extrapolations, state the exceptional unique-extension case, and compute sharp population identified intervals through Hausdorff principal representations. On the public 10,000-rollout-per-problem release of Brown et al., counterfactual n = 16 evaluations leave failure at k = 1000 ambiguous by factors from 1.5 to over 2,600 across four MATH/GSM8K/CodeContests configurations. The calibration shows that intermediate-scale failure share alone does not determine width. Our result does not reject parametric inference-time scaling laws; it supplies the nonparametric baseline against which their assumptions can be evaluated. We give an exact, conservative one-coordinate finite-task confidence certificate and a reporting standard separating direct estimates, identified sets, and model-conditioned forecasts.2026-09-08T10:27:52Z13 pages, 3 figures, 7 tablesPranav SinghPrashant Singhhttp://arxiv.org/abs/2609.08537v1The Exact Time-Uniform Rate Frontier for Stochastic Gradient Descent on Smooth Convex Objectives2026-09-08T10:24:19ZWe study the time-uniform convergence of the raw iterate of standard stochastic gradient descent (SGD) for unconstrained smooth convex objectives. We prove that, under standard noise assumptions, the time-uniform convergence rate gets arbitrarily close to $\sqrt{\log n / n}$ but never reaches it. More specifically, we prove that for every positive, eventually nondecreasing sequence $h$ satisfying $h(n) = o(\sqrt{n})$, a bound of order $h(n)/\sqrt{n}$, holding simultaneously for all $n$ with probability at least $1-α$ and uniformly over the problem class, is achievable if and only if \[
\sum_{j = 1}^{\infty} \frac{1}{h(2^j)^2} < \infty. \] The constructive sufficiency result follows from a dyadic horizon-free schedule together with an additive conditional-restart inequality. The necessity counterpart applies to every deterministic nonnegative schedule and holds even for a one-dimensional analytic smooth convex objective with Gaussian noise.2026-09-08T10:24:19Z28 pages, including appendixRuijie LiKang ChenTianyu Wanghttp://arxiv.org/abs/2609.09244v1Critical initialization destabilizes higher input derivatives in wide scalar-input networks2026-09-08T10:17:45ZThe edge-of-chaos condition preserves first-order input perturbations in wide randomly initialized networks, but physics-informed losses, score matching and derivative regularization depend on higher input derivatives. For smooth scalar-input fully connected networks, using a joint Gaussianity of the finite derivative jet that holds in the infinite-width limit at each fixed depth, we derive mean-field recursions through third order that are exact at the variance fixed point, with finite-depth corrections that decay geometrically. At criticality, the first-derivative variance is depth-invariant, whereas the second-derivative variance grows linearly whenever the activation has nonzero curvature. The resulting third-order system closes on mean-field susceptibilities. For residual networks with branch scale L^{-1/2}, we prove that every fixed finite derivative order has uniformly bounded variance under explicit regularity assumptions. Simulations verify the critical growth laws, the residual bound, and the closed recursion. The results concern initialization, not trained-network performance.2026-09-08T10:17:45Z34 pages, 4 figuresPrashant SinghPranav Singhhttp://arxiv.org/abs/2410.16004v4Are Bayesian networks typically faithful?2026-09-08T09:20:23ZFaithfulness is a common assumption in causal inference, often motivated by the fact that the faithful parameters of linear Gaussian and discrete Bayesian networks are typical, and the folklore belief that this should also hold for other classes of Bayesian networks. We address this open question by showing that among all Bayesian networks over a given DAG, the faithful Bayesian networks are indeed `typical': they constitute a dense, open set with respect to the total variation metric. This does not directly imply that faithfulness is typical in restricted classes of Bayesian networks that are often considered in statistical applications. To this end we consider the class of Bayesian networks parametrised by conditional exponential families, for which we show that under regularity conditions, the faithful parameters constitute a dense and open set, the unfaithful parameters have Lebesgue measure zero, and the induced faithful distributions are open and dense in the weak topology. This extends the existing results for linear Gaussian and discrete Bayesian networks. We also show for nonparametric classes of Bayesian networks with uniformly equicontinuous and uniformly bounded conditional densities that the faithful Bayesian networks are open and dense in the weak topology. All these results also hold for Bayesian networks with latent variables, if faithfulness is only required to hold with respect to the latent projection. Finally, for the considered conditional exponential family parametrisations and nonparametric conditional density models, the topological properties of conditional independence imply the existence of a consistent conditional independence test. Together with the topological properties of faithfulness, this implies that sound constraint-based causal discovery algorithms like PC and FCI are consistent on an open and dense -- and hence `typical' -- set of Bayesian networks.2024-10-21T13:38:04ZPhilip BoekenPatrick ForréJoris M. Mooijhttp://arxiv.org/abs/2410.11500v2On Generalisation Error Bounds for Transformers2026-09-08T09:15:56ZIn this paper, we establish a collection of covering number bounds for linear function classes under various norm constraints on the inputs and matrices. We then combine these results with existing covering number bounds to derive improved estimates and, based on these estimates, develop generalization error bounds for single-layer Transformers. The resulting generalization bounds improve upon several existing results in the literature and, in particular, are independent of the input sequence length. Moreover, our generalization error bound decays at the rate $O(1/\sqrt{n})$, where $n$ denotes the sample size, thereby improving upon existing bounds that scale as $O((\log n)/\sqrt{n})$. Furthermore, our covering number analysis explicitly incorporates rank constraints on the underlying matrix classes, allowing us to characterize how low-rank structures affect the metric entropy and, consequently, the resulting generalization bounds for Transformer architectures.2024-10-15T11:14:04Z29 pagesLan V. Truonghttp://arxiv.org/abs/2605.18276v2Geometric Dictionary Learning of Dynamical Systems with Optimal Transport2026-09-08T08:22:13ZLearning dynamical systems through operator-theoretic representations provides a powerful framework for analyzing complex dynamics, as spectral quantities such as eigenvalues and invariant structures encode characteristic time scales and long-term behavior. However, dynamical operators are typically estimated independently for each system, preventing the discovery of shared structure across related dynamics. To address this limitation, we posit that related dynamical systems lie near a low-dimensional manifold in spectral operator space. Based on this hypothesis, we introduce DOODL (Dynamical OperatOr Dictionary Learning), a framework that learns a dictionary of characteristic spectral dynamics whose combinations approximate this manifold and yield compact, interpretable embeddings of individual systems. Beyond representation learning, DOODL enables fast and interpretable operator estimation from short and partially observed trajectories by constraining the estimation to the learned operator manifold. Experiments on metastable Langevin dynamics and turbulent plasma simulations demonstrate that DOODL scales to highly complex multiscale regimes while capturing characteristic spectral structure governing the dynamics rather than merely fitting trajectories, achieving errors one to two orders of magnitude lower than independent operator estimation methods in challenging low-data regimes.2026-05-18T12:10:35ZThibaut GermainSami ChemlalRémi FlamaryVladimir R. KosticKarim Lounicihttp://arxiv.org/abs/2609.08234v1Distribution-free inference on the number of changepoints2026-09-08T04:26:54ZSuppose we are given an ordered sequence of independent data whose distribution changes $K$ times at unknown locations, for some unknown $K \geq 0$. In this paper, we study the problem of performing distribution-free inference on $K$. First, we show an impossibility result: any distribution-free upper confidence bound on $K$ must be trivial and uninformative. Then, using conformal $p$-values, and under only the assumption that the data segments induced by the changepoints are exchangeable (within themselves) and mutually independent, we construct a finite-sample valid lower confidence bound on $K$, which we call the Conformal LOwer bound on Changepoint Count (CLOCC). We show that CLOCC is the only feasible way to provide a lower bound on $K$ under the stated assumptions, a property we refer to as its universality. We provide practical guidelines for choosing score functions that yield efficient and tight lower bounds. We evaluate CLOCC in several synthetic and real-data experiments, where it provides informative lower bounds on $K$, demonstrating its practical applicability.2026-09-08T04:26:54Z34 pages, 3 figures, 1 tableRohan HoreAaditya Ramdashttp://arxiv.org/abs/2602.04761v2Improved Dimension Dependence for Bandit Convex Optimization with Gradient Variations2026-09-08T04:12:30ZGradient-variation online learning has drawn increasing attention due to its deep connections to game theory and optimization. It has been studied extensively in the full-information setting, but is underexplored with bandit feedback. In this work, we focus on gradient variation in Bandit Convex Optimization (BCO) with two-point feedback. By proposing a refined analysis of the non-consecutive gradient variation, a fundamental quantity in gradient variation with bandit feedback, we improve the dimension dependence for both convex and strongly convex functions compared with the best known results (Chiang et al., 2013). Our improved analysis of the non-consecutive gradient variation also implies other favorable problem-dependent guarantees, such as gradient-variance and small-loss regret bounds. Beyond the two-point setup, we demonstrate the versatility of our technique by achieving the first gradient-variation bound for one-point bandit linear optimization over hyper-rectangular domains. Finally, we validate the effectiveness of our results in more challenging tasks such as dynamic and universal regret minimization, establishing the first gradient-variation dynamic and universal regret bounds for two-point BCO.2026-02-04T16:58:53ZICML 2026Hang YuYu-Hu YanPeng Zhaohttp://arxiv.org/abs/2609.08219v1Speed Limit for Information Acquisition in Stochastic Learning Dynamics2026-09-08T04:02:52ZNeural networks acquire internal representations through learning. In this work, we formulate stochastic gradient descent (SGD) as a Markovian stochastic process and derive a Fisher-information flow speed limit that bounds the rate at which trainable parameters can acquire information about latent variables in the data-generating process. The resulting inequality decomposes the information flow into drift and noise contributions, thereby quantifying the roles of deterministic learning forces and SGD-induced fluctuations from an information-theoretic perspective. We verify the bound in analytically tractable basis-function linear regression, where the information budget predicted by the bound reproduces the ordering and characteristic time scales with which different latent variables are encoded in the learned parameters. These results establish Fisher-information speed limits as a quantitative framework for diagnosing when and how different aspects of the data-generating mechanism are acquired during stochastic learning.2026-09-08T04:02:52Z14 pages, 4 figuresShuta KobayashiAndreas Dechant