https://arxiv.org/api/o3UGu9oBFttxfMo2f65LmFMo1bE2026-09-11T17:47:52Z10572015http://arxiv.org/abs/2609.11539v1Extended One-Liners for the Gamma, Poisson, and Binomial Distributions2026-09-10T13:38:31ZWe demonstrate explicit transformations of three independent uniforms with exact Gamma$(a,1)$, $a>0$, Poisson$(λ)$, $λ>0$, and Binomial$(N,p)$, $N\geq 1$, $0\leq p \leq 1$ laws using only elementary operations.2026-09-10T13:38:31Z18 pagesDylan Greaveshttp://arxiv.org/abs/1810.04449v4Faster Hamiltonian Monte Carlo by Learning Leapfrog Scale: an offline randomized solution2026-09-10T13:03:16ZWe introduce a Hamiltonian Monte Carlo (HMC) methodology based on an offline empirical calibration of randomized leapfrog parameters. The approach, referred to as eHMC, where \textit{e} stands for empirical, leverages importance sampling to construct an empirical distribution on discretization parameters, thereby eliminating the need for manual burn-in diagnostics and online adaptation. The proposal distribution used in the calibration stage is obtained via a Population Monte Carlo scheme with tempering and relies on flexible parametric variational families such as normalizing flows. Once the calibration stage complete, the resulting algorithm defines a homogeneous Markov chain via a mixture of HMC kernels with a fixed mixing distribution, and hence preserves the target distribution. Numerical experiments indicate that eHMC can achieve competitive or improved sampling efficiency compared to the No-U-Turn Sampler (NUTS) in the case useful integration times can be summarized by the offline distribution. The comparison is assessed by standard efficiency metrics normalized by the number of leapfrog steps during the post-calibration sampling phase.2018-10-10T10:34:48ZCode available on https://github.com/jstoehr/eHMCChangye WuCEREMADEPierre PudloI2MChristian P. RobertCEREMADEJulien StoehrCEREMADEhttp://arxiv.org/abs/2609.11440v1Optimizing GEDI Simulator Configuration for European Temperate Forests2026-09-10T12:10:45ZAccurate estimation of aboveground biomass density is essential for quantifying forest carbon stocks. NASA's GEDI mission provides valuable canopy structure data, but its sparse sampling necessitates the use of simulators to calibrate biomass models at field inventory locations. The widely used simulator of Hancock et al. (2019) emulates GEDI waveforms from airborne LiDAR point clouds, yet it has never been validated over European temperate forests. Here, we compare approximately 9,500 pairs of observed and simulated GEDI relative height (RH) profiles across French forests using the national airborne LiDAR program as input. We separate two sources of error: waveform modeling differences, assessed by referencing both simulated and real RH metrics to a common ALS-derived ground elevation, and ground detection bias, evaluated by comparing each GEDI L2A processing algorithm against the ALS reference. Under the baseline configuration, the mean absolute bias across the full RH profile reaches 0.69 m in leaf-on and 1.28 m in leaf-off conditions. Switching to intensity-based return weighting and selecting the a3 L2A algorithm reduces these biases to 0.44 m and 0.40 m respectively. The a3 algorithm also achieves near-unbiased ground detection (-0.03 m versus -0.88 m for the default), directly reducing a previously overlooked source of error. We also show that leaf-off acquisitions and low-sensitivity shots, both typically excluded from standard biomass products, are simulated as reliably as their counterparts, substantially expanding the potential calibration and inference datasets.2026-09-10T12:10:45ZSubmitted to AGU Earth and Space Science (ESS). Presented at EGU General Assembly 2026, Vienna, Austria, abstract EGU26-9762. DOI: https://doi.org/10.5194/egusphere-egu26-9762Selim BehloulNikola BesicSteven HancockCedric VegaSylvie DurrieuJean-Pierre RenaudIbrahim FayadPhilippe Ciaishttp://arxiv.org/abs/2609.11278v1A unified framework for spatially resolved cortical activation analysis2026-09-10T09:12:28ZCluster-based permutation tests are widely used for analyzing MEG data, even though they are limited to cluster-level inference and do not provide spatially resolved effect estimates.
We propose a regression-based framework for modeling brain activity directly on the cortical surface. Spatial effects are represented using Wendland radial basis functions, and model-based gradient boosting is employed for data-driven selection and estimation of localized activation patterns. This yields interpretable, spatially resolved effect estimates while mitigating the need for extensive multiple testing correction. In a simulation study with heterogeneous signal structures, the proposed approach recovers localized effects that are difficult to detect using cluster-based methods. An application to experimental MEG data further illustrates its ability to reveal spatially specific activation patterns.
Overall, this unified framework provides a flexible and interpretable approach for modeling cortical surface data within a unified statistical setting.2026-09-10T09:12:28ZLars KnieperNadia Müller-VoggelTobias HeppAnna von PlessenElisabeth Bergherrhttp://arxiv.org/abs/2609.10865v1The Elliptically Optimal Confidence Interval: A Bivariate Extension of Wilson's Score Method2026-09-09T22:06:19ZConstructing a confidence interval for the difference between two independent binomial proportions involves a nuisance direction that is not identified by the estimand. The one-sample Wilson score interval inverts a scalar score test, but has no direct bivariate analogue isolating the difference: inverting the joint normal approximation yields an elliptical region in the unit square, whereas the estimand \(p_1-p_2\) is one-dimensional. We define the Elliptically Optimal (EO) confidence interval as the range of \(p_1-p_2\) over this region and solve the resulting optimization problem in closed form, obtaining explicit bounds in six mutually exclusive and exhaustive cases. The solution admits a compact characterization: the EO interval is the score interval obtained by maximizing over the nuisance variance rather than estimating it. It is therefore the shortest interval obtained by projecting the elliptical region, and inherits its coverage guarantee. We derive the exact coverage excess, \(2[Φ(z\mathcal{R})-Φ(z)]\), where \(\mathcal{R}\) is the ratio of the least-favourable to the true standard deviation. The excess vanishes on an explicit line through the parameter space, is bounded by \(α\), and is invariant under proportional scaling of the sample sizes. Exact enumeration of the binomial coverage shows that the Wald interval, whose variance estimator is downward biased by a factor \(1-1/n\) under balanced allocation, falls below nominal coverage almost everywhere. The EO interval never under-covers under the normal approximation and always yields admissible, non-degenerate bounds. Its price is over-coverage when both proportions are extreme, which we quantify exactly.2026-09-09T22:06:19ZNawaf Mohammedhttp://arxiv.org/abs/2309.15735v5Estimating MCMC convergence rates using common random number simulation2026-09-09T17:26:15ZThis paper presents how to use common random number (CRN) simulation to evaluate Markov chain Monte Carlo (MCMC) convergence to stationarity. We provide an upper bound on the Wasserstein distance of a Markov chain to its stationary distribution after $N$ steps in terms of averages over CRN simulations. We apply our bound to Gibbs samplers on a model related to James-Stein estimators, a variance component model, and a Bayesian linear regression model. Using our examples, we show that the CRN-based simulation combined with a coalescing condition to generate a total variation bound converges to zero much more faster than the available drift and minorization bounds, while also converging at the same rate as the one-shot coupling bound.2023-09-27T15:48:10ZSabrina SixtaJeffrey S. RosenthalAustin Brown10.1080/15326349.2026.2688991http://arxiv.org/abs/2609.10477v1Multivariate linear regression without prior assumptions2026-09-09T17:18:04ZRecovering the linear relationships that govern a system from noisy measurements is a basic task across the physical and engineering sciences. Because every measured variable may carry an unknown amount of noise, classical regression must commit in advance to a set of structural assumptions: ordinary least squares requires a declared input-output partition with input variables being noise-free, total least squares assumes equal noise variance across all variables, and generalized total least squares additionally requires the noisy-variable partition and variances to be known beforehand. Kalman~\cite{Kalman:1982} showed that any procedure returning a unique linear model from inexact data must rest on such unverifiable a priori assumptions -- ``prejudices'' -- that cannot be checked against the data itself, and that removing them leaves the identification problem fundamentally indeterminate. Whether these prejudices can instead be resolved directly from the data has remain unresolved. Here we show that an iterative generalized-eigenvalue algorithm, QZ-IPCA, recovers the noisy-variable partition, noise variances, number of linear relations, and regression coefficients of a multivariate linear system simultaneously, using only the raw data. Across all possible exhaustive noise configurations of a five-variable benchmark network, QZ-IPCA correctly identifies model structure and recovers coefficients with error below 6.4\%. It outperforms ordinary least squares even when given the best partition, and succeeds in rank identification precisely where standard total least squares falls once noise variances differ across variables. These results show that the assumptions conventionally required for multivariate regression are not necessary, recasting model identification as a problem solvable from data geometry alone.2026-09-09T17:18:04ZThe manuscript is under further processing. Important modifications might take place in further versions. There are 32 pages containing 10 figures and 3 tablesMayank S. K. GuptaDeepanjhan DasArun K. TangiralaShankar Narasimhanhttp://arxiv.org/abs/2609.09944v1Estimating Hierarchically Rank Structured Covariance Matrices2026-09-09T09:35:05ZWe consider the problem of estimating a high-dimensional covariance matrix from a very limited number of samples. This problem is ubiquitous in computational fluid dynamics, where a small number of fluid snapshots must be used to construct a Gramian matrix determining a reduced-order model, as well as in computational geoscience, where a small ensemble of Earth system forecasts must be used to estimate the covariance matrix associated with the forecast uncertainty. It is common practice to regularize the small-sample covariance by imposing a "localization" structure that enforces a physically realistic correlation length scale, imposing a sparsity constraint, "shrinking" towards a prescribed target, or attenuating small correlations. We propose an alternate technique that regularizes the small-sample covariance by imposing hierarchical rank structure. Compared to regularization methods that assume sparsity such as spatial localization, hierarchical rank structure accommodates a wider range of covariance matrices, roughly corresponding to situations where long-range correlations vary more smoothly than short-range ones. It also results in a data-sparse matrix format that permits highly efficient matrix-vector products. We present theory and algorithms which show how to efficiently estimate a high-dimensional, hierarchically rank structured covariance matrix from limited samples. Through an error analysis and numerical experiments with a variety of model problems, we demonstrate that these techniques are effective at reducing sampling errors, and that in many cases they achieve smaller estimation error than conventional techniques.2026-09-09T09:35:05ZRobin ArmstrongAnil DamleSamuel E. Ottohttp://arxiv.org/abs/2609.02557v3TrunX: A massively parallel, differentiable implementation of the 3-PG forest growth model in JAX2026-09-09T07:00:55ZProcess-based forest models are widely used to simulate forest growth and responses to environmental change, but their calibration and application often require many computationally expensive model evaluations. We present an implementation of the Physiological Processes Predicting Growth (3-PG) model in JAX that uses just-in-time compilation, vectorization, and GPU acceleration to reduce execution time. The implementation also supports automatic differentiation, providing gradients of model outputs and calibration objectives with respect to model parameters. This enables efficient gradient-based optimization and gradient-informed Bayesian calibration, extending 3-PG beyond conventional gradient-free approaches. The implementation produced results numerically consistent with r3PG for the evaluated configuration. Overall, the JAX implementation provides a faster and differentiable framework for calibrating and applying the 3-PG model.2026-09-02T13:09:57ZGlory Mary GiviCédric TravellettiGrégory Mermoudhttp://arxiv.org/abs/2603.25806v3Context Tree Prior Distributions based on Node Weighting with exact Bayes Factors2026-09-08T18:31:24ZVariable-length Markov chains (VLMCs) are a flexible class of higher-order Markov models that admit a natural representation as context trees. Existing Bayesian methods for specifying prior distributions on trees rely on branching processes, but these suffer from a fundamental limitation: the connection between node-branching probabilities and the structural properties of the induced tree distribution is not straightforward, making it difficult to encode specific structural beliefs. We address this issue by introducing a novel representation of prior distributions on tree spaces, characterized by assigning weights to individual contexts through a function on nodes. In this way, our approach provides an intuitive mechanism for incorporating structural hypotheses into the prior while preserving computational tractability, allowing marginal likelihoods and posterior mode trees to be computed exactly via generalizations of the Context Tree Weighting (CTW) and Context Tree Maximizing (CTM) algorithms. By enabling exact Bayes factor calculations, our methodology provides a principled framework for model comparison over structural priors. We demonstrate the flexibility and effectiveness of our approach by comparing different prior specifications through simulation studies and an application to financial markets.2026-03-26T18:12:06Z35 pages, 6 figuresThiago PaulichenVictor Fregugliahttp://arxiv.org/abs/2609.09053v1Quadratic Point Estimate Method for Uncertainty Quantification with Dependent Non-Gaussian Inputs2026-09-08T17:10:28ZAs an extension of the Point Estimate Method (PEM) to evaluate probabilistic moments of quantities of interest (QoI) in general $n$-dimensional spaces, the Quadratic Point Estimate Method (QPEM) has been recently developed. This new method is defined to fully represent up to fifth-order input moments in the Gaussian space, providing general analytical expressions for sample locations and weights, without requiring any numerical optimization. The QPEM can significantly improve the estimation accuracy of the output QoI moments, in relation to PEM-based methods whose numbers of sigma points grow linearly with the problem dimension, while at the same time having an affordable and competitive computational cost up to a considerable number of dimensions. The QPEM is further enhanced in this work by enabling copula integration into the framework, which enables effective modeling of the joint input probability density function by estimating marginals and the dependence structure of the involved random variables. The validity and efficient performance of the copula-based QPEM are showcased against numerous other sampling methods in various examples considering two practical scenarios: (i) when the joint dependence structure can be inferred from data, and (ii) when only marginal distributions and correlation matrices are known.2026-09-08T17:10:28ZMinhyeok KoKonstantinos G. Papakonstantinouhttp://arxiv.org/abs/2609.08981v1Transformers as In-Context Samplers: From Closed-Form Diffusion to Estimation-Free Sampling2026-09-08T16:25:11ZA growing body of work establishes that large language models are not mere statistical memorizers, but are capable of in-context learning: performing inference at test time using only examples provided in the prompt, without any parameter updates. Prior theoretical work has shown that this capability extends to supervised learning tasks such as linear regression. We prove that in-context learning extends further to \emph{data generation}: frozen transformers can simulate iterative generative samplers from in-context samples. We first show that transformers can realize closed-form and smoothed closed-form diffusion samplers. The construction identifies a concrete generative role for softmax attention: it computes responsibility weights and weighted empirical averages, while feedforward layers implement Euler updates.
To empirically relate these constructions to pretrained language models, we study \emph{semantic-topic sampling}: prompts consisting of words drawn from a common semantic category, such as animals, foods, or cities. Across transformer layers, the normalized hidden states exhibit a two-stage geometry: they move toward a uniform spherical reference in intermediate layers and then return to structured, topic-dependent representations near the output. We further measure an interacting-particle energy on these hidden-state clouds and observe the same U-shape pattern. We then prove that transformers can approximate an energy-based sampler, constructing the same U-shape energy across the layers.2026-09-08T16:25:11ZArman AdibiAlireza JafariMohammad GhavamzadehHadi Daneshmandhttp://arxiv.org/abs/2604.22391v2Conformalized Super Learner2026-09-08T15:28:39ZThe Super Learner (SL) is a widely used ensemble method that combines point predictions from a library of learners based on their predictive performance. Interval predictions are of considerable practical interest because they allow uncertainty in predictions produced by an individual learner or an ensemble to be quantified. Several methods have been proposed for constructing interval predictions based on the SL, however, these approaches are typically justified using asymptotic arguments or rely on computationally intensive procedures such as the bootstrap. Conformal prediction (CP) is a machine learning framework for constructing prediction intervals with finite-sample and asymptotic coverage guarantees under mild conditions. We propose coupling CP with the SL through a natural construction that mirrors the original SL framework, using individual learner weights and combining learner-specific conformity scores via a weighted majority vote. We characterize the properties of the resulting SL-based prediction intervals for continuous outcomes. We cover settings under exchangeability, potential violations of exchangeability, and data-generating mechanisms exhibiting heteroscedasticity, sparsity, and other forms of distributional heterogeneity. A comprehensive simulation study shows that the conformalized SL achieves valid finite-sample coverage with competitive performance relative to the true data-generating mechanism. A central contribution of this work is an application to predicting creatinine levels using socio-demographic, biometric, and laboratory measurements. This example demonstrates the benefits of an ensemble with carefully selected learners designed to capture key aspects of complex regression functions, including non-linear effects, interactions, sparsity, heteroscedasticity, and robustness to outliers.2026-04-24T09:28:46ZR codes and data can be found at: https://github.com/ZWU-001/CSLZhanli WuFabrizio LeisenMiguel-Angel Luque-FernandezF. Javier Rubiohttp://arxiv.org/abs/2609.08866v1Bayesian palaeoclimate reconstruction from zero-inflated count-compositional pollen data: A case study of Lago Grande di Monticchio in southern Italy2026-09-08T15:10:50ZBayesian palaeoclimate reconstruction from fossil pollen counts relies on a modern pollen-climate calibration data set to infer the pollen-climate relationships used to reconstruct past climates. While geographically large calibration data sets improve coverage of climate space and reduce unreliable extrapolation, they also introduce substantial heterogeneity, structural zeros, and complex pollen-climate relationships. We propose a Bayesian modular framework for palaeoclimate reconstruction from count-compositional pollen data that addresses these challenges, and provides coherent uncertainty quantification. The framework employs the zero-and-$N$-inflated multinomial logistic-normal distribution to describe the compositional pollen counts coupled with Bayesian additive regression tree priors to model the nonlinear effects and interactions among the climate covariates. Inference is formulated through a cut posterior distribution that modularises the analysis into forward and reconstruction modules. The forward module is fitted once using a large modern calibration data set and its posterior uncertainty is subsequently propagated to reconstruct climate variables from fossil pollen counts. For the reconstruction module, we develop and compare three inverse posterior sampling schemes. Simulation studies and empirical validation on the modern data set demonstrate that a combination of sampling importance resampling with a multiple imputation technique and a continuous uniform prior over the domain of the modern climate variables achieves the best predictive performance, with well-calibrated uncertainty quantification for the climate reconstruction. In our motivating case study, we further illustrate the proposed methodology by reconstructing a three-dimensional climate vector from fossil pollen records collected at Lago Grande di Monticchio in southern Italy.2026-09-08T15:10:50Z35 pages; 9 figuresAndré F. B. MenezesAndrew C. ParnellBrian HuntleyKeefe Murphyhttp://arxiv.org/abs/1601.08057v5On the Geometric Ergodicity of Hamiltonian Monte Carlo2026-09-08T11:06:59ZWe establish general conditions under which Markov chains produced by the Hamiltonian Monte Carlo method will and will not be geometrically ergodic. We consider implementations with both position-independent and position-dependent integration times. In the former case we find that the conditions for geometric ergodicity are essentially a gradient of the log-density which asymptotically points towards the centre of the space and grows no faster than linearly. In an idealised scenario in which the integration time is allowed to change in different regions of the space, we show that geometric ergodicity can be recovered for a much broader class of tail behaviours, leading to some guidelines for the choice of this free parameter in practice.2016-01-29T11:13:46Z29 pages + supplement (included in arXival as Appendix), 1 figure. Corrigendum related to Theorem 5.14 and Corollary 2.3 included as Appendix DBernoulli 25(4A) (2019), 3109-3138Samuel LivingstoneMichael BetancourtSimon ByrneMark Girolami10.3150/18-BEJ1083