https://arxiv.org/api/7jFedv28FmyKaETxxj6Qur1c9xY2026-09-10T17:24:43Z242171515http://arxiv.org/abs/2609.08981v1Transformers as In-Context Samplers: From Closed-Form Diffusion to Estimation-Free Sampling2026-09-08T16:25:11ZA growing body of work establishes that large language models are not mere statistical memorizers, but are capable of in-context learning: performing inference at test time using only examples provided in the prompt, without any parameter updates. Prior theoretical work has shown that this capability extends to supervised learning tasks such as linear regression. We prove that in-context learning extends further to \emph{data generation}: frozen transformers can simulate iterative generative samplers from in-context samples. We first show that transformers can realize closed-form and smoothed closed-form diffusion samplers. The construction identifies a concrete generative role for softmax attention: it computes responsibility weights and weighted empirical averages, while feedforward layers implement Euler updates.
To empirically relate these constructions to pretrained language models, we study \emph{semantic-topic sampling}: prompts consisting of words drawn from a common semantic category, such as animals, foods, or cities. Across transformer layers, the normalized hidden states exhibit a two-stage geometry: they move toward a uniform spherical reference in intermediate layers and then return to structured, topic-dependent representations near the output. We further measure an interacting-particle energy on these hidden-state clouds and observe the same U-shape pattern. We then prove that transformers can approximate an energy-based sampler, constructing the same U-shape energy across the layers.2026-09-08T16:25:11ZArman AdibiAlireza JafariMohammad GhavamzadehHadi Daneshmandhttp://arxiv.org/abs/2609.08866v1Bayesian palaeoclimate reconstruction from zero-inflated count-compositional pollen data: A case study of Lago Grande di Monticchio in southern Italy2026-09-08T15:10:50ZBayesian palaeoclimate reconstruction from fossil pollen counts relies on a modern pollen-climate calibration data set to infer the pollen-climate relationships used to reconstruct past climates. While geographically large calibration data sets improve coverage of climate space and reduce unreliable extrapolation, they also introduce substantial heterogeneity, structural zeros, and complex pollen-climate relationships. We propose a Bayesian modular framework for palaeoclimate reconstruction from count-compositional pollen data that addresses these challenges, and provides coherent uncertainty quantification. The framework employs the zero-and-$N$-inflated multinomial logistic-normal distribution to describe the compositional pollen counts coupled with Bayesian additive regression tree priors to model the nonlinear effects and interactions among the climate covariates. Inference is formulated through a cut posterior distribution that modularises the analysis into forward and reconstruction modules. The forward module is fitted once using a large modern calibration data set and its posterior uncertainty is subsequently propagated to reconstruct climate variables from fossil pollen counts. For the reconstruction module, we develop and compare three inverse posterior sampling schemes. Simulation studies and empirical validation on the modern data set demonstrate that a combination of sampling importance resampling with a multiple imputation technique and a continuous uniform prior over the domain of the modern climate variables achieves the best predictive performance, with well-calibrated uncertainty quantification for the climate reconstruction. In our motivating case study, we further illustrate the proposed methodology by reconstructing a three-dimensional climate vector from fossil pollen records collected at Lago Grande di Monticchio in southern Italy.2026-09-08T15:10:50Z35 pages; 9 figuresAndré F. B. MenezesAndrew C. ParnellBrian HuntleyKeefe Murphyhttp://arxiv.org/abs/2405.09989v4A Gaussian process model for chemoinformatics with application to the hazard classification of organic solvents2026-09-08T15:10:24ZWith the proliferation of screening tools for chemical testing, it is now possible to create vast databases of chemicals easily. However, rigorous statistical methodologies employed to analyse these databases are in their infancy, and further development to facilitate chemical discovery is imperative. In this paper, we address the challenge of predicting an organic solvent's water pollution class based on its chemical structure. We hypothesise that solvents with similar chemical structure are more likely to belong in the same hazard class. This chemical similarity is measured by the Tanimoto distance, a non-Euclidean metric on the chemical space. To incorporate the similarity between chemical structures in the model, we propose a Gaussian process model on the chemical space, with the kernel being a function of the Tanimoto distance. A novel feature of the proposed model is the inclusion of a scaling parameter in the kernel, which controls the strength of the correlation between compounds and offers additional flexibility. We find that accounting for correlation between chemical compounds substantially improves predictive performance over the uncorrelated model, where compound similarity is unaccounted for. Our model also compares favourably against other established models in the literature. Furthermore, we present a genetic algorithm designed to identify important features to a compound's efficacy, with the goal of facilitating chemical discovery. The algorithm operates based on two criteria derived from the proposed model. Simulation studies are conducted to demonstrate the suitability of the proposed methods.2024-05-16T11:18:32ZArron GosnellEvangelos Evangelouhttp://arxiv.org/abs/2609.08813v1Dynamic Latent Space Modeling of Inhomogeneous Poisson Network Processes with Applications to International Relations2026-09-08T14:39:10ZWe study continuous-time relational event data, where time-stamped dyadic interactions reflect both individual node propensities and evolving relational proximity. We propose a dynamic latent space model for inhomogeneous Poisson processes, where event intensities depend on node-specific activity parameters and time-varying latent distances modeled via flexible B-splines. We prove model identifiability by decoupling baseline activity from latent position, ensuring high interaction volumes do not warp the spatial map. For scalability, we develop a minibatch stochastic gradient algorithm with stable initialization and geometric anchoring, alongside an effective-degrees-of-freedom BIC for tuning model complexity. Simulations confirm accurate parameter recovery and out-of-sample prediction. Applied to cooperative diplomatic events among 60 major economies (1995--2022), the model uncovers shifting patterns of international cooperation and isolates mobile geopolitical actors from stationary institutional anchors.2026-09-08T14:39:10ZJie JianOwen G. WardJiguo Caohttp://arxiv.org/abs/2609.08797v1The Rater Ising-Potts Model with LLM-Derived Weights: An Application to Multi-Category Scoring Reliability2026-09-08T14:26:19ZThe Ising model is extended to the Potts model for multinomial data. We introduce a Rater Ising-Potts model that uses agreement indicators between pairs of raters and category labels, with weights derived from LLM embeddings. The model does not presuppose ordered category thresholds or equidistant scoring; instead, it focuses directly on pairwise agreement among raters and assigns category-specific positive weights, making it particularly suited for multi-category scoring reliability when raters evaluate responses using a scoring guide. We demonstrate the model's effectiveness on diverse constructed-response tasks, including balanced short-answer items and more challenging, imbalanced essay prompts from the AERA dataset. Across these settings, the model achieves strong agreement with human scores, with the vast majority of misclassifications occurring between adjacent score levels, confirming its ability to preserve the ordinal structure of scoring rubrics without imposing rigid assumptions. A practical similarity normalization and optional power transformation is introduced as a tunable preprocessing step that sharpens semantic distinctions and can be adapted to different datasets. These findings suggest that LLM-derived semantic similarities, combined with this parsimonious Potts-type formulation and flexible similarity scaling, offer a robust and interpretable framework for reliability auditing in educational assessment contexts. Extensions to multiple raters and hierarchical rating processes are discussed.2026-09-08T14:26:19ZMatthias von Davierhttp://arxiv.org/abs/2609.08780v1Structure-Informed Bayesian Inference of Anomalous Transport and Hidden Molecular Trapping in Amorphous Media2026-09-08T14:12:43ZMolecular diffusion in fluctuating amorphous and macromolecular media governs key transport processes across soft-matter physics, energy storage, and biological membranes. Extracting localized trapping states from single-particle tracking trajectories remains a fundamental challenge; because thermal structural breathing continuously reconfigures pore boundaries, conventional geometric algorithms suffer from severe systematic biases, erroneously merging distinct localized states during cyclic molecular returns. Here, we address this deadlock by shifting the paradigm from local geometric recurrence to a structure-informed Bayesian regularization. Leveraging discrete Morse theory, we extract the time-invariant topological skeleton of the fluctuating host matrix to construct robust, gas-specific physical priors that account for individual molecular dimensions. Trajectory steps are sequentially partitioned via a two-stage probabilistic refinement that dynamically adapts to the transport landscape. Benchmarked against a rigorous environment where synthetic particles explore the actual interconnected matrix graph, our approach eliminates systemic biases, restricting macroscopic trapping parameter deviations to just a few percent under optimal linear $O(N)$ computational scaling. Applied to hydrogen and methane transport within a type-I kerogen matrix, serving as a prototype for highly tortuous, flexible macromolecular networks, the method successfully decodes the hidden microscopic mechanisms of confined diffusion. To ensure immediate broad impact, the documented open-source code and data are made publicly available, offering an accessible strategy readily adaptable to a broad spectrum of tracking phenomena, from ion transport in battery polymers to protein trafficking within cellular environments.2026-09-08T14:12:43Z26 pages, 13 figures, 4 tables. Open-source code and reproducibility data are publicly available on GitHub and ZenodoAndrey AnanevMaria PotapovaNikolay KondratyukTimur VostroknutovAleksey Khlyupinhttp://arxiv.org/abs/2609.08577v1Instability in Patient Clustering: A Multiverse Analysis of Unsupervised Clustering in the CENTER-TBI cohort2026-09-08T11:12:59ZUnderstanding patient heterogeneity is key to improving prognostic modeling in traumatic brain injury (TBI). Unsupervised clustering is widely used to explore patterns in patient characteristics that may define subgroups. However, it involves a multitude of decisions, including the choice of algorithm, the distance metric, and the method used to determine the "optimal" number of clusters. The aim of this study is to investigate how these choices influence the resulting clustering solution.
We analyzed data from 4,509 patients enrolled in the Collaborative European NeuroTrauma Effectiveness Research in TBI (CENTER-TBI) study. K-medoids, agglomerative, and spectral clustering were applied in a complete 3 X 2 X 2 factorial design, in combination with Euclidean or Gower's distances, and silhouette score or gap statistic to choose the number of clusters. We investigated the agreement of clustering solutions with UpSet Plots and stability with the (adjusted) Rand index. Comparisons were made both across approaches using the original dataset and within approaches using bootstrap resampling.
Clustering results varied substantially depending on the analysis choices. The number of suggested clusters varied widely, from one to twenty-five. Adjusted Rand indices confirmed low concordance between methods. Moreover, none of the clustering solutions demonstrated discriminatory performance comparable to a supervised logistic regression model in classifying patient recovery illustrating the limited usefulness of clustering for this purpose. The high instability in clustering results compromises interpretability and underscores that such solutions should not be blindly interpreted as underlying structure.2026-09-08T11:12:59ZSean R E A BagcikAneeta Merlin ChackoEwout W SteyerbergMaarten van SmedenAndrew I R MaasErik van ZwetNicole S Erlerhttp://arxiv.org/abs/2609.03523v2Random mixtures in Bayes Hilbert spaces2026-09-08T08:35:19ZWe present a framework for the analysis and unmixing of random density mixtures in the Bayes Hilbert space. General identifiability results for mixtures in Hilbert spaces are established and applied to the Bayes Hilbert space setting. Building on these results, we propose a penalised maximum likelihood approach for the unmixing of Bayes Hilbert mixtures aimed at recovering the statistically space-efficient representation, together with a computationally efficient coordinate-wise maximisation algorithm for its implementation. The methodology is illustrated through a hyperspectral data application, where observations can be naturally embedded in the Bayes Hilbert space and analyzed in terms of distributional shape rather than amplitude. A complementary simulation study demonstrates the interpretability and practical performance of the proposed approach.2026-09-03T08:22:03ZGiulia PatanèSonja GrevenAlessandra Menafogliohttp://arxiv.org/abs/2609.08383v1GCMagicc v1: a fast generative emulator for multivariate climate-impact ensembles2026-09-08T07:54:07ZProjecting the impacts of climate change requires large ensembles of climate variables that match historical observations, align with the warming ranges assessed by the IPCC, and can efficiently run new future emissions scenarios, including the newest generation of climate model scenarios (CMIP7) and pathways consistent with countries' Paris Agreement pledges. Generating such ensembles at the scale needed for impact studies is normally computationally prohibitive. We close this gap with GCMagicc, a hybrid model that pairs a simple physical climate model with machine learning to generate ensembles of 10 climate variables at the resolution of full-scale Earth system models, without relying on GPU resources or retraining for new scenarios. Trained on 32 CMIP6 Earth system models and observational/reanalysis data, GCMagicc complements rather than replaces Earth system models. We apply it to a range of future pathways: the canonical SSP scenarios of the latest IPCC report (1.2-6.1°C warming, min-max across scenarios of 5-95 percentile ranges), current policies (2.3-4.0°C), national pledges under the Paris Agreement (1.5-3.3°C) and the CMIP7 range from the 'VL' to 'H' scenarios (1.2-4.2°C), releasing a large public dataset. As an illustration, we perform an attribution analysis of the severe 2025 Iranian drought using GCMagicc ensembles, with three CMIP6 large ensembles for comparison, with and without anthropogenic forcings. The results suggest a strong anthropogenic signal: a median probability of drought at least as severe as observed of 29% with anthropogenic forcing, and zero under natural-forcing-only simulations. In the future, drought conditions are projected to materially worsen, amplifying the potential for agricultural and food security impacts and geopolitical conflicts that use water scarcity as a weapon. GCMagicc data is available at https://gcmagicc.org.2026-09-08T07:54:07ZNicolai MeinshausenMalte MeinshausenJared LewisZebedee NichollsSarah SchöngartAlister SelfXinwei ShenKarla SpillerElisabeth Vogelhttp://arxiv.org/abs/2609.08060v1Pre-game paired-comparison modeling of professional League of Legends map outcomes2026-09-07T23:51:43ZWe build and evaluate a pre-game win-probability forecaster for individual maps (``games'') in professional \emph{League of Legends} (LoL). The proposed model is a one-stage logistic regression fit end-to-end on the win/loss log-loss: each team's exponentially-weighted moving average of past same-side results, a ridge-shrunk stable strength that is the maximum-a-posteriori estimate of a logistic mixed model, and a first-pick draft covariate, natively calibrated out of sample (walk-forward slope $0.995$). It augments a purely dynamic Bradley--Terry specification with stable team strengths. A second, independently built two-stage composite mixed model under restricted maximum likelihood (REML) and best linear unbiased prediction (BLUP) shrinkage, with Platt calibration, serves as the strongest rival the authors could build. On $5{,}135$ games across six regional leagues and three international events (2024--2026), under paired per-game Diebold--Mariano inference, the two architectures are statistically indistinguishable on every protocol and window (global holdout $0.2230$ vs.\ $0.2257$; walk-forward $0.2207$ vs.\ $0.2215$), so the simpler model is preferred on parsimony, not accuracy; both improve on the classical dynamic benchmark ($0.2351$) by a clear margin and on the static fits ($0.2301$/$0.2268$) more modestly. Against Polymarket on $928$ matched maps, the forecasts are statistically indistinguishable from the market on its own per-game contracts, with a modest market edge concentrated on cross-region Worlds and series-decider maps.2026-09-07T23:51:43Z26 pages, 2 figuresMin-Ren GuanShen-Ning Tunghttp://arxiv.org/abs/2609.07987v1When Can LLM Digital Twins Reduce Human Measurement? From Behavioral Fidelity to Statistical Substitutability2026-09-07T21:05:51ZLLM-based digital twins promise to reduce repeated human data collection by generating person- specific responses, yet existing evaluations provide little evidence about whether they can reduce human measurement while preserving valid inference. To address this, we introduce statistical substitutability, an inferential criterion that evaluates the extent to which twin predictions can reduce human measurement for a particular estimand while preserving valid inference. We develop a framework, grounded in mixed-subject and prediction-powered inference, that evaluates statistical substitutability along four dimensions: aggregate fidelity, paired respondent-level signal, finite-sample human-label recovery, and stability across populations. Across two empirical evaluations spanning behavioral experiments, multiple models, and alternative respondent representations, we find that digital twins can reproduce average human effects while providing little information about which individuals differ from those averages. Newer models and richer respondent information improve some dimensions of performance but do not reliably translate into human-data savings. Human calibration can reduce aggregate prediction error, yet limited labeled samples often fail to produce stable precision gains. Importantly, these findings demonstrate that behavioral fidelity is neither necessary nor sufficient for statistical substitutability. More broadly, they suggest that AI-generated evidence should be evaluated based on its ability to support valid scientific inference rather than its ability to reproduce human outcomes alone. Digital twins should therefore be judged for confirmatory use by whether they reduce uncertainty about human quantities, not merely by whether they reproduce human means, distributions, or effects.2026-09-07T21:05:51ZSteven WangKyle HuntShaojie TangKenneth Josephhttp://arxiv.org/abs/2509.12066v3On the universal calibration of heavy-tailed combination tests2026-09-07T19:59:45ZIt is often of interest to test a global null hypothesis using multiple, possibly dependent $p$-values by combining their strengths while controlling the type-I error. Recently, several heavy-tailed combination tests, such as the harmonic mean test and the Cauchy combination test, have been proposed: they transform $p$-values into heavy-tailed random variables before combining them into a single test statistic. The resulting tests, which are calibrated under some form of independence assumption among the $p$-values, have been shown to be rather robust to dependence asymptotically as the $α$ level gets small. Yet, it has remained an open problem to understand this general phenomenon and characterize how such tests behave under dependence. Using the framework of multivariate regular variation from extreme value theory, we show that for a class of combination tests that are homogeneous, the asymptotic level of the test can be expressed using the angular measure under multivariate regular variation. This measure characterizes the dependence of the transformed heavy-tailed variables in their upper tails, or equivalently, the dependence of the $p$-values near zero. We use this result to study several tests. The harmonic mean test, which coincides with the Pareto linear combination test, is shown to be universally calibrated regardless of the tail dependence; further, this test is shown to be the only one that achieves universal calibration among all homogeneous heavy-tailed combination tests. In contrast, the Cauchy combination test is shown to be universally honest but often conservative; the Dunn-Šidák correction, also known as Tippett's method, while being honest, is calibrated if and only if the underlying $p$-values are independent near zero. These theoretical findings are corroborated with simulations and an application to independence testing with survey data.2025-09-15T15:49:18Z7 figures, 49 pagesParijat ChakrabortyF. Richard GuoKerby SheddenStilian Stoevhttp://arxiv.org/abs/2602.16195v3Phase Transitions in Collective Damage of Civil Structures under Natural Hazards2026-09-07T19:28:20ZThe fate of cities under natural hazards depends not only on hazard intensity but also on the coupling of structural damage, a collective process that remains poorly understood. Here we show that urban structural damage exhibits phase-transition phenomena. As hazard intensity increases, the system can shift abruptly from a largely safe to a largely damaged state, analogous to a first-order phase transition in statistical physics. Higher diversity in the building portfolio smooths this transition, but multiscale damage clustering traps the system in an extended critical-like regime, analogous to a Griffiths phase, suppressing the emergence of a more predictable disordered (Gaussian) phase. These phenomenological patterns are interpreted through an effective random-field Ising model, with the external field, disorder strength, and temperature interpreted as the effective hazard demand, structural diversity, and modeling uncertainty, respectively. Applying this framework to real urban inventories reveals that widely used engineering modeling practices can shift urban damage patterns between synchronized and volatile regimes, systematically biasing exceedance-based risk metrics by up to 50% under moderate earthquakes ($M_w \approx 5.5$-$6.0$), equivalent to a several-fold gap in repair costs. This phase-aware description turns the collective behavior of civil infrastructure damage into actionable diagnostics for urban risk assessment and planning.2026-02-18T05:31:35ZSebin OhJinyan ZhaoRaul RinconJamie E. PadgettZiqi Wanghttp://arxiv.org/abs/2609.07833v1Neural Posterior Estimation for Tomographic Weak Lensing Mass Mapping2026-09-07T18:00:10ZWeak gravitational lensing shear and convergence trace the distribution of baryonic and dark matter across space, making them a powerful probe of cosmic structure. Inferring shear and convergence from images is a challenging inverse problem. The prevailing approach to this task estimates shear from weighted averages of galaxy ellipticities, calibrates these estimates to account for systematic biases, and transforms them to reconstruct convergence, a multistage procedure that requires substantial computational resources and meticulous handling of statistical uncertainties. As an alternative, we propose a probabilistic approach to field-level weak lensing inference in which we train a deep neural network to directly map a multiband image to a variational distribution over the underlying tomographic shear and convergence fields. This neural posterior estimation (NPE) procedure implicitly marginalizes over nuisance variables in the cosmological forward model and does not require evaluating the likelihood function. It is also amortized, so it enables rapid posterior inference for astronomical surveys once the neural network is trained. When evaluated on synthetic images from the LSST-DESC DC2 Simulated Sky Survey, NPE produces well-calibrated variational distributions for shear and convergence that are consistent with the ground truth. We describe how maps sampled from these variational distributions could be used in a subsequent simulation-based inference procedure to approximate the posterior distribution over cosmological parameters.2026-09-07T18:00:10Z17 pages, 10 figures, 1 tableTim WhiteShreyas ChandrashekaranCamille AvestruzJeffrey Regierthe LSST Dark Energy Science Collaborationhttp://arxiv.org/abs/2609.07794v1Two-resolution state modelling of intermittent online gambling activity during Sweden's temporary deposit-cap period2026-09-07T17:36:59ZIntermittent online gambling records create two distinct representation problems: calendar-time analyses must retain represented days with no recorded play, whereas analyses of active behaviour must preserve variation within active episodes. We pair calendar-time and active-bout models in a two-resolution analysis of 14.2 million player-days from a single operator during 2019-2023. The calendar-time layer retains assignable represented days in five observed categories and propagates a pre-period reference. The active-episode layer fits a Bayesian mixed-emission hidden semi-Markov model to turnover, playing time, approved deposits and session counts. In the reference four-regime fit, two regimes share a rounded model-implied turnover median of 153.5 EUR, while their model-implied playing-time and approved-deposit medians differ by factors of 4.6 and 5.7, respectively, and their expected total session counts by a factor of 2.8. This separation remains under an ordinary hidden Markov model, two routing perturbations and a residual-dependence sensitivity; the tested turnover-only fits were diagnostically unstable. Over the temporary Swedish deposit-cap period, deposits and turnover per represented player-day were roughly 55% below the propagated reference and active-category occupancy was 7.8% lower. The largest negative deviations in a prespecified Sweden-assigned cohort occurred among high pre-policy deposit-intensity customers. These are operator-specific descriptive comparisons, not causal policy effects.2026-09-07T17:36:59Z67 pages, including supplementary materialSam AnderssonTimo KoskiHelga WesterlindKeenan LyonPer CarlbringOlof Molander