https://arxiv.org/api/DEAktAycGsbamIwoPp4Tc2BNWic 2026-09-11T20:43:47Z 24226 45 15 http://arxiv.org/abs/2602.16195v3 Phase Transitions in Collective Damage of Civil Structures under Natural Hazards 2026-09-07T19:28:20Z The fate of cities under natural hazards depends not only on hazard intensity but also on the coupling of structural damage, a collective process that remains poorly understood. Here we show that urban structural damage exhibits phase-transition phenomena. As hazard intensity increases, the system can shift abruptly from a largely safe to a largely damaged state, analogous to a first-order phase transition in statistical physics. Higher diversity in the building portfolio smooths this transition, but multiscale damage clustering traps the system in an extended critical-like regime, analogous to a Griffiths phase, suppressing the emergence of a more predictable disordered (Gaussian) phase. These phenomenological patterns are interpreted through an effective random-field Ising model, with the external field, disorder strength, and temperature interpreted as the effective hazard demand, structural diversity, and modeling uncertainty, respectively. Applying this framework to real urban inventories reveals that widely used engineering modeling practices can shift urban damage patterns between synchronized and volatile regimes, systematically biasing exceedance-based risk metrics by up to 50% under moderate earthquakes ($M_w \approx 5.5$-$6.0$), equivalent to a several-fold gap in repair costs. This phase-aware description turns the collective behavior of civil infrastructure damage into actionable diagnostics for urban risk assessment and planning. 2026-02-18T05:31:35Z Sebin Oh Jinyan Zhao Raul Rincon Jamie E. Padgett Ziqi Wang http://arxiv.org/abs/2609.07833v1 Neural Posterior Estimation for Tomographic Weak Lensing Mass Mapping 2026-09-07T18:00:10Z Weak gravitational lensing shear and convergence trace the distribution of baryonic and dark matter across space, making them a powerful probe of cosmic structure. Inferring shear and convergence from images is a challenging inverse problem. The prevailing approach to this task estimates shear from weighted averages of galaxy ellipticities, calibrates these estimates to account for systematic biases, and transforms them to reconstruct convergence, a multistage procedure that requires substantial computational resources and meticulous handling of statistical uncertainties. As an alternative, we propose a probabilistic approach to field-level weak lensing inference in which we train a deep neural network to directly map a multiband image to a variational distribution over the underlying tomographic shear and convergence fields. This neural posterior estimation (NPE) procedure implicitly marginalizes over nuisance variables in the cosmological forward model and does not require evaluating the likelihood function. It is also amortized, so it enables rapid posterior inference for astronomical surveys once the neural network is trained. When evaluated on synthetic images from the LSST-DESC DC2 Simulated Sky Survey, NPE produces well-calibrated variational distributions for shear and convergence that are consistent with the ground truth. We describe how maps sampled from these variational distributions could be used in a subsequent simulation-based inference procedure to approximate the posterior distribution over cosmological parameters. 2026-09-07T18:00:10Z 17 pages, 10 figures, 1 table Tim White Shreyas Chandrashekaran Camille Avestruz Jeffrey Regier the LSST Dark Energy Science Collaboration http://arxiv.org/abs/2609.07794v1 Two-resolution state modelling of intermittent online gambling activity during Sweden's temporary deposit-cap period 2026-09-07T17:36:59Z Intermittent online gambling records create two distinct representation problems: calendar-time analyses must retain represented days with no recorded play, whereas analyses of active behaviour must preserve variation within active episodes. We pair calendar-time and active-bout models in a two-resolution analysis of 14.2 million player-days from a single operator during 2019-2023. The calendar-time layer retains assignable represented days in five observed categories and propagates a pre-period reference. The active-episode layer fits a Bayesian mixed-emission hidden semi-Markov model to turnover, playing time, approved deposits and session counts. In the reference four-regime fit, two regimes share a rounded model-implied turnover median of 153.5 EUR, while their model-implied playing-time and approved-deposit medians differ by factors of 4.6 and 5.7, respectively, and their expected total session counts by a factor of 2.8. This separation remains under an ordinary hidden Markov model, two routing perturbations and a residual-dependence sensitivity; the tested turnover-only fits were diagnostically unstable. Over the temporary Swedish deposit-cap period, deposits and turnover per represented player-day were roughly 55% below the propagated reference and active-category occupancy was 7.8% lower. The largest negative deviations in a prespecified Sweden-assigned cohort occurred among high pre-policy deposit-intensity customers. These are operator-specific descriptive comparisons, not causal policy effects. 2026-09-07T17:36:59Z 67 pages, including supplementary material Sam Andersson Timo Koski Helga Westerlind Keenan Lyon Per Carlbring Olof Molander http://arxiv.org/abs/2609.07683v1 Modeling Brain MRI Using Persistent Homology and Multilevel Functional Data Analysis 2026-09-07T16:05:59Z Persistent homology provides a multiscale representation of biomedical images by capturing higher-order topological features that reflect their underlying structural organization. However, the resulting topological summaries are typically used as predictors or features for classification and group comparisons rather than treated as primary variables of interest. We construct a generalized multilevel functional framework for analyzing persistent-homology summaries as longitudinal functional responses in repeated three-dimensional structural magnetic resonance imaging (MRI). Specifically, we represent topological features using Betti curves and model these curves as count-valued functional responses. A negative-binomial distribution accommodates the discrete and potentially overdispersed nature of Betti counts, while the multilevel formulation accounts for the dependence induced by repeated measurements and separates between-subject and within-subject sources of functional variation. Functional principal component analysis further evaluates the dominant modes of variation at each level. A Bayesian approach is used for joint estimation of the functional regression and multilevel functional principal components. We apply this modeling framework to longitudinal structural MRI data from the OASIS-2 study to investigate associations between brain topology and demographic and clinical characteristics, including age, gender, follow-up time, and dementia severity. The results demonstrate that the framework can capture covariate-associated variation across the filtration continuum while evaluating distinct sources and patterns of longitudinal variation across homology dimensions. 2026-09-07T16:05:59Z Shashipraba N. K. Rajakaruna Asim K. Dey A. Alexandre Trindade http://arxiv.org/abs/2609.07631v1 Structural Analysis of a Dynamic Multilayer Network via Matrix Autoregressive Models: A Case Study of International Interactions between Countries 2026-09-07T15:28:14Z Dynamic and multilayer networks have been widely studied separately, but their joint analysis remains comparatively underdeveloped. Because the relational information of a dynamic multilayer network can be represented as a tensor at each time point $t$, each layer can be summarized through a set of structural statistics, yielding a matrix-valued observation and, consequently, a matrix-valued time series. To exploit this structure, we propose the use of matrix autoregressive (MAR) models, which simultaneously characterize temporal dependence across relational layers and structural statistics. We apply this framework to the ICEWS dataset, which records international interactions among countries under four relational domains and therefore naturally defines a dynamic multilayer network. The results indicate that negative verbal interactions (\textit{Verbal-}) play a prominent role in the subsequent structural reconfiguration of the material-interaction layers, while mean strength exhibits the strongest temporal persistence and reciprocity the broadest cross-statistic influence. These findings illustrate the usefulness of MAR models for providing a parsimonious and interpretable characterization of temporal and cross-layer dependence in dynamic multilayer networks. 2026-09-07T15:28:14Z 30 pages, 11 figures, 3 tables Camila Pinzón Mario Arrieta Juan Sosa http://arxiv.org/abs/2609.07610v1 Translation of Black-Box Clinical Prediction Models into Standalone Transparent Nomograms: Temporal External Validation in Heart Transplantation 2026-09-07T15:17:26Z We convert black-box clinical prediction models for tabular data into standalone nomograms that can be audited term by term. PRiSM (Partial Responses in Structured Models) takes the shape of each effect and interaction from the source model, not merely which variables mattered, and lets the outcome select and weight them. We tested this in 50,356 heart transplant recipients, with validation in a later era than training. Nomograms from all 5 source models - a public clinical risk score, logistic regression, neural networks, random forests and extreme gradient boosting - met a prespecified noninferiority criterion for discrimination before any further simplification, and generally preserved calibration and clinical net benefit. Those from the 3 machine-learning models showed no detectable difference in discrimination from de novo generalized additive and explainable boosting models, exceeded neural additive models, and carried fewer terms than the explainable boosting model. PRiSM is released as an open-source Python package. 2026-09-07T15:17:26Z 33 pages of main text with 4 figures and 3 tables; supplemental information (supplementary methods, Figures S1-S4, Tables S1-S26) appended, 86 pages total Henry Pigot Paulo J. G. Lisboa Sandra Ortega-Martorell Ivan Olier Joseph Mahon Johan Nilsson http://arxiv.org/abs/2609.07531v1 Simulation-Supervised Foundation Models for Retention Time Prediction in High-Performance Liquid Chromatography beyond Experimental Data Coverage 2026-09-07T14:09:39Z Accurate prediction of high-performance liquid chromatography (HPLC) retention times (RTs) across diverse molecules and chromatographic methods remains challenging because experimental training data cover only a limited region of chemical and method spaces. Here, we develop FUSE-RT (Foundation model Unifying Simulation and Experimental supervision for Retention Time), a multitask foundation model that integrates RT data from 179 chromatographic methods and adapts to unseen molecules and methods using limited target-domain data. To extend transferability beyond experimental coverage, we introduce simulation-to-real (Sim2Real) transfer learning, in which molecular representations learned from large-scale computational data are transferred to experimental RT prediction. Specifically, we use PolyOmics, comprising 39 properties for approximately 21,400 molecules generated by molecular dynamics and density-functional theory calculations, as auxiliary supervision. We evaluate generalization under molecular, method, and joint molecular--method distribution shifts. Simulation-derived supervision substantially improves transfer beyond the experimental molecular domain, particularly under pronounced coverage gaps and few-shot adaptation. Moreover, RT-prediction error decreases systematically with increasing simulation-data size, following a significant power-law relationship. These results establish Sim2Real transfer as a scalable strategy for extending RT prediction beyond the finite coverage of experimental chromatographic data. 2026-09-07T14:09:39Z Stephen Wu Yufeng Han Yasuhiro Mito Yoshiyuki Watabe Yoshihiro Hayashi Hikaru Takaya Takuya Kubo Ryo Yoshida http://arxiv.org/abs/2609.07512v1 Statistical versus machine learning-based spatial interpolation of post-processed ensemble weather forecasts 2026-09-07T14:00:29Z Statistical post-processing improves ensemble weather forecasts, but generating calibrated predictions at locations without observations remains challenging. This study compares statistical and machine-learning-based methods for post-processing ECMWF 2-m temperature and 10-m wind speed forecasts at observed and unobserved stations in Germany. We consider EMOS-based approaches, distributional regression networks, Transformers, and graph neural networks under both limited and extended predictor settings. For temperature, we also investigate linear forecast combinations and propose an altitude-aware linear pool (ALP). The results show that post-processing improves upon the raw ensemble in most settings, but no single method performs best across all variables, station groups, and evaluation metrics. The proposed ALP provides a small but significant improvement over the standard linear pool at unobserved locations. 2026-09-07T14:00:29Z 25 pages, 3 figures, 15 tables Mária Lakatos http://arxiv.org/abs/2608.21128v2 Structural Estimation of Marketing Mix Model Parameters from Geo-Experiments 2026-09-07T12:09:21Z Marketing Mix Models (MMMs) are widely used for marketing measurement and budget allocation, but face fundamental identification challenges: due to endogenous marketing spend decisions, MMM estimation on observational time-series data cannot recover the true causal effects of marketing. On the other hand, geo-experiments provide causal identification through randomization, but it is not clear how to use them efficiently to calibrate marketing mix models. We propose a novel structural estimation approach that recovers the complete set of MMM parameters - adstock decay ($α$), saturation ($λ$), and effectiveness ($β$) - directly from geo-experimental time-series. By differencing outcomes between treatment and control regions, our method eliminates observed and unobserved confounding factors while preserving the temporal variation that identifies each parameter. We demonstrate on synthetic data that this approach recovers the true ROAS and the response curve over the range of spending covered by the experiments, together with credible estimates of the underlying parameters. Our framework enables efficient pooling across multiple experiments and provides a principled foundation for MMM calibration that fully utilizes the information experiments contain. 2026-08-21T14:06:53Z Niklas Heusch http://arxiv.org/abs/2608.21130v2 A Synthetic Benchmark Dataset with Endogenous Marketing Spend for Validating Marketing Mix Models 2026-09-07T12:04:15Z Marketing Mix Models (MMMs) estimate the incremental sales effect of advertising from observational time series, yet they are rarely validated against ground truth, because ground truth is unobservable in real data. Synthetic data closes that gap in principle, but existing generators produce marketing spend exogenously - omitting the central difficulty of the estimation problem, since real budgets are planned around promotional calendars, seasons, and recent performance. This paper presents a parameterized generator, and a fixed reference instance, of a synthetic weekly retail dataset (156 weeks, three media channels) in which spend arises from four documented coordination mechanisms - quarterly budget feedback, anticipatory spending ahead of a promotional calendar, scheduled TV bursts, and algorithmic performance chasing - on a demand baseline with seasonal, quality, price, and unobserved sentiment components. Spend translates into incremental sales through two transformations, carryover and diminishing returns, instantiated here as geometric adstock and logistic saturation with known parameters; the true causal decomposition of every week's sales is recorded alongside the error-contaminated variables a practitioner would observe. Every mechanism is a parameter that can be varied or switched off, and a companion procedure simulates go-dark geo-experiments with exact treatment effects. The seeded generator and reference instance are publicly released with notebooks that reproduce every number in this paper. 2026-08-21T14:07:08Z Niklas Heusch http://arxiv.org/abs/2609.07381v1 Distributed Lag Neural Additive Models 2026-09-07T11:58:03Z We introduce Distributed Lag Neural Additive Models (DLNAMs), neural-additive analogues of Distributed Lag Non-linear Models (DLNMs) for learning nonlinear effects distributed over lags. DLNAMs replace a prespecified spline cross-basis with neural components that learn exposure--lag response surfaces, avoiding choices of basis family, dimension, and knot placement while preserving additive interpretability and familiar distributed-lag summaries. Exp-centered input layers, smooth activations, and learned subnetwork mixtures produce smooth, locally adaptive representations; pointwise uncertainty combines a conditional last-layer Laplace approximation with between-member ensemble variation. In simulations, DLNAMs generally outperformed DLNM comparators, including penalized and treed variants, in recovering known response functions, with lower bias, stronger boundary recovery, and better-calibrated cumulative intervals; gains were largest for more demanding functions. The architecture performed consistently across sample sizes, outcome families, lag horizons, and jointly fitted multi-exposure settings, retaining recovery performance as exposures were added; fit-specific changes were largely confined to optimization, and applications recovered established empirical patterns. 2026-09-07T11:58:03Z Calle Helmersson Shivang Pandey Leonardo Olivetti Elena Raffetti 10.5281/zenodo.22288964 http://arxiv.org/abs/2609.07363v1 Bayesian Emulation of Multi-fidelity Earth System Modelling Using Hierarchical Gaussian Processes 2026-09-07T11:31:06Z Multi-fidelity Earth system models provide simulations at different levels of complexity and computational cost, but exhaustive exploration of the parameter space at the highest fidelity is often prohibitively expensive. Multi-fidelity emulators can reduce this burden by combining abundant lower-fidelity simulations with limited high-fidelity evaluations. We compare four Gaussian-process-based multi-fidelity approaches: the Kennedy--O'Hagan autoregressive model (K&O), hierarchical kriging (HK), Bayesian hierarchical emulation for multi-level models (BayHEm), and multi-fidelity deep Gaussian processes (MF-DGP). We evaluate the methods using two contrasting applications: a three-fidelity tsunami simulator and a two-fidelity implementation of the Joint UK Land Environment Simulator (JULES). Performance is assessed using leave-one-out predictive accuracy, uncertainty representation, design requirements, and computational characteristics. In the tsunami application, BayHEm gives the lowest normalised root mean square error (NRMSE = 0.031) and highest SCORE (3.025), while MF-DGP performs worse than the single-fidelity baseline when only 10 high-fidelity simulations are available. In the JULES application, MF-DGP gives the lowest NRMSE (0.079) and highest SCORE (3.032), with all 30 held-out high-fidelity observations lying within their nominal 95\% predictive intervals. These contrasting results show that no single multi-fidelity emulator is uniformly superior. Instead, method choice should reflect the complexity of the inter-fidelity relationship, the amount of high-fidelity information available, and the importance placed on predictive accuracy and uncertainty quantification. 2026-09-07T11:31:06Z Xiaoyu Xiong Louise Kimpton Huiyi Yang Mian Xu James Salter Peter Challenor http://arxiv.org/abs/2506.00680v2 Understanding the European energy crisis through structural causal models 2026-09-07T11:17:02Z Natural gas supplies in Europe were disrupted and energy prices soared in the context of Russia's invasion of Ukraine. Electricity prices in France experienced the largest relative increase among European countries, even though the share of natural gas in the electricity mix is small compared to its neighbours. In this article, we demonstrate the importance of causal statistical methods and propose causal graphs to investigate the French and Spanish electricity markets and pinpoint key influencing factors on electricity prices and net exports. We demonstrate that a causal approach resolves paradoxical results of simple correlation studies and enables a quantitative analysis of indirect causal effects and what-if scenarios. We introduce a linear structural causal model as well as non-linear tree-based machine learning combined with Shapley Flow values. The models elucidate the interplay of gas prices and the unavailability of nuclear power plants during the energy crisis as the high unavailability made France dependent on imports. 2025-05-31T19:14:46Z Revised and Extended Version Nature Communications 17, 6539 (2026) Sarah Schreyer Anton Tausendfreund Florian Immig Ulrich Oberhofer Julius Trebbien Aaron Praktiknjo Benjamin Schäfer Dirk Witthaut 10.1038/s41467-026-75433-7 http://arxiv.org/abs/2608.17573v2 Feature Priming in Online Linear Regression: Sparse-Regret Lower Bounds and Tight Coordinatewise Rates 2026-09-07T10:34:02Z In high-dimensional online prediction, sparse comparators motivate regret bounds that depend on sparsity rather than ambient dimension. Feature priming seeks such adaptation by reweighting features using past data and refitting a minimum-norm predictor. At COLT 2023, Warmuth and Amid posed the open problem of whether the univariate, Pearson, or multivariate priming rules admit competitive online regret guarantees. Under the natural past-only Moore--Penrose protocol, we establish sparse-regret lower bounds that refute the corresponding sparse-logarithmic guarantee. The key obstruction is cheap nuisance interpolation, which permits exact interpolation of the history while assigning insufficient weight to the truly predictive coordinate. An exact target-mass identity and a two-sign argument convert this obstruction into clipped prediction loss. Hadamard constructions yield $Ω(\min\{T,\sqrt d\})$ clipped regret for each of the three unit-power rules against a zero-loss one-sparse comparator. For every fixed power $α\ge1$, one shared paired construction further yields linear regret simultaneously for all three powered rules and selectors among them in sufficiently high dimension. A rank upper bound is tight for powered univariate priming, even with Euclidean-unit inputs, and for unit-power Pearson priming with coordinatewise bounded inputs and target-preserving totalization. A separate algebraic construction gives $Ω(\min\{T,d^{1/4}\})$ regret for unit-power multivariate priming under Euclidean-unit inputs. The univariate lower bound persists under any nonnegative second-stage ridge schedule, while a paired ridge construction yields linear lower bounds for all three powered rules. Exploratory diagnostics on frozen language-model activations are consistent with the same qualitative mechanism. The exact multivariate frontier remains open. 2026-08-18T09:32:03Z 56 pages, 2 figures. Added the unit-power Pearson exact frontier and a Euclidean-unit multivariate lower bound, with full proofs Huibo Xu Shi Fu Qixin Zhang Dacheng Tao http://arxiv.org/abs/2609.07246v1 Structure-Adaptive E-Value Filter for Detecting Regional Signals in Brain Imaging 2026-09-07T09:06:49Z Structural MRI provides a noninvasive view of the neuroanatomical differences associated with cognitive impairment and dementia. Using data from the Alzheimer's Disease Neuroimaging Initiative, we investigate which anatomical regions exhibit widespread gray-matter differences and how these regional patterns vary across the clinical spectrum. Addressing this goal requires translating spatially dependent voxel-level evidence into regional conclusions while controlling multiplicity across anatomical regions. We propose the Structure-adaptive E-value FilTer (SEFT), which uses flexible working models to construct spatially adaptive voxel-level scores and aggregates them into regional partial-conjunction e-values. When combined with the e-value Benjamini--Hochberg (e-BH) procedure, these e-values provide finite-sample control of the set-wise false discovery rate under arbitrary interregional dependence. The ADNI analysis reveals a coherent neuroanatomical pattern: the exploratory analysis shows that differences between normal cognition and mild cognitive impairment are concentrated in medial-temporal regions, whereas the differences between mild cognitive impairment and dementia extend more broadly into temporal--limbic and posterior association regions. Both patterns largely overlap the normal-cognition--dementia benchmark, identifying a shared anatomical core across the clinical comparisons. 2026-09-07T09:06:49Z Jingcheng He Hangjin Jiang Wenguang Sun for the Alzheimer's Disease Neuroimaging Initiative