https://arxiv.org/api/bJeJODwVKhbyr2EwUUWSTgm9z90 2026-09-10T17:23:03Z 12145 0 15 http://arxiv.org/abs/2609.10341v1 Surrogate-accelerated parameterisation of physics-based Li-ion battery models 2026-09-09T15:37:09Z Physics-based lithium-ion battery models provide access to physically meaningful internal electrochemical states and processes, but cell-specific parameter inference from terminal current-voltage data is computationally expensive and limited by identifiability. We present a surrogate-accelerated inverse framework based on a single-particle model with electrolyte dynamics (SPMe). Its forward map uses our Artiphy surrogate framework for rapid, differentiable evaluation of voltage and selected internal states. After rescaling to remove exact structural redundancies, we infer non-redundant transport, kinetic and capacity parameter groups, including concentration-dependent solid and electrolyte diffusivities. Synthetic voltage data from a Doyle-Fuller-Newman (DFN) model under a WLTP-like current protocol provide a benchmark with known reference parameters and controlled model discrepancy. The inferred SPMe reproduces the benchmark voltage with an error of order 1 mV and recovers electrode capacities well. Positive-electrode diffusivity is recovered accurately over much of the probed stoichiometric range. Local sensitivity and Fisher-information analysis identifies correlated kinetic-Ohmic and electrolyte-transport directions, and shows how localised information and the global diffusivity parameterisation can yield narrow Fisher-curvature envelopes despite weak voltage sensitivity to negative-electrode diffusion over much of the drive cycle. These results represent a step towards rapid physics-based in-silico parameterisation and reduced reliance on destructive cell characterisation. 2026-09-09T15:37:09Z 37 pages, 11 figures A. Emir Gumrukcuoglu Josh Pearson Jamie M. Foster James Burridge http://arxiv.org/abs/2609.10216v1 Information capacity of quantum statistics: Fock-state tests of a discrete binary-sequence model on cloud photonic quantum processors 2026-09-09T14:15:51Z Our central premise is that quantum mechanics may be the statistical limit of a more fundamental discrete theory: any such theory equips a physical system with a finite information capacity, and its departure from quantum statistics is controlled by how much of that capacity the system uses. We show that commercial cloud photonic quantum processors have reached the precision required to bound this capacity from below, using the binary-sequence model of Powers et al. as the concrete test theory: outcome probabilities arise from counting discrete sequences of length $n$, quantum mechanics is recovered as $n \to \infty$, and $n$ measures the information capacity of the register behind a prepared state. Photon Fock states $|1\rangle$, $|1,1\rangle$, heralded $|2\rangle$, and cascaded beam-splitter pairs are measured on programmable interferometers with dominant systematics determined in situ. The model's composition-consistent parametrization, singled out by requiring that rotations compose, recovers quantum mechanics with deviations $1.24/n$; a random-effects likelihood analysis calibrated by parametric bootstrap excludes all $n \le 100$: the information capacity of the register carrying the two-photon state, if finite, exceeds $10^2$. Cascaded beam splitters test the composition law directly: the data are split-invariant, excluding naive count composition at $8σ$ and confirming the interference-sign rule. Model-independently, curve-averaged deviations from the quantum partition law larger than $2.3\times10^{-2}$ are excluded at 95% CL, and the originally published linear parametrization is excluded outright. Because the compilation offset is frozen per circuit it is calibratable, opening the $10^{-3}$ floor ($n \sim 10^3$) to current hardware: cloud photonic processors are quantitative instruments for quantum foundations, and information capacity an experimentally boundable quantity. 2026-09-09T14:15:51Z 10 pages, 10 figures Chiran Wijesundara Octavia T. Volpe Dejan Stojkovic Herbert Fotso Tim Thomay http://arxiv.org/abs/2405.00636v4 Robustness of shallow graph embedding methods for community detection 2026-09-09T05:37:55Z This study investigates the robustness of shallow graph embedding methods for community detection in the face of network perturbations, specifically node deletions. Graph embedding techniques, which represent nodes as low-dimensional vectors, are widely used for various graph machine learning tasks due to their ability to capture structural properties of networks effectively. However, the impact of perturbations on the performance of these methods remains relatively understudied. The research considers state-of-the-art shallow graph embedding methods from two families: matrix factorization (e.g., LE, LLE, HOPE, M-NMF) and random walk-based (e.g., DeepWalk, LINE, node2vec). Through experiments conducted on both synthetic and real-world networks, the study reveals varying degrees of robustness within each family of shallow graph embedding methods. The robustness is found to be influenced by factors such as network size, initial community partition strength, and the type of perturbation. Notably, node2vec and LLE consistently demonstrate higher robustness for community detection across different scenarios, including networks with degree and community size heterogeneity. These findings highlight the importance of selecting an appropriate shallow graph embedding method based on the specific characteristics of the network and the task at hand, particularly in scenarios where robustness to perturbations is crucial. 2024-05-01T17:04:20Z Accepted manuscript. Published in Applied Network Science Appl Netw Sci (2026) Zhi-Feng Wei Pablo Moriano Ramakrishnan Kannan 10.1007/s41109-026-00825-z http://arxiv.org/abs/2501.10886v3 Linear scaling causal discovery from high-dimensional time series by dynamical community detection 2026-09-08T21:35:48Z Understanding which parts of a dynamical system cause each other is extremely relevant in fundamental and applied sciences. However, inferring causal links from observational data, namely without direct manipulations of the system, is still computationally challenging, especially if the data are high-dimensional. In this study we introduce a framework for constructing causal graphs from high-dimensional time series, whose computational cost scales linearly with the number of variables. The approach is based on the automatic identification of dynamical communities, groups of variables which mutually influence each other and can therefore be described as a single node in a causal graph. These communities are efficiently identified by optimizing the Information Imbalance, a statistical quantity that assigns a weight to each putative causal variable based on its information content relative to a target variable. The communities are then ordered starting from the fully autonomous ones, whose evolution is independent from all the others, to those that are progressively dependent on other communities, building in this manner a community causal graph. We demonstrate the computational efficiency and the accuracy of our approach on time-discrete and time-continuous dynamical systems including up to 80 variables. 2025-01-18T21:50:43Z Matteo Allione Vittorio Del Tatto Alessandro Laio 10.1103/kd73-93cg http://arxiv.org/abs/2201.01096v2 Wavescan: Multiresolution Time-Frequency Transform of Gravitational-Wave Data 2026-09-08T21:21:42Z Identifying transient signals embedded in non-stationary noise requires analyzing the time-dependent spectral components of the observed time series. The time-frequency distribution of the signal power can be estimated with Gabor atoms, or wavelets, localized in time and frequency by a window function. Such analysis is constrained by the Heisenberg-Gabor uncertainty principle, which limits the simultaneous time and frequency localization achievable with individual wavelets. Moreover, the resulting time-frequency distribution is subject to temporal and spectral leakage, limiting the identification of sharp or rapidly varying features in the power spectrum. This paper introduces a time-frequency transform that uses a stack of wavelets to scan local power across multiple scales. At each time-frequency location, a wavelet least affected by the leakage is selected from the stack to obtain high-resolution localization of power. The resulting wavelet scan ("wavescan") extends conventional multiresolution analysis by enhancing time-frequency localization and mitigating local power distortions caused by temporal and spectral leakage. The paper describes the principal components of the wavescan framework, including the estimation of the time-varying spectrum, identification of transient signals in the time-frequency data, and reconstruction of the corresponding time-domain waveforms. To demonstrate the performance of the method, the wavescan transform is applied to gravitational-wave data from the LIGO detectors. 2022-01-04T11:47:25Z 8 pages, 6 figures Sergey Klimenko http://arxiv.org/abs/2609.08767v1 Quantity, quality, and timing: Guiding glacier data assimilation strategies in the high Arctic 2026-09-08T14:02:24Z Accurate simulation of glacier surface mass balance is essential for predicting sea level rise and freshwater resources, but it is constrained by uncertainties in meteorological forcing and model parameters. Here, we deploy glacier data assimilation strategies to assess the value of observations for improving surface mass balance simulation, focusing on observation quantity, quality, and timing. We perform synthetic twin experiments on Kongsvegen glacier, Svalbard, using a Particle Batch Smoother with 1000 ensemble members. Synthetic observations of albedo, snow depth, and surface temperature are assimilated at two quality levels, under two climatic scenarios, and over 12 years. Assimilation benefit is measured as the percentage improvement in the continuous ranked probability score of the posterior glacier surface mass balance relative to the prior. A single optimally timed high quality observation yields mean improvements of up to 80\%. Larger numbers of low quality observations partially compensate for lower improvement. In the accumulation zone, however, additional snow depth observations degrade performance through particle degeneracy. Optimal timing is governed by the seasonal transitions of the truth trajectory rather than by prior ensemble spread alone. The optimal windows shift by up to six weeks between early and late melting years. Joint assimilation adds value through temporal diversity rather than observational diversity, while independently timed observations outperform same day combinations. The asynchronously optimally timed combined assimilation of three variables sustains improvements of 85 to 97\% across all years in the ablation zone. These findings provide guidelines for adaptive observation scheduling in glacier monitoring and reanalysis. 2026-09-08T14:02:24Z 26 pages, 7 figures. Prepared for submission to Frontiers Wenxue Cao Kristoffer Aalstad Louise S. Schmidt Thomas V. Schuler http://arxiv.org/abs/2609.08477v1 Nonequilibrium stochastic thermodynamics of boundary functionals: From Zubarev's ensemble to stochastic particle separation 2026-09-08T09:20:37Z This paper develops a generalization of Zubarev's nonequilibrium statistical operator method for the case of simultaneous inclusion of additive and nonlocal boundary functionals of trajectories. Using a unified thermodynamic approach, a three-parameter model of nonequilibrium systems is constructed, including the first-passage time, the dwell time above a given level, and the absolute extremum of the process. A correspondence is demonstrated between the maximum information entropy method for trajectories and the Donsker-Varadan large deviation formalism. Using the Doob`s h-transform, it is demonstrated that fixing extremal functionals of history induces efficient non-Markovian transport in the system with dynamic adaptation to historical records. Criteria for the applicability of the developed apparatus in the "large time" domain and the possibilities of its use for optimizing stochastic particle separation in periodic potentials are discussed. 2026-09-08T09:20:37Z 33 pages, 5 figures V. V. Ryazanov http://arxiv.org/abs/2609.08460v1 Finding the distribution of matter using lenses - II: deconvolution-based reconstruction with 3x2pt measurements 2026-09-08T08:57:44Z We present a deconvolution-based framework for testing scale-dependent departures of the late-time matter power spectrum from a fiducial cosmological model using $3\times2$pt measurements. We introduce a free-form scale-dependent modulation $A(k)$ of the fiducial nonlinear matter power spectrum and construct the linear response of binned galaxy-clustering, galaxy-galaxy-lensing, and cosmic-shear spectra to the discretized modulation $A(k)$. The response is evaluated with full-sky, beyond-Limber kernels including density, redshift-space-distortion, gravitational-shear, and intrinsic-alignment contributions. We reconstruct $A(k)$ using a regularized modified Richardson-Lucy algorithm, with weak diffusion in $\ln k$ and selection of the minimum-$χ^2$ solution along the iteration history. Using Rubin/LSST Year 10-like synthetic data, we find that oscillatory modulations with amplitudes $\gtrsim1\%$ can be recovered over $0.1\lesssim k\lesssim0.5\,{\rm Mpc}^{-1}$, provided the oscillation frequency $f\lesssim10$ on $\log_{10}[k/(0.2\,{\rm Mpc}^{-1})]$. We further introduce a posterior-weighted consistency statistic calibrated with posterior-predictive null mocks, thereby accounting for cosmological and nuisance-parameter uncertainties without relying on Wilks' theorem. The null case is consistent with $A(k)=1$, while a $1\%$ oscillatory modulation is detected at $\sim 2.6σ$. These results demonstrate the potential of regularized deconvolution as a model-independent consistency test of the matter power spectrum in future $3\times2$pt surveys. 2026-09-08T08:57:44Z 13+7 pages, 9 figures Jun-Qian Jiang Ajoy Dawn Dhiraj Kumar Hazra Benjamin L'Huillier Arman Shafieloo http://arxiv.org/abs/2609.08457v1 Finding the distribution of matter using lenses - I: deconvolution-based reconstruction with CMB lensing 2026-09-08T08:56:32Z The matter power spectrum is one of the primary statistical descriptors of the large-scale distribution of matter in the Universe and provides a powerful probe of cosmic structure formation. Measurements of cosmic microwave background (CMB) lensing offer an integrated view of the matter distribution over a wide range of redshifts, enabling the reconstruction of the underlying matter power spectrum. In this work, we reconstruct the reference linear matter power spectrum $ P_\text{lin}(k,0)$ from the baseline joint CMB lensing measurements of Planck PR4, ACT DR6, and SPT-3G using a covariance-weighted modified Richardson-Lucy(MRL) deconvolution algorithm. The reconstructed spectrum is found to be consistent with the fiducial linear prediction on large scales, while exhibiting a systematic enhancement for $k \gtrsim 0.1\,{\rm Mpc}^{-1}$, where nonlinear gravitational evolution becomes important. To investigate this behavior, we introduce a scale-dependent correction factor, $A(k)$, defined through $P(k)=A(k)\,P_{\rm nl}(k),$ where $P_{\rm nl}(k)$ is the fiducial nonlinear matter power spectrum obtained from 2LPT simulations. The reconstructed correction factor remains consistent with unity within $2σ$ confidence over the reconstructed range, indicating that the observed enhancement is well explained by the standard nonlinear evolution of the matter power spectrum. In addition, the reconstruction shows agreement with the fiducial BAO template around the BAO feature at $k\sim(0.04-0.06)\ {\rm Mpc}^{-1}$, indicating that some BAO-scale information survives the lensing projection. 2026-09-08T08:56:32Z v1: 18 pages and 5 figures Ajoy Dawn Jun-Qian Jiang Dhiraj Kumar Hazra Benjamin L'Huillier Arman Shafieloo http://arxiv.org/abs/2510.23717v2 Robust and Generalizable Background Subtraction on Images of Calorimeter Jets using Unsupervised Generative Learning 2026-09-08T01:48:21Z Accurate separation of signal from background is one of the main challenges for precision measurements across high-energy and nuclear physics. Conventional supervised learning methods are insufficient here because the required paired signal and background examples are impossible to acquire in real experiments. Here, we introduce an unsupervised unpaired image-to-image translation neural network that learns to separate the signal and background from the input experimental data using cycle-consistency principles. We demonstrate the efficacy of this approach using images composed of simulated calorimeter data from the sPHENIX experiment, where physics signals (jets) are immersed in the extremely dense and fluctuating heavy-ion collision environment. Our method outperforms conventional subtraction algorithms in fidelity and overcomes the limitations of supervised methods. Furthermore, we evaluated the model's robustness in an out-of-distribution test scenario designed to emulate modified jets as in real experimental data. The model, trained on a simpler dataset, maintained its high fidelity on a more realistic, highly modified jet signal. This work represents the first use of unsupervised unpaired generative models for full detector jet background subtraction and offers a path for novel applications in real experimental data, enabling high-precision analyses across a wide range of imaging-based experiments. 2025-10-27T18:00:07Z The code is publicly available at https://github.com/LS4GAN/uvcgan-s, and the data used for training are available on Zenodo at https://zenodo.org/records/17783990 Go, Y., Torbunov, D., Huang, Y. et al. Robust and generalizable background subtraction on images of calorimeter jets using unsupervised generative learning. Scientific Reports (2026) Yeonju Go Dmitrii Torbunov Yi Huang Shuhang Li Timothy Rinn Haiwang Yu Brett Viren Meifeng Lin Yihui Ren Dennis Perepelitsa Jin Huang 10.1038/s41598-026-59672-8 http://arxiv.org/abs/2609.07722v1 Real-time blind indexing by cross-frame consensus 2026-09-07T16:34:34Z Classical crystallography determines orientation and unit-cell geometry from a rotation series collected from a single specimen, whereas serial crystallography acquires one still exposure from each of many randomly oriented microcrystals and must reconstruct a complete dataset by pooling partial measurements across the ensemble. In serial experiments, each frame typically contains only a subset of reflections, and frames with few reliable peaks (the small-N regime) are especially difficult to index blindly, because multiple incorrect lattices or orientations can explain sparse observations. The approach implemented in GLINT addresses this small-N regime by treating the dataset, rather than the individual frame, as the unit of inference. Candidate solutions are generated for many frames, weak but recurrent lattice hypotheses are pooled across the ensemble to identify a shared unit cell, and each frame is then registered against that consensus cell through a known-cell orientation search with lower dimensionality than blind indexing. In benchmark tests, GLINT matched the strongest blind indexer in the comparison on indexing yield at roughly 40x lower per-frame blind-indexing latency against XGANDALF's fastest configuration, and batched known-cell registration ran entirely on one GPU, allowing cell discovery on the fly during serial data collection. On experimental serial data, frames indexed blind by GLINT merged to crystallographic quality (CC* = 0.90), while the method rejected non-crystal images rather than forcing unsupported indexing assignments. These results demonstrate that indexing performance can be evaluated not only by indexing yield and throughput, but also by the quality of the resulting merged crystallographic data. 2026-09-07T16:34:34Z 32 pages including Supporting Information (Secs. S1-S16), 8 figures, 12 tables Stefano Marchesini Yuan Ni http://arxiv.org/abs/2507.08867v3 Mind the Gap: Navigating Inference with Optimal Transport Maps 2026-09-06T15:24:22Z Machine learning (ML) techniques have recently enabled enormous gains in sensitivity to new phenomena across the sciences. In particle physics, much of this progress has relied on excellent simulations of a wide range of physical processes. However, due to the sophistication of modern machine learning algorithms and their reliance on high-quality training samples, discrepancies between simulation and experimental data can significantly limit their effectiveness. In this work, we present a solution to this ``misspecification'' problem: a model calibration approach based on optimal transport, which we apply to high-dimensional simulations for the first time. We demonstrate the performance of our approach through jet tagging, using a dataset inspired by the CMS experiment at the Large Hadron Collider. A 128-dimensional internal jet representation from a powerful general-purpose classifier is studied; after calibrating this internal ``latent'' representation, we find that a wide variety of quantities derived from it for downstream tasks are also properly calibrated: using this calibrated high-dimensional representation, powerful new applications of jet flavor information can be utilized in LHC analyses. This is a key step toward allowing the unbiased use of ``foundation models'' in particle physics. More broadly, this calibration framework has broad applications for correcting high-dimensional simulations across the sciences. 2025-07-09T16:28:21Z 31 pages, 13 figures Malte Algren Tobias Golling Francesco Armando Di Bello Christopher Pollard http://arxiv.org/abs/2609.06085v1 Explainable Deep Learning for Price-Trade Dynamics: From Black-Box Forecasts to Effective Parametric Models 2026-09-05T13:25:40Z Understanding the joint dynamics of prices and trades is central to market microstructure, where returns and order flow interact through nonlinear and state-dependent mechanisms. Linear models are interpretable but may miss these effects, while deep neural networks improve forecasting at the cost of transparency. We use neural networks as tools for structural discovery rather than only for prediction. A deep feed-forward network is trained on high-frequency returns and signed volumes for large- and small-tick stocks and compared with a linear VAR benchmark. The neural network improves predictive performance, especially for returns, revealing nonlinear dependencies beyond the linear specification. Using Shapley-based explainability, we show that the dominant contributions are concentrated at the most recent lags. Model-implied responses are consistent with conditional averages reconstructed from the data. Unlike empirical averages, however, the neural-network decomposition isolates individual regressor contributions to the aggregate dependence. Lagged signed volume generates sign-preserving and saturating effects, consistent with nonlinear price impact and order-flow persistence. Lagged returns act as state variables: when the previous trade does not move the price, the model predicts continuation in the direction of past order flow, whereas non-zero returns generate attenuation or reversal. Building on these findings, we introduce a parsimonious SHAP-inspired nonlinear parametric model. It reproduces the main return-volume dependencies, outperforms the linear VAR benchmark, and achieves performance comparable to the neural network. A multi-lag extension captures residual longer-memory effects while preserving interpretability. Overall, explainability offers a route from black-box prediction to economically meaningful parametric models of price and trade dynamics. 2026-09-05T13:25:40Z Manuel Naviglio Fabrizio Lillo http://arxiv.org/abs/2609.05207v1 FluxDisco: Symbolic Regression for Stoichiometric Dynamical Systems via Monte Carlo Graph Search 2026-09-04T14:38:59Z Dynamical symbolic regression methods identify governing differential equations from noisy data, balancing interpretability and predictive accuracy. However, standard methods often produce expressions that violate known physical laws. To address this, we propose FluxDisco, a physics-informed framework tailored for flux-based, stoichiometric ODE systems. By leveraging a known stoichiometry, we reduce the expression search space and ensure physical adherence. Our framework adapts the Monte Carlo Graph Search algorithm for the unique challenges associated with joint flux discovery of stoichiometric systems. We evaluate our method across a range of physical and biological systems, demonstrating its ability to accurately recover governing dynamics through interpretable equations. 2026-09-04T14:38:59Z Cassandra Durr Lancaster University Alvaro Köhn-Luque University of Oslo Chris Jewell Lancaster University Lloyd A. C. Chapman Lancaster University http://arxiv.org/abs/2607.16035v3 Dynamic models with $p$ parameters are identified by $2p+1$ random features 2026-09-04T10:19:03Z A foundational principle in nonlinear dynamics is that the structure of a dynamical system can be recovered from a small number of generic measurements or coordinates. We develop an analogous principle for the identification of dynamic models for time series with noise, which builds on previous identification results for noiseless dynamical systems. The noise is allowed to be non-iid, non-Gaussian, and dependent on the state. Our results cover noisily observed differential equations and discrete-time dynamical systems, as well as stochastic models with process noise. We illustrate the utility of this identification principle using a Lorenz-63 model and a Hénon map model, both with observational noise. 2026-07-17T15:08:20Z Michael Wieck-Sosa Cosma Rohilla Shalizi