https://arxiv.org/api/BlVHMulGhIFuysoFKtkesf6pCLA 2026-09-10T18:13:25Z 12145 15 15 http://arxiv.org/abs/2609.04914v1 Derivation of the Sample-Size Scaling of TWO-NN Intrinsic-Dimension Estimates from Molecular Dynamics Trajectories 2026-09-04T09:18:18Z The intrinsic dimension of a dataset is the number of independent directions needed to describe the space occupied by its data. Estimators based on nearest neighbors infer this number from how the probability to find a neighbor point grows around each sampled point. Because the distances $r$ between neighbor points decrease as the sample grows, the estimated dimension can depend strongly on the number of available points. Here, we derive the large-sample behavior of the TWO-NN estimator for data drawn from a smooth $d$-dimensional space. The typical nearest-neighbor distance scales as $N^{-1/d}$, and smooth deviations from a locally uniform distribution produce successive corrections proportional to $N^{-2/d}$. We test this result using the trajectories coming from ten independent $100~μ$s simulations of alanine dipeptide. Configurations are represented by all pairwise distances among the ten heavy atoms. This representation has a known geometric dimension of $3n_{\mathrm{at}}-6=24$. Over the investigated range, the TWO-NN estimate shows no systematic dependence on the temporal spacing between configurations, but increases from approximately $7.5$ to $15.6$ as the sample size grows from $10^2$ to $2\times10^5$. Extrapolations that retain corrections through $r^2$, $r^4$, and $r^6$ give limiting dimensions of $25.23$, $22.89$, and $27.00$, respectively. All three estimates lie close to the known dimension and collectively bracket it, supporting the proposed scaling. Their spread provides a direct estimate of the systematic uncertainty associated with the truncation. The derived scaling therefore explains the strong sample-size dependence of TWO-NN and provides a practical route from finite sample estimates to the underlying geometric dimension. 2026-09-04T09:18:18Z Riccardo Capelli http://arxiv.org/abs/2604.25948v3 Causal Edge Rees Algebras for Spatiotemporal Graphs 2026-09-04T08:18:38Z Understanding the evolution of connectivity in spatiotemporal systems requires mathematical frameworks capable of encoding not only instantaneous interactions but also their cumulative causal structure. In this work, we introduce the \emph{Causal Edge Rees Algebra} (CERA), a new algebraic construction associated with causal spatiotemporal graphs. Given a temporal filtration induced by causal constraints, we associate a sequence of edge ideals whose Rees algebra encodes the full history of connectivity evolution in a single graded object. This construction establishes a bridge between dynamic graph topology and commutative algebra. In particular, we show that successive quotients of the filtration capture the emergence of new structural connections, allowing the identification of critical edges responsible for the fusion of previously disconnected components. This leads to the definition of temporal bridge modules and to a bridge detection theorem, which relates the dimension of these modules to the reduction in the number of connected components over time. Unlike existing algebraic approaches in topological data analysis, which are primarily based on geometric filtrations, the proposed framework is driven by intrinsic causal constraints. As a result, the CERA captures not only topological features but also their temporal organization and mechanisms of coalescence. The theory provides a new algebraic perspective on causal network dynamics, connecting edge ideals, Rees algebras, and temporal graphs. Beyond its theoretical significance, the framework opens new directions for the analysis of spatiotemporal systems, including epidemic networks, transport systems, and information propagation processes. 2026-04-17T22:27:58Z The authors have identified a fundamental methodological issue in the manuscript: the mapping from directed spatiotemporal causal edges to the proposed Rees-algebraic construction is not sufficiently justified. Consequently, the algebraic interpretation and the downstream claims require substantial reformulation. The authors therefore withdraw the paper Marcilio Ferreira dos Santos Cleiton de Lima Ricardo http://arxiv.org/abs/2601.17724v3 Quantum-Inspired Algorithms beyond Unitary Circuits: the Laplace Transform 2026-09-04T08:00:23Z Quantum-inspired algorithms can deliver substantial speedups over classical state-of-the-art methods by executing quantum algorithms with tensor networks on conventional hardware. Unlike circuit models restricted to unitary gates, tensor networks naturally accommodate non-unitary maps. This flexibility lets us design quantum-inspired methods that start from a quantum algorithmic structure, yet go beyond unitarity to achieve speedups. Here we introduce a tensor-network approach to compute the discrete Laplace transform, a non-unitary, aperiodic transform (in contrast to the Fourier transform). We encode a length-$N$ signal on two paired $n$-qubit registers and decompose the overall map into a non-unitary exponential Damping Transform followed by a Quantum Fourier Transform, both compressed in a single matrix-product operator. This decomposition admits strong MPO compression to low bond dimension resulting in significant acceleration. We demonstrate simulations up to $N=2^{30}$ input data points, with up to $2^{60}$ output data points, and quantify how bond dimension controls runtime and accuracy, including precise and efficient pole identification. 2026-01-25T07:19:56Z 9 pages Noufal Jaseem Sergi Ramos-Calderer Gauthameshwar S. Dingzu Wang José Ignacio Latorre Dario Poletti http://arxiv.org/abs/2605.13627v3 SINAPSE: A lightweight deep learning framework for accurate and explainable neutron-$γ$ discrimination 2026-09-04T07:30:56Z Traditionally, neutron-$γ$ discrimination in organic scintillators relies on techniques such as time-of-flight (ToF) selection and pulse-shape discrimination (PSD). However, particle identification through graphical cuts remains challenging in the low-charge regime due to poor signal-to-noise ratios (SNR). In this work, we propose SINAPSE, a lightweight deep learning framework for accurate and explainable neutron-$γ$ discrimination in the low-charge regime. The framework employs a dual-branch architecture that combines a 1-dimensional convolutional autoencoder for waveform denoising with a classifier for particle identification. Random augmentations are applied to high-SNR waveforms to simulate low-charge conditions, enabling robust extrapolation into regimes where conventional PSD labels are unreliable. We show that SINAPSE achieves superior denoising performance compared to conventional digital signal processing techniques, and outputs well-calibrated probabilities, consistent with traditional graphical cuts. Finally, we apply SHAP (SHapley Additive exPlanations) values to show that model decisions are driven by physically meaningful pulse-shape features, confirming consistency with established PSD principles. 2026-05-13T14:53:47Z 14 pages, 15 figures Thomas Carreau Adrien Matta Owen Syrett Benoît Mauss David Etasse Cyril Lenain Pierre Morfouace Julien Taieb David Regnier Patrick Copp Matthew Devlin Charlène Surault Jason Surbrook http://arxiv.org/abs/2510.24054v3 Algorithmic Randomness, Exchangeability, and the Principal Principle 2026-09-04T04:23:59Z We introduce a framework uniting algorithmic randomness with exchangeable credences to address foundational questions in philosophy of probability and philosophy of science. To demonstrate its power, we show how one might use the framework to derive the Principal Principle -- the norm that rational credence should match known objective chance -- without circularity. The derivation brings together de Finetti's exchangeability, Martin-Löf randomness, Lewis's and Skyrms's chance-credence norms, and statistical constraining laws (arXiv:2303.01411). Laws that constrain histories to algorithmically random sequences naturally pair with exchangeable credences encoding inductive symmetries. Using the de Finetti representation theorem, we show that this pairing directly entails the Principal Principle of this framework. We extend the proof to partial exchangeability and provide finite-history bounds that vanish in the infinite limit. The Principal Principle thus emerges as a mathematical consequence of the alignment between nomological constraints and inductive learning. This reveals how algorithmic randomness and exchangeability can illuminate foundational questions about chance, frequency, and rational belief. 2025-10-28T04:26:19Z Accepted version, The British Journal for the Philosophy of Science. 22 pages. (Compared to v2, footnote 9 is expanded to clarify a relation between computable measures and limiting relative frequencies.) Jeffrey A. Barrett Eddy Keming Chen http://arxiv.org/abs/2512.18769v6 Quantitative mobile gamma-ray spectrometry through Bayesian inference 2026-09-04T02:10:40Z Accurate quantitative mapping of gamma-ray emitters is critical for applications ranging from radiological emergency response and environmental monitoring to nuclear security and deep space exploration. Here, we show that such mapping can be achieved by combining mobile gamma-ray spectrometry with high-fidelity Monte Carlo simulations and full-spectrum Bayesian inference. Using 6 s of single-pass mobile spectrometry data benchmarked against independent in-situ and laboratory assays, we demonstrate decisive source mixture identification ($>\!\!5σ$), meter-scale source localization, and recovery of source activities with percent-level accuracy. The developed method marks a critical advance in quantitative gamma-ray sensing, enabling improved radiological situational awareness, enhanced terrestrial geophysical and geochemical mapping, as well as more robust constraints on radionuclide abundances on extraterrestrial bodies across the Solar System. 2025-12-21T15:17:52Z 26 pages, 6 figures, 1 ancillary file David Breitenmoser Alberto Stabilini Malgorzata Magdalena Kasprzak Sabine Mayer http://arxiv.org/abs/2605.21530v2 Pairwise Distance-Diffusion Analysis (PDDA): A Geometric Framework for Estimating Hurst Exponents in Multivariate Long-Memory Processes 2026-09-03T23:49:36Z We introduce Pairwise Distance-Diffusion Analysis (PDDA), a geometric framework that connects the Hurst exponent to the scaling of pairwise distances in long-memory stochastic processes. From a single distance-plot representation, PDDA yields three complementary routes to persistence: R/S-PDDA, based on the growth of geometric extrema; MSD-PDDA, based on the scaling of the second moment of lagged distances; and recurrence-volume scaling, in which the decay of close returns with increasing temporal separation reflects the expansion of the lagged displacement cloud. The framework extends naturally to multivariate isotropic and anisotropic processes, where the local spatial dimension specifies the available degrees of freedom while the Hurst exponents govern their temporal expansion. These results establish the distance plot as a common geometric representation of persistence, providing a unified distance-based foundation for Hurst analysis. 2026-05-19T20:22:37Z Revised version: double-column format adopted; PDDA assumptions clarified. The former fractal-range treatment was replaced by recurrence-probability scaling, with a more precise derivation based on volume expansion of multivariate ARFIMA displacement clouds. Supplemental material updated accordingly Diogo C. Soriano Frederique Vanheusden Slawomir J. Nasuto http://arxiv.org/abs/2609.04520v1 Crowding controls the scaling of bus frequency with demand 2026-09-03T22:18:56Z Cities must allocate limited resources to maintain mobility, with uncertainties about the resulting state of the system. Analyzing roughly 3,000 bus routes with more than 4 billion yearly riders across 19 metropolitan areas worldwide, we uncover a robust scaling law of the form $f \sim (d/t)^α$ with exponent $α\in [1/2,\,2/3]$, linking the service frequency $f$ to passenger demand $d$ and route duration $t$. We show that this scaling emerges from a simple optimization principle: cities implicitly minimize total passenger waiting time under a fixed operational budget when both schedule frequency and crowding are taken into account. This mechanism produces two universal regimes: a frequency-dominated regime with $α= 1/2$ when crowding is negligible, and a capacity-dominated regime with $α= 2/3$ when most routes are overloaded. Intermediate exponents arise when only part of the network operates near capacity. Furthermore, we find that the benefits of additional investment are highly uneven across systems. For instance, our model suggests that a $20\%$ budget increase yields nearly a 5-minute reduction in daily waiting time per passenger in Boston, compared to only about 1 minute in Paris. These findings place urban transit within a broader class of constrained capacity-allocation problems, while highlighting a distinct regime in which prescribed route demands shape the allocation of limited service resources. The resulting scaling laws show how simple optimization principles can generate systematic exponents in complex transport systems, beyond the dissipation-based frameworks usually considered in physical and biological flow networks. 2026-09-03T22:18:56Z Proc. Natl. Acad. Sci. U.S.A. 123 (29) e2535998123 (2026) Siddharth Patwardhan Şirag Erkol Filippo Radicchi Marc Barthelemy 10.1073/pnas.2535998123 http://arxiv.org/abs/2609.04449v1 Learning multistate kinetics with a variational multistate committor network 2026-09-03T20:13:13Z The long-time dynamics of complex molecular systems often involves rare transitions across networks of metastable states. Building on transition-path theory, which provides a rigorous framework for describing rare transitions between two metastable states, we introduce the variational multistate committor network (VMCN), a neural framework that learns the probabilities of reaching each metastable state directly from molecular simulation data. From this representation, VMCN identifies state-specific commitment, candidate transition regions and committor-consistent pathways between state pairs, and an effective kinetic network characterized by transition rates. The model is trained using finite time-lag trajectory data together with boundary conditions defined on conservative state cores. Applications to a triple-well potential, trialanine isomerization, and the $c$--ring rotation in the V$_{\rm o}$ domain of a vacuolar ATPase show that VMCN recovers metastable organization, provides committor-consistent descriptions of transition mechanisms, and estimates state-to-state kinetics. VMCN further provides diagnostics for incomplete state decompositions and enables adaptive exploration of candidate metastable states and their connecting regions. By integrating VMCN with generative committor-guided path sampling (Gen-COMPAS) for chignolin, we start from two end point structures, identify a misfolded state and a candidate folding intermediate, and we direct subsequent sampling toward the resulting multistate transition network. 2026-09-03T20:13:13Z Chenyu Tang Cheng Giuseppe Chen Benoît Roux Christophe Chipot http://arxiv.org/abs/2608.07123v2 Thermodynamic Human-Computer Interaction 2026-09-03T13:43:12Z Target acquisition is often modeled separately for desktop, mobile, and other interaction modalities. We present Thermodynamic HCI, a framework that splits interaction into thermal equilibrium and non-equilibrium regimes. The theory generalizes across interaction modalities by representing agent-target interaction using kinetic and potential energies. We derive the movement time of Fitts' law and the speed-accuracy tradeoff observed in Schmidt's law from the principles of thermal physics. Furthermore, we develop theorems that describe how target properties, such as the color of a button, affect user accuracy. The target acquisition model, derived from the theory, when evaluated on desktop and mobile website prefetching experiments, achieved an accuracy of 98% for both cursor and touchscreen based interaction. For every clicked link, it produced a fetch:click ratio of 1.37 for desktop and 1.75 for mobile. 2026-08-07T11:32:27Z Uzafir Ahmad Rafaq Muaz Hassan Ali Muzaffar http://arxiv.org/abs/2511.12637v2 Universalities in the Avalanche Dynamics of Novelties and Non-Novelties and some Notes on the Heaps law 2026-09-03T11:00:09Z Unprecedented events intertwine with the repetition of the past in natural phenomena and human activities. Key statistical patterns, such as Heaps' and Taylor's laws and Zipf's law, have been identified as characterizing the dynamical processes that govern the emergence of novelties and the abundance of repeated elements. Observing these statistical regularities has been pivotal in motivating the search for modeling schemes that can explain them and clarify key mechanisms underlying the appearance of new elements and their subsequent recurrence. In this study, we analyze sequences of novel and non-novel elements, referred to as avalanches, in real-world systems. We show that avalanche statistics provide a complementary characterization of innovation dynamics, extending beyond the three fundamental laws mentioned above. Although arising from collective dynamics, some systems behave as a single instance of a stochastic process. Others, such as natural language, exhibit features that we can only explain by a superposition of different dynamics. This distinction is not apparent when considering Heaps' law alone, while it clearly emerges in the avalanche statistics. By interpreting these empirical observations, we also advance the theoretical understanding of urn-based models that successfully reproduce the observed behaviors associated with Heaps', Zipf's, and Taylor's laws. We derive analytical expressions that accurately describe the probability distributions of avalanches and the Heaps law beyond its asymptotic regime. Building on these results, we derive a scaling relation that we show also holds in real-world systems, indicating a form of universality in the dynamics of novelty. 2025-11-16T15:07:48Z 15 pages, 11 figures Filippo Santoro Alberto Petri Francesca Tria http://arxiv.org/abs/2609.03603v1 Neural-Network Maxent: a general extension with learned nonlinearity, applied to time-series for Desert Locust distribution modelling 2026-09-03T09:48:28Z Species Distribution Modelling (SDM) is essential for understanding how environmental conditions shape biodiversity, particularly for destructive pests such as the Desert Locust (Schistocerca gregaria), whose breeding dynamics are tightly coupled to rapidly evolving environmental conditions. Maxent has become the dominant method for presence-only data, but its reliance on a linear combination of hand chosen feature transforms limits its ability to capture the nonlinear, temporal relationships common in ecological monitoring, where covariates such as precipitation, soil moisture, and vegetation indices evolve meaningfully over time. Standard implementations flatten time-series covariates into independent features, discarding sequential structure that carries critical signal. We introduce RNN Maxent, an extension of the Maxent framework that replaces the fixed feature dictionary with a neural network, specifically a Gated Recurrent Unit (GRU), trained end to end via backpropagation. The approach preserves Maxent's presence only statistical foundations, background normalization, and probability calibration, differing only in that the nonlinearity is learned from data rather than fixed in advance. We apply RNN Maxent to map suitable habitat for the Desert Locust using 50 day environmental time series derived from ERA5 Land, MODIS, and Sentinel 3, maintaining a 7 day gap between covariates and presence records to yield forecasting behavior. Compared against standard Maxent, RNN Maxent improves performance across metrics (ROC AUC 0.862 std 0.036 vs. 0.792; F1 0.671 std 0.056 vs. 0.590). 2026-09-03T09:48:28Z 22 pages, 8 figures, preprint Alessandro Grassi Edoardo Kimani Bellotto Wassim El Azami Sabrina Outmani Maximilien Houel http://arxiv.org/abs/2609.03599v1 Reduced latent leakage does not reliably predict lower likelihood bias in collider inference 2026-09-03T09:47:10Z Reusable collider representations can be evaluated through downstream discrimination and probes of retained information, but neither quantity directly tests the behaviour of score templates in a profiled likelihood. We test a specific prediction in a controlled two-channel routing protocol: if reduced physics-label readability in a nuisance branch indicates a more inference-robust representation, it should accompany a smaller profiled signal-strength bias under fixed unmodelled shifts. In a public Compact Muon Solenoid $H\rightarrow ZZ\rightarrow4\ell$ workflow, a downstream split of fixed EveNet embeddings preserves signal/background area under the receiver operating characteristic curve ($0.9894\pm0.0004$) while reducing nuisance-branch physics readability from $0.961\pm0.013$ to $0.593\pm0.030$. Probe-sensitivity and effective-rank controls exclude a failed readout and branch collapse. In a separate top quark jet-tagging workflow, the leakage reduction recurs with preserved task performance. Across two development event shards, however, its Spearman association with maximum absolute profiled bias is $0.036$, and three of six material leakage-improving transitions do not reduce that bias. A one-shot preregistered confirmation on an independently accessed shard produces material leakage reductions in all three paired seeds, while the maximum absolute bias increases in two. Thus, within the tested protocol, latent readability is a useful routing diagnostic but not a likelihood-robustness certificate. The result supports a practical validation rule: claims about inference robustness require a prespecified likelihood-facing stress test and held-out confirmation. 2026-09-03T09:47:10Z 27 pages, 6 figures Tong Pan http://arxiv.org/abs/2609.03581v1 The Greedy Bump Bias: Local Profiling Geometry and the Look-Elsewhere Effect 2026-09-03T09:27:58Z When fitting a localized signal whose position or shape is not known in advance, one typically allows these parameters to vary together with the signal amplitude and chooses the values that maximize the likelihood. This freedom has two related statistical consequences. If a genuine signal is present, its fitted amplitude will be affected by a positive bias; otherwise, the same freedom increases the chance of finding an unusually signal-like background fluctuation, giving rise to the look-elsewhere effect. We show that these two effects can be understood as consequences of the same local geometry of the family of signal templates. We study this connection in a Gaussian matched-filter model, where a smooth D-dimensional family of normalized templates describes the unknown signal location or shape. In the normalized matched-filter problem, the curvature of a genuine signal peak and the fluctuations that determine the curvature of a high background peak are governed by the same template metric. This allows us to derive an explicit asymptotic relation. We then follow the problem away from the strong-signal and high-threshold limits. Separating the signal-associated maximum from the best competing maximum gives an exact decomposition of the global bias into a local profiling contribution and a contribution from remote-peak competition. In a one-dimensional Gaussian example, the second factorial cumulant accounts for most of this correction, while the third brings the prediction into close agreement with simulation. A two-point Kac--Rice calculation reproduces the second cumulant and reveals a quartic short-distance suppression of nearby maxima. The resulting picture separates the roles of local dimension, model-dependent curvature, and global extremal competition within a common framework. 2026-09-03T09:27:58Z 42 pages, 5 figures, 2 appendices Tommaso Dorigo http://arxiv.org/abs/2607.21916v2 Estimating dynamic models by matching random features 2026-09-02T22:28:40Z Scientists increasingly express their ideas as dynamic models of complex processes. It is often much easier to simulate these models than to calculate the probability of their generating a particular outcome, making likelihood-based estimation infeasible. Existing likelihood-free approaches rely either on manually chosen summary statistics or on representations learned by neural networks. The former is error-prone and laborious, while the latter is computationally intensive, leaving many scientists in a difficult position. We show that, for a large class of dynamic models, parameters can be estimated by matching a small number of random features of the observed and simulated data. Specifically, we show that models with a $p$-dimensional parameter can be identified from just $2p+1$ generic random Fourier features. We introduce two estimators for stationary and nonstationary processes, respectively, and we establish their consistency under mild regularity conditions. More broadly, our results serve as the foundation for a new class of random feature methods for simulation-based estimation and inference. 2026-07-24T02:43:40Z Michael Wieck-Sosa Cosma Rohilla Shalizi