https://arxiv.org/api//kf+O7v3seK1e6Iv486fLEkvQbY2026-09-10T16:34:30Z24217015http://arxiv.org/abs/2609.10477v1Multivariate linear regression without prior assumptions2026-09-09T17:18:04ZRecovering the linear relationships that govern a system from noisy measurements is a basic task across the physical and engineering sciences. Because every measured variable may carry an unknown amount of noise, classical regression must commit in advance to a set of structural assumptions: ordinary least squares requires a declared input-output partition with input variables being noise-free, total least squares assumes equal noise variance across all variables, and generalized total least squares additionally requires the noisy-variable partition and variances to be known beforehand. Kalman~\cite{Kalman:1982} showed that any procedure returning a unique linear model from inexact data must rest on such unverifiable a priori assumptions -- ``prejudices'' -- that cannot be checked against the data itself, and that removing them leaves the identification problem fundamentally indeterminate. Whether these prejudices can instead be resolved directly from the data has remain unresolved. Here we show that an iterative generalized-eigenvalue algorithm, QZ-IPCA, recovers the noisy-variable partition, noise variances, number of linear relations, and regression coefficients of a multivariate linear system simultaneously, using only the raw data. Across all possible exhaustive noise configurations of a five-variable benchmark network, QZ-IPCA correctly identifies model structure and recovers coefficients with error below 6.4\%. It outperforms ordinary least squares even when given the best partition, and succeeds in rank identification precisely where standard total least squares falls once noise variances differ across variables. These results show that the assumptions conventionally required for multivariate regression are not necessary, recasting model identification as a problem solvable from data geometry alone.2026-09-09T17:18:04ZThe manuscript is under further processing. Important modifications might take place in further versions. There are 32 pages containing 10 figures and 3 tablesMayank S. K. GuptaDeepanjhan DasArun K. TangiralaShankar Narasimhanhttp://arxiv.org/abs/2501.10675v3Recovering Unobserved Network Links from Aggregated Relational Data: Bayesian Latent Surface Modeling and Penalized Regression2026-09-09T16:35:52ZAggregated relational data (ARD) record counts of ties to attribute-defined groups while leaving individual edges unobserved. We compare latent-geometry and regularized network estimators through a common observation map. The comparison distinguishes the realized adjacency matrix, conditional edge probabilities, and model parameters. We study roster-based ARD with known node-level group memberships, giving both estimators the same roster and aggregate counts.
We relate the aggregate means to a Poisson working likelihood and a Huber loss, and give their derivatives. Overlapping groups, shared edges, and reporting error affect the interpretation of these objectives. Geometry restricts the representation of edge probabilities, while regularization selects among candidate fits. Identification depends on the observation map and model restrictions rather than uniqueness of a numerical optimizer.
A reproducible synthetic experiment specifies the data-generating process, estimation algorithms, and evaluation targets under matched information. The matrix estimator gives better realized-edge rankings and aggregate fit, while the geometric estimator gives lower error for generating probabilities. The resulting framework organizes ARD reconstruction around the interaction of observation design, structural assumptions, and computation.2025-01-18T06:51:51Z14 pages, 2 figures. Substantially revised replacement of the withdrawn version. Clarified observation regime and targets; corrected likelihood and loss calculations; added a reproducible matched-input synthetic experiment. Code and saved results are included as ancillary filesYen-hsuan Tsenghttp://arxiv.org/abs/2609.10275v1The Impact of a Gridded Streamflow Measure on Drought Variation in the Conterminous United States2026-09-09T14:56:07ZModels for droughts draw on a wide range of meteorological and hydrological inputs. Stakeholders classify droughts according to different purposes and priorities, and accordingly rely on different measures to explain and predict the onset of drought. While many meteorological inputs are available as gridded data products, hydrological streamflow measurements are often only available as point-referenced gauge measurements. This leads to an issue of misalignment for studies relying on both areal meteorological and point-referenced hydrological data. Such gauge data can also have notable spatial and/or temporal missingness. Many areas remain ungauged, and where gauges exist, equipment malfunctions cause temporal gaps. In this study, we document the value of an existing gridded streamflow measure for the conterminous United States, one which was specifically designed to match the spatio-temporal support of publicly available meteorological and ordinal drought measurements. We use this homogenized database to assess the relative importance of this streamflow measure in explaining US drought variability. This assessment first requires that we address autocorrelation and variability in the variance of observed droughts, two extensions to existing statistical methodology. After suitably controlling for both, our results show that the streamflow measure is often the most statistically important explanatory variable from among a wide set of meteorological variables. We explore the spatial variation and some drivers of this result.2026-09-09T14:56:07ZRob ErhardtCourtney Di VittorioStaci HeplerMostafa ShamsWendy WeiJames Zhaohttp://arxiv.org/abs/2609.10174v1A spatiotemporal negative binomial model with dynamic dispersion: An application to Tuberculosis infections2026-09-09T13:43:05ZTuberculosis (TB) remains a critical public health concern in Brazil, characterized by pronounced spatial heterogeneity and fluctuating temporal volatility. In this paper, we study monthly TB notifications across 61 microregions of Sao Paulo state from 2001 to 2024. To do this, we introduce a negative binomial spatial integer-valued generalized autoregressive conditional heteroskedastic (INGARCH) model featuring jointly dynamic conditional means and time-varying dispersion. To capture inter-regional spillovers, we incorporate both discrete adjacency structures and a novel continuous distance-based formulation leveraging the Matern correlation function. Parameter estimation via conditional maximum likelihood employs a two-step profile-likelihood iterative scheme, demonstrating solid finite-sample performance in simulation studies. Applied to the Sao Paulo TB surveillance data, the framework substantially outperforms standard Poisson and fixed-dispersion spatiotemporal baselines in empirical fit and uncertainty quantification, maintaining nominal 95% predictive coverage across both dense metropolitan centers and rural microregions. Our results reveal marked spatial heterogeneity in baseline incidence, dynamic overdispersion driven by localized outbreaks, and short-range spatial interaction decay. By accurately modeling spatiotemporal volatility, the proposed methodology provides a robust statistical tool to support public health surveillance, policy-making, and resource allocation.2026-09-09T13:43:05Z21 pages, 9 figures, 2 tablesRodrigo B. SilvaLuiza S. C. PiancastelliWagner Barreto-Souzahttp://arxiv.org/abs/2601.09673v3A probabilistic match classification model for low-scoring sports2026-09-09T11:18:58ZAll existing match classification models in the tournament design literature suffer from two major limitations: a contestant is considered indifferent only if uncertain future results do never affect its prize, and competitive matches are not distinguished with respect to the incentives of the contestants. We propose a probabilistic framework to address both issues. For each match, our approach relies on simulating all other matches played simultaneously or later to compute the qualifying probabilities for the three main outcomes (win, draw, loss), thereby classifying each match into six categories. The suggested model is applied to the last round of the previous group stage and the new incomplete round-robin league, introduced in the 2024/25 season of UEFA club competitions. The incomplete round-robin tournament is found to contain fewer unimportant matches with two indifferent teams, and substantially more matches where both teams should play offensively. However, the robustly higher proportion of potentially collusive matches can threaten with serious scandals.2026-01-14T18:15:32Z25 pages, 4 tables, 8 figuresLászló CsatóAndrás Gyimesihttp://arxiv.org/abs/2609.10020v1Cointegration by Parts: Locating Cointegration in Time2026-09-09T10:52:40ZTests for cointegration are typically applied to a single window spanning the entire sample, assuming that the long-run relationship holds throughout. When it holds over only a part of the sample, such tests lose power, because the stationary episode is diluted by periods without cointegration. We propose three statistics for testing whether two or more series cointegrate only over a part of the sample, each an infimum of the Engle-Granger statistic over recursive, backward-expanding, or doubly-flexible windows. We derive their limiting distributions and establish which alternatives each is consistent against. Only the doubly-flexible statistic has power against both break directions. Inspired by the seminal work of James G. MacKinnon, critical values are obtained by simulation and summarized through response surface regressions. We apply the tests to global mean sea level and global mean surface temperature anomalies. All three reject the null of no cointegration at the 5% level, locating it in a sub-period of the 1880--2019 record that coincides with documented discontinuities in sea surface temperature data collection.2026-09-09T10:52:40ZOlivia KvistJ. Eduardo Vera-Valdéshttp://arxiv.org/abs/1907.06994v2Regularized Estimation and Feature Selection in Mixtures of Generalized Linear Experts2026-09-09T08:53:00ZMixtures of experts (MoE) are conditional mixture models in which both the mixing proportions and the component densities depend on the predictors, and are widely used for regression, classification and model-based clustering of heterogeneous data. Fitting MoE by maximum likelihood becomes unstable, and sometimes infeasible, when the predictors are numerous or correlated. We propose a regularized maximum likelihood framework for simultaneous parameter estimation and feature selection in MoE whose experts belong to the generalized linear model family, covering Gaussian, Poisson and multinomial responses within a single formulation. Sparsity is induced in both the gating network and the experts through $\ell_1$ penalties, and the penalized log-likelihood is maximized by a proximal Newton-EM algorithm whose M-step reduces to weighted Lasso problems with closed-form coordinate-ascent updates. Unlike existing penalized MoE procedures, the algorithm requires neither a local quadratic approximation of the penalty nor any matrix inversion, it returns exactly sparse estimates without thresholding, and a proximal Newton-type variant guarantees a monotone increase of the penalized objective at every iteration. On simulated data and five real data sets, the method recovers the actual sparsity support and delivers prediction and clustering accuracy that is competitive with, and often better than, state-of-the-art regularized MoE. The source codes of our developed algorithms and their documentation are publicly available on Github at https://github.com/nv-thin/GLM-RMoE.2019-07-14T10:58:31ZThin Nguyen-VanFaicel ChamroukhiHa Hoang VanBao Tuyen Huynhhttp://arxiv.org/abs/2609.09573v1Geometric organization of olfactory descriptor data in the Poincaré disk2026-09-09T00:54:34ZOdor quality is commonly represented using high dimensional descriptor profiles, yet their low dimensional organization remains unclear. We investigated whether a two-dimensional hyperbolic embedding can provide an interpretable representation of this structure. We applied hyperbolic metric multidimensional scaling to two complementary datasets: 480 Sagar rating profiles from three participants rating 160 odorants on 15 continuous descriptors, and 4983 GoodScents--Leffingwell molecules annotated with 138 binary descriptors. The embeddings substantially preserved pairwise descriptor distances, supporting subsequent analyses of radial and angular organization. In Sagar, rating profile entropy was strongly and negatively associated with hyperbolic radius, with diffuse profiles closer to the center and concentrated profiles closer to the boundary. This radial organization emerged primarily at the level of the full descriptor profile, rather than any individual descriptor, and remained robust across alternative descriptor representations, participant specific analyses, and averaged ratings. Sweet, musky, fruity, pleasantness showed the strongest directional trends. In GoodScents--Leffingwell, active label entropy, reflecting descriptor multiplicity, increased with radius, whereas orthogonalized descriptor entropy, reflecting spread across orthogonal modes, decreased with radius. Related binary descriptors occupied coherent localized high-density regions. These findings reveal complementary radial and angular organization in the hyperbolic representation of olfactory descriptor data. They support hyperbolic mapping as an interpretable descriptive framework in which radius summarizes global profile properties, while the angular component captures continuous descriptor gradients and categorical organization.2026-09-09T00:54:34ZSubmitted to Chemical SensesAniss Aiman MedbouhiFarzaneh TalebGiovanni Luca MarchettiDanica Kragichttp://arxiv.org/abs/2609.09516v1Differential Privacy Guarantees in Small Area Estimation2026-09-08T22:57:14ZStatistical agencies increasingly rely on small area estimation to produce reliable estimates for subpopulations with limited sample sizes. These estimates are built from individual survey responses, so agencies must ensure that releasing them does not reveal information about any single respondent. We show that when a single draw from the posterior distribution of the Bayesian Fay-Herriot model is released, pure $\varepsilon$-differential privacy is unattainable, but the release satisfies formal privacy guarantees under Rényi differential privacy and zero-concentrated differential privacy without any noise being added, provided we treat the variance components as fixed. The key insight is that the posterior draw equals the posterior mean plus the Gaussian noise whose variance equals the posterior variance. The guarantee is thus governed by the sensitivity of the direct survey estimate and the posterior variance, and applies equally to a release of the posterior mean with that amount of noise added. For binary outcomes estimated with the Hájek estimator, the sensitivity equals the largest survey weight in the area divided by the sum of the weights. For the intercept-only model we derive exact coefficients describing how a change in one record propagates to every area's posterior mean, giving finite-sample per-area guarantees and a joint guarantee for releasing all areas at once that exceeds the largest per-area guarantee by at most a few percent in our applications. Two applications, poverty prevalence across 2,462 Public Use Microdata Areas in the American Community Survey and smoking prevalence across 52 substrata in the Washington state Behavioral Risk Factor Surveillance System, show that the guarantee is driven far more by the inequality of the survey weights than by the sample size, and that the shrinkage of the model tightens it substantially.2026-09-08T22:57:14ZSoumojit DasJörg Drechslerhttp://arxiv.org/abs/2609.09389v1A Counterfactual Framework for Estimating Infectious Disease Prevalence under Repeated Testing with Symptomatic and Contact-Tracing Components2026-09-08T19:40:48ZThis paper addresses the problem of estimating infectious disease prevalence under longitudinal testing programs that include scheduled, symptomatic, and contact-tracing testing. Our study is motivated by data from The Ohio State University, where a mandatory once-per-week COVID-19 testing and isolation program was implemented during the Fall 2020 semester, supplemented by additional testing for symptomatic individuals and identified contacts. In this setting, the probability of being tested depends on symptoms or contact-tracing status, creating a complex observation process. We develop a counterfactual framework that links the observation process to a hypothetical process in which infection is prevented. This formulation enables unbiased estimation of disease prevalence by modeling the testing process, possibly nonparametrically, without requiring explicit modeling of transmission dynamics, even though the testing and infection processes are jointly dependent.2026-09-08T19:40:48ZAccepted for publication in The Annals of Applied StatisticsJeongjin LeeJunke YangGrzegorz A. RempalaPatrick M. Schnellhttp://arxiv.org/abs/2603.22188v2Generalized Sequential Monte Carlo Sampling for Redistricting Simulation2026-09-08T19:26:56ZSimulation methods have become important tools for quantifying partisan and racial bias in redistricting plans. We generalize the Sequential Monte Carlo (SMC) algorithm of McCartan and Imai (2023), one of the commonly used approaches. First, our generalized SMC (gSMC) algorithm can split off regions of arbitrary size, rather than a single district as in the original SMC framework, enabling the sampling of multi-member districts with a varying number of representatives. Second, the gSMC algorithm can operate over various sampling spaces, providing additional computational flexibility. Third, we derive optimal-variance incremental weights and show how to compute them efficiently for each sampling space, leading to more efficient sampling. Finally, we propose a hybrid gSMC-MCMC algorithm by incorporating Markov chain Monte Carlo (MCMC) steps to handle large-scale redistricting applications without changing the target distribution. We demonstrate the effectiveness of the proposed methodology through analyses of the Irish Parliament, which uses multi-member districts of varying sizes, and the Pennsylvania House of Representatives, which has more than 200 single-member districts.2026-03-23T16:48:43ZPhilip O'SullivanKosuke ImaiCory McCartanhttp://arxiv.org/abs/2609.09325v1Kalman Filtering and Smoothing for Improving Precision in Horvitz--Thompson Estimation of Infectious Disease Prevalence2026-09-08T18:10:48ZHorvitz--Thompson (HT) estimators can provide unbiased daily estimates of infectious disease prevalence under repeated surveillance by correcting for nonrandom testing induced by scheduled, symptom-based, and contact-tracing components. However, because each HT estimate is based on the testing data available for that day and may involve highly variable inverse probability weights, it can be noisy, have precision that varies over time, and become unavailable during temporary interruptions in testing. The daily HT estimator is modeled as a noisy observation of an underlying prevalence process, with day-specific observation variances estimated using a delete-a-group jackknife. Our primary specification is a joint local linear trend state-space model that extends the standard level-only random walk by adding a latent slope. The Kalman filter improves precision by borrowing information from past estimates. At each time $t$, the process variances are estimated or carried forward using only observations available through time $t$, so the resulting filtered estimate is available in real time. We also describe the corresponding Kalman smoother as a retrospective extension based on the full observed series. When daily HT estimates are missing, the Kalman filter proceeds through prediction-only updates, whereas the corresponding smoother retrospectively reconstructs those periods using later observations. In simulations, the joint Kalman filter substantially improves precision relative to the raw daily HT estimator while preserving the main temporal pattern, and the smoother provides a more stable retrospective summary. In The Ohio State University's fall 2020 SARS-CoV-2 surveillance data, the filter and smoother provide estimates on no-testing days, when the HT estimator provides neither point nor interval estimates, and yield narrower confidence intervals than HT intervals on observed days.2026-09-08T18:10:48ZAbstract shortened to comply with arXiv's 1,920-character limitJeongjin LeeGrzegorz A. RempalaPatrick M. Schnellhttp://arxiv.org/abs/2604.10641v2On the Capacity of Distinguishable Synthetic Identity Generation under Face Verification2026-09-08T18:05:03ZSynthetic face generators can produce many nominal identities, but nominal count does not determine how many are jointly distinguishable under a specified verification rule. We define finite-dimensional capacity as the supremum of codebook sizes over distinct latent identity codes whose induced identity-conditional embedding distributions satisfy per-identity genuine acceptance and pairwise impostor non-match constraints. For deterministic view-invariant pipelines, fixed-code capacity equals the spherical-code cardinality over the realizable embedding set and reduces to the classical spherical-code cardinality when every sphere direction is realizable. For stochastic identity-conditional embedding distributions concentrated with probability at least $1-η$ in spherical caps of angular radius $ρ$, we derive a sufficient center-separation condition, spherical-code capacity lower bounds under full angular expressivity, and positive asymptotic lower-bound exponents for dimension-indexed pipeline families. We also derive prior-constrained random-code lower bounds from pairwise center-separation failure probabilities. When each identity-conditional embedding distribution has support equal to a spherical cap of angular radius $ρ$, we derive necessary zero-error geometric conditions and, for $2ρ<\arccos(τ)$ under full $ρ$-cap angular expressivity, show that the restricted zero-error capacity equals the classical spherical-code cardinality at minimum angle $\arccos(τ)+2ρ$. For finite repeated-view samples, a maximum clique in the resulting compatibility graph identifies the largest sampled subset satisfying all empirical genuine and pairwise impostor constraints. We evaluate this sample-restricted quantity on a deterministically selected DigiFace-1M subset under three fixed recognizers with identity-disjoint in-domain threshold calibration.2026-04-12T13:42:39ZBehrooz Razeghihttp://arxiv.org/abs/2609.09039v1Covariate Adjustment in Randomized Experiments: A Unified Framework for Decision and Practice2026-09-08T17:03:57ZShould researchers adjust for covariates in randomized experiments, and if so, how? The literature offers three distinct prescriptions: do not adjust because randomization guarantees unbiasedness; adjust for outcome-prognostic covariates to improve precision; or adjust for covariates imbalanced between treatment arms. These competing prescriptions create confusion and uncertainty. We develop a unified framework for decision and practice. Given available information, we show that the optimal correction is what we call ex-post bias. The only relevant criterion for adjustment is prognosticity for ex-post bias; neither raw covariate imbalance nor outcome prognosticity is sufficient by itself. We also show that correcting imbalance and improving precision are two sides of the same decision problem. We develop two estimation approaches, one of which recovers familiar adjustment estimators and provides a new theoretical justification for them. Simulations compare alternative covariate-selection and adjustment strategies. Overall, our framework provides a unified foundation for covariate adjustment in randomized experiments.2026-09-08T17:03:57ZJiawei FuDonald P. Greenhttp://arxiv.org/abs/2608.22223v2Spending Scarce Confirmatory PET Measurements: Target-Aligned Validation in A4/LEARN2026-09-08T16:30:30ZAnti-amyloid therapies and blood-based biomarkers are changing Alzheimer disease workups into a two-stage measurement workflow: screen broadly with cheaper information, then spend scarce confirmatory amyloid measurements where they support the decision that will be reported. Amyloid positron-emission tomography (PET) remains one such protocol measurement for amyloid burden, but PET slots, trial budgets, and payer-facing evidence packages are finite. This paper asks a deliberately operational question: when is simple transparent PET validation enough, and when is a fitted residual-uncertainty score worth the added complexity? For a weighted protocol target, the first-order value of validating subject i is the product of target influence and residual protocol uncertainty. Generic uncertainty sampling uses only the second factor and can spend PET measurements on subjects that are hard to predict but weak for the scientific, clinical, or commercial claim. We apply this rule to the A4/LEARN PET archive, treating observed PET as a design laboratory for scarce-confirmation studies. For the primary APOE4 carrier versus non-carrier contrast in Centiloid 24-or-higher PET positivity, simple APOE4-balanced validation recovers nearly all of the target-specific gain: at PET budget 200, the confidence-interval width ratio relative to random validation is 0.923 for APOE4 balancing and 0.914 for target-specific scoring, while generic uncertainty sampling is 0.980. Other targets behave differently: target-specific scoring gives larger gains for an age-slope analysis and for cutoff-indexed PET positivity. The practical message is simple: spend scarce protocol measurements according to the claim being validated, not only according to prediction uncertainty.2026-08-23T05:17:12Z11 pages, 3 figures, 1 table; supplementary material and code/results are included as ancillary filesEliuvish Han Cui