https://arxiv.org/api/yU4cr/zYWWDPShUBCacrIYF6DMk 2026-09-11T21:03:35Z 24226 60 15 http://arxiv.org/abs/2609.07186v1 Bayesian nonparametric inference for modal missing species and features 2026-09-07T08:15:26Z Species and feature sampling problems arise naturally whenever each observed unit is associated with one or more labels from a countable alphabet, and inference focuses on the unobserved portion of the distribution over such an alphabet. Within this framework, recent contributions have shifted attention from inference on the total probability mass over unseen labels to distribution-free confidence intervals for the largest unobserved label probability (i.e., the modal missing probability), thereby providing more refined information on whether the missing mass is concentrated on a few high-prevalence unseen labels or distributed across many negligible ones. Besides lacking a unified framework for species and features, these approaches employ a worst-case perspective that causes substantial information loss and produces overly-conservative intervals. We address these limitations through a unified model-based framework for inference on modal missing probabilities in both species and feature settings, which leverages a flexible Bayesian nonparametric formulation to localize uncertainty around models compatible with the observed data. This leads to sharper closed-form credible intervals that effectively exploit prior information and observed data, while preserving the theoretical frequentist properties and robustness of distribution-free intervals. Simulation studies confirm these improvements, while an organized crime application illustrates how our contribution has the potential to reshape law-enforcement decision-making in investigations. 2026-09-07T08:15:26Z Alessandro Colombi Mario Beraha Daniele Durante Stefano Favaro http://arxiv.org/abs/2603.18074v2 Adapting Technical-Service LLM Agents with Latent Logic Augmentation, Robust Noise Reduction, and Hybrid Reward Modeling 2026-09-07T06:25:56Z Technical-service LLM agents are entering production workflows, where value depends on whether engineers adopt generated replies. Service tickets hide decision logic, contain noisy single-reference responses, and make reward evaluation costly, making standard post-training brittle. Existing post-training and LLM-as-a-Judge approaches improve grounding or feedback, but do not jointly model latent decision logic, response diversity, and reward cost. We address this gap by coupling latent logic augmentation, robust noise reduction, and hybrid reward modeling. The framework augments supervised fine-tuning data with Planning-Aware Trajectory Modeling and Reasoning Augmentation, builds dual-filtered Multiple Ground Truths, and trains the policy with a hybrid reward that combines a Reranker with an LLM-as-a-Judge. On real Cloud technical-service tasks, the adapted Qwen3-4B achieves the highest Multi-ECS (0.441), lower reward cost, and the highest production adoption rate (46.63%). 2026-03-18T05:01:17Z 36 pages, 6 figures, 14 tables. Camera-ready version accepted to the EMNLP 2026 Industry Track. Title, author order and metadata, experiments, analysis, references, and appendices have been updated; the set of authors is unchanged Junzhuo Ma Chenghuang Shen Yi Yu Xingyan Liu Jing Gu Hangyi Sun Guangquan Hu Jianfeng Liu Weiting Liu Pu Mingyue Wang Yu Zhengdong Xiao Rui Xie Longjiu Luo Qianrong Wang Gurong Cui Honglin Qiao Wenlian Lu http://arxiv.org/abs/2609.07033v1 Choosing the Dictionary and Penalty for IV-LASSO 2026-09-07T04:41:41Z Estimating the first stage of an instrumental variables (IV) model with the least absolute shrinkage and selection operator (LASSO) requires choosing a dictionary of technical instruments and a penalty level. First-order asymptotic theory offers no guidance on these choices, as any consistent implementation yields a structural parameter estimator with the same limiting distribution. In finite samples, however, these choices can have a substantial impact on the resulting structural parameter estimate. Working in a model with a single endogenous regressor and homoskedastic Gaussian errors, we use first- and second-order Stein identities to derive the approximate mean squared error (AMSE) of the instrumental-variables LASSO (IV-LASSO) estimator, which can be consistently estimated and used to rank a prespecified list of dictionary-penalty candidates. The AMSE reveals a bias-variance trade-off: more complex first-stage fits better approximate the conditional mean of the endogenous variable but are also more correlated with the structural errors, with complexity measured by the degrees of freedom of the LASSO fit. The weight on this bias rises with the endogeneity of the regressor, a quantity that neither plug-in nor cross-validation penalty rules take into account. Despite the AMSE being derived in a Gaussian model, penalty selection by minimizing the feasible AMSE criterion delivers up to a one-third lower mean squared error compared to cross-validation and plug-in penalty rules in Gaussian and non-Gaussian simulation designs calibrated to the data of Gilchrist and Sands (2016). 2026-09-07T04:41:41Z Yukun Ma Manu Navjeevan Bogdan Salahub http://arxiv.org/abs/2607.27678v2 Generalized Bayesian Inference using the Bayesian Bootstrap for Survival Models 2026-09-07T04:15:13Z Survival inference often requires uncertainty quantification under censoring and limited sample sizes, while prior information may be available but difficult to incorporate without specifying a full likelihood. Likelihood-based Bayesian survival methods provide prior-informed inference but may be sensitive to distributional assumptions and computationally demanding for complex survival models. The Bayesian bootstrap provides a likelihood-free, nonparametric approach to uncertainty quantification, but by itself does not provide a general mechanism for incorporating priors on model parameters. We propose a general Bayesian method for survival models that combines the generalized Bayesian (Gibbs) updating with Bayesian bootstrap. The Bayesian bootstrap generates a distribution over a target survival estimator through Dirichlet weights, while generalized Bayesian updating incorporates parameter-specific prior information through a loss function. The resulting posterior provides a flexible alternative to likelihood-based Bayesian inference and can be applied to a broad class of survival estimators. We develop the method for the Cox proportional hazards model and obtain posterior inference for regression coefficients and hazard ratios. Simulation studies demonstrate uncertainty quantification and prior-data learning across varying sample sizes. An application to right-censored survival data illustrates its practical utility. The methodology is implemented in the open-source \texttt{R} package \texttt{BayesBoots}. 2026-07-30T04:49:14Z K Shuvo Bakar Armando Teixeira-Pinto http://arxiv.org/abs/2609.06827v1 Beyond Aggregate VARs: A Bayesian Benchmark for HANK Models 2026-09-06T20:48:01Z Heterogeneous-agent New Keynesian (HANK) models characterize how entire cross-sectional distributions respond to structural shocks. Traditional representative-agent models are routinely disciplined by impulse responses from aggregate vector autoregressions (VARs). HANK models have no comparable established empirical benchmark because they make predictions not only about aggregates, but also about distributions of micro-level data. We propose a Bayesian benchmark that jointly models macroeconomic aggregates and several marginal distributions from repeated cross sections, including distributions observed in different surveys. Our approach can use both standard structural VAR identification approaches on macroeconomic aggregates and identification restrictions imposed on micro-level data. The model delivers a joint posterior of the distributional effects of shocks, without the need for household panel data or a separate first-stage density estimate. 2026-09-06T20:48:01Z Florian Huber Gary Koop Christian Matthes http://arxiv.org/abs/2609.06739v1 The profit-bias identity in sports betting: bookmaker profit as the public's prediction error 2026-09-06T17:22:25Z Sports betting moves money continuously from a large public to a small number of firms. The most influential account of that flow, due to Levitt (2004), holds that books price away from the market-clearing point to exploit predictable public biases. It computes profit from two numbers (the probability that a side wins the proposition and the fraction of handle it attracts), treating the share on a side as independent of the outcome. Here we relax that assumption and derive a profit-bias identity: profit is affine and increasing in the expected share of handle on the losing side, with Levitt's expression as the special case of independence. The identity resolves the book's margin into exactly three channels: its hold, the product of its price shading and the public's lean, and the covariance between bet share and outcome. A lean is thus worthless without shading, and shading is worthless without a lean. Under a public-belief model, the profit driver is the public's Bayes error: the probability of a representative bettor selecting the losing side. We provide necessary and sufficient conditions for a "Goldilocks Zone": prices at which book and bettor both profit. Testing these predictions on 1,139 Major League Baseball games, we find that the apparent dependence between bet share and outcome is a Simpson's paradox: present when games are pooled, but absent once they are separated by which side the book favored. The public leans heavily toward favorites, but we detect no matching shading, and the realized margin is indistinguishable from the hold. 2026-09-06T17:22:25Z Jacek P. Dmochowski http://arxiv.org/abs/2603.17866v3 NFL step-and-turn: A generative framework for evaluating player movement in American football 2026-09-06T15:29:17Z In sports analytics, player tracking data have driven significant advancements in the task of player evaluation. We present a novel generative framework for evaluating the observed frame-by-frame player positioning against a distribution of hypothetical alternatives. We illustrate our approach by modeling the within-play movement of an individual ball carrier in the National Football League (NFL). Specifically, we develop Bayesian multilevel models for frame-level player movement based on two components: step length (distance between successive locations) and turn angle (change in direction between successive steps). Using the step-and-turn models, we perform posterior predictive simulation to generate hypothetical ball carrier steps at each frame during a play. This enables comparison of the observed player movement with a distribution of simulated alternatives using common valuation measures in American football. We apply our framework to tracking data from the first nine weeks of the 2022 NFL season and derive novel player performance metrics based on hypothetical evaluation. 2026-03-18T15:54:28Z Quang Nguyen Ronald Yurko http://arxiv.org/abs/2501.06180v2 The role of probabilistic load and renewable prediction in enhancing day-ahead electricity price forecasts 2026-09-06T07:00:34Z The increasing penetration of renewable energy sources (RES) has amplified volatility in electricity markets, while demand variability continues to challenge system stability. Traditional day-ahead electricity price forecasting (EPF) models rely on point forecasts of load and RES generation, which fail to capture the uncertainty inherent in weather-driven supply and dynamic demand. This study introduces a novel framework that integrates probabilistic forecasts of load, wind, and solar generation into EPF models. Using data from the German EPEX market between 2015 and 2023, we generate quantile forecasts of fundamental variables via historical simulation, conformal prediction and quantile regression, and incorporate them into both parsimonious expert and high-dimensional LASSO-based models. Empirical results show that probabilistic inputs substantially enhance forecast accuracy, reducing root mean square errors by up to 13% compared with point-forecast benchmarks. Notably, extreme quantiles of load and RES forecasts emerge as the most influential predictors, underscoring the importance of rare but system-critical scenarios. These findings demonstrate that accounting for uncertainty in both load and renewables is crucial for reliable electricity price forecasting and offers practical value for system operators, traders, and policymakers navigating renewable integration. 2025-01-10T18:58:38Z Renewable Energy, Volume 269, 125844, 2026 Bartosz Uniejewski Florian Ziel 10.1016/j.renene.2026.125844 http://arxiv.org/abs/2609.06413v1 Shrinkage invalidates the Hosmer-Lemeshow test: goodness of fit for penalized logistic regression, with an application to glaucoma diagnosis 2026-09-06T06:13:48Z Clinical prediction models are increasingly fitted by penalized logistic regression, because collinearity or many candidate predictors makes maximum likelihood unstable or impossible. Calibration is then almost always assessed by a grouped goodness-of-fit test such as the Hosmer-Lemeshow test. We show that this combination is invalid. Under ridge regression the grouped standardized residuals acquire a non-centrality induced by shrinkage, so the reference distribution used in practice is wrong, and at the penalty that most improves the fitted probabilities the test rejects correctly specified models between 92 and 100 per cent of the time. We derive the corrected law and define the shrinkage-corrected Hosmer-Lemeshow test, which subtracts an estimate of that non-centrality, restoring the maximum likelihood reference exactly to first order, and is made valid by prepivoting at a power cost we measure. We also give the attenuation law governing what any such test can detect once the linear predictor must be estimated. In glaucoma diagnosis by confocal laser tomography, where the maximum likelihood estimate does not exist, the corrected test finds the evidence for misfit weaker by more than three orders of magnitude: the fitted risks are too flat rather than mis-ordered, so the model needs recalibration rather than rebuilding. 2026-09-06T06:13:48Z 18 pages, 4 figures, 5 tables; supplementary material (13 pages) included as an ancillary file. Submitted to the Journal of the Royal Statistical Society Series C. Reproduction archive: https://doi.org/10.5281/zenodo.21903202. Methods implemented in the R package ebrahim.gof (CRAN), function shrink.gof() Ebrahim Khaled Ebrahim http://arxiv.org/abs/2609.06267v1 From Discrete Trailing Returns to a Continuous Graphical Profile: Return-to-Present Curves 2026-09-05T21:28:26Z Investment performance is commonly presented either as a conventional cumulative-return chart, which fixes a historical starting date and traces performance forward, or as a trailing-return table, which fixes the current endpoint but reports only a small set of prespecified horizons. These two displays have complementary limitations: fixed-start comparisons are conditional on the selected origin, whereas trailing returns provide only discrete snapshots of the underlying fixed-endpoint return function. We present the Return-to-Present (RTP) curve as a continuous fixed-endpoint representation that brings these perspectives together by holding the evaluation date fixed while allowing the hypothetical historical purchase date to vary over the available history. Familiar 1-month, 3-month, 6-month, 1-year, and longer trailing returns therefore become selected points on a continuous curve. When multiple investments are overlaid, RTP directly displays entry-date sensitivity, persistent relative advantage, crossings, and the timing and magnitude of separation without requiring selection of a single historical origin. The same endpoint-based construction naturally accommodates investments with unequal inception dates, recurring purchases, and retrospective portfolio rotation decisions in which sale and replacement-purchase dates may differ. We illustrate these uses with real investment data and discuss its relationship to momentum. RTP does not define a new return measure; its contribution is a simple graphical organization of familiar realized returns for historical comparison and decision support rather than prediction or statistical inference. 2026-09-05T21:28:26Z Lei Liu http://arxiv.org/abs/2602.06301v2 Design-Conditional Prior Elicitation for Dirichlet Process Mixtures: A Unified Framework for Cluster Counts and Weight Control 2026-09-05T20:17:47Z Dirichlet process mixtures describe variation in effects or latent traits across sites, studies, or examinees. The concentration parameter controls the number and relative sizes of the clusters, but its hyperprior is often chosen by default. We describe a method for choosing this hyperprior when the number of units is fixed by design. Analysts specify an expected number of clusters and their uncertainty about that count. We translate these judgments into a hyperprior and examine what it implies about the sizes of the clusters. When the count and size judgments conflict, the Dual-Anchor procedure chooses and reports the trade-off. Simulations show that the default hyperprior can favor a single cluster when estimates for individual units are noisy. Applications to a multisite trial, a meta-analysis, and a vocabulary test show sensitivity in heterogeneity summaries despite relatively small changes in individual estimates. The DPprior R package implements the calibration and diagnostics. 2026-02-06T01:49:16Z JoonHo Lee http://arxiv.org/abs/2609.05970v1 Machine Learning for Pre-Culture ESBL Risk Stratification to Guide Empiric Antibiotic Selection: A 12-Hospital Study of Enterobacteriaceae Cultures 2026-09-05T08:19:41Z Empiric antibiotic therapy for suspected ESBL-producing Enterobacteriaceae must be selected 48-72 hours before culture results, forcing clinicians to choose between undertreating resistant infections and overusing carbapenems that drive further resistance. We developed a cost-sensitive XGBoost model predicting an ESBL phenotype (resistance to ceftriaxone, ceftazidime, cefepime or piperacillin-tazobactam) at culture ordering using 45 pre-culture EHR features across 132,955 cultures from 72,217 patients at 12 hospitals (14.41% with the ESBL phenotype). Cultures were partitioned at the patient level. At 90% sensitivity, the model achieved 95.8% NPV, reducing post-test ESBL probability to 4.2%, a threshold that may support safe carbapenem-sparing in non-ICU settings, while sparing 307 of every 1,000 cultures an unnecessary broad-spectrum course at the cost of 14 missed ESBL cases per 1,000. SHAP analysis identified prior ESBL colonization as the dominant predictor, ahead of prior organism burden and neighborhood deprivation; removing deprivation features caused minimal performance loss ($Δ\text{AUROC} = -0.020$), enabling equitable bedside deployment. Discrimination was unchanged under a strict IDSA ESBL-E definition (AUROC 0.766), with specimen type added as a predictor (0.764) and without any class-imbalance correction (0.762), and ranged from 0.71 to 0.78 across organism strata. 2026-09-05T08:19:41Z Accepted for presentation at American Medical Informatics Association (AMIA) Annual Symposium 2026 Aravind V. Kuruvikkattil Lalitha Pranathi Pulavarthy Rashmita Kudamala Saptarshi Purkayastha http://arxiv.org/abs/2609.05941v1 Moment-Matching Probabilistic Data Association for Optimization-Based SLAM 2026-09-05T07:19:31Z Optimization-based simultaneous localization and mapping (SLAM) makes it possible to reduce accumulated navigation errors of sensing platforms by returning to known areas (loop closure). In this paper, we present an approach to combine probabilistic data association (PDA) with optimization-based SLAM. Instead of associating a single measurement with each landmark, we follow the PDA paradigm from the multiobject tracking community. In particular, in a processing stage performed in addition to the nonlinear least-squares solver of optimization-based SLAM, our method (i) assigns multiple measurements to landmarks probabilistically, (ii) computes the mean and covariance of landmark distributions via moment matching by taking multiple measurement-to-landmark associations into account, and (iii) establishes a virtual landmark measurement and a corresponding linear-Gaussian measurement model that leads to the mean and covariance matrix as moment-matching PDA in (ii). By converting the PDA update step into an equivalent linear-Gaussian measurement update step, PDA can be performed effectively within any optimization-based SLAM method. Our preliminary numerical evaluation in a scenario with false negatives and false positives indicates that incremental smoothing and mapping 2 (iSAM2), combined with the proposed PDA approach, can improve agent localization performance compared to conventional iSAM2. 2026-09-05T07:19:31Z 7 pages, 3 figures, ISIF FUSION 2026 Conference Khoa Nguyen Mitchell Turton Florian Meyer http://arxiv.org/abs/2609.05784v1 A Network-Structured Bayesian Hierarchical Model for Sparse Mutation-Drug Response Associations: Application to Cancer Pharmacogenomics 2026-09-05T00:37:22Z We develop a network-structured Bayesian hierarchical model for sparse association mapping between genomic alterations and quantitative treatment-response phenotypes. The framework combines a Gaussian Markov random field prior that borrows strength across pathway-connected genes, a global-local horseshoe prior inducing sparsity, and a conjugate Gibbs sampler requiring no Metropolis-Hastings steps. Though broadly applicable to high-dimensional settings with known predictor networks, we validate it using cancer cell-line drug-sensitivity data. Applied to GDSC2 ($N=951$ cell lines, $G=219$ driver genes, $D=295$ drugs), the model identifies 126 gene-drug associations (0.195\% of 64{,}605 pairs), concentrated in EZH2 (45 drugs, all sensitivity-direction, mean effect $-0.911$ $\ln$IC50) and KMT2D (36 drugs, all sensitivity-direction, mean effect $-0.496$ $\ln$IC50). These markers show external support in an independent PRISM screen (1{,}518 compounds), with KMT2D achieving complete directional replication (36/36) and EZH2 partial replication (8/12). Five-fold cross-validated predictive log-likelihood confirms each prior layer's value: the full model outperforms the no-network ablation by $+3{,}109$ log-units per fold and the no-horseshoe ablation by $+14{,}039$ log-units, consistently across folds. Simulations under three scenarios show the full model achieves the highest precision and lowest false-discovery rate throughout, while the network prior improves sensitivity recovery under network-structured signal. A tissue-stratified extension identifies coherent subgroup refinements, including lung-specific EGFR-inhibitor sensitivity and skin-specific BRAF-Dabrafenib sensitivity. These results show the framework identifies sparse, interpretable, externally supported drug-sensitivity markers while enabling principled investigation of tissue-specific departures from shared effects. 2026-09-05T00:37:22Z The manuscript is under review with Biostatistics (Oxford Academic) Hammed A. Olayinka Saheed O. Olayemi http://arxiv.org/abs/2508.11920v2 A Wavelet-Based Framework for Mapping Long Memory in Resting-State fMRI: Age-Related Changes in the Hippocampus from the ADHD-200 Dataset 2026-09-04T23:32:37Z Functional magnetic resonance imaging (fMRI) time series are known to exhibit long-range temporal dependencies that challenge traditional modeling approaches. In this study, we propose a novel computational pipeline to characterize and interpret these dependencies using a long-memory (LM) framework, which captures the slow, power-law decay of autocorrelation in resting-state fMRI (rs-fMRI) signals. The pipeline involves voxelwise estimation of LM parameters via a wavelet-based Bayesian method, yielding spatial maps that reflect temporal dependence across the brain. These maps are then projected onto a lower-dimensional space via a composite basis and are then related to individual-level covariates through group-level regression. We applied this approach to the ADHD-200 dataset and found significant positive associations between age in children and the LM parameter in the hippocampus, after adjusting for ADHD symptom severity and medication status. These findings complement prior neuroimaging work by linking long-range temporal dependence to developmental changes in memory-related brain regions. Overall, the proposed methodology enables detailed mapping of intrinsic temporal dynamics in rs-fMRI and offers new insights into the relationship between functional signal memory and brain development. 2025-08-16T05:55:32Z 29 pages, 24 figures Yasaman Shahhosseini Cédric Beaulac Farouk S. Nathoo Michelle F. Miranda