https://arxiv.org/api/lhAXdxzLCd6xYUXW23IhDDxObfY2026-09-11T20:44:48Z303244515http://arxiv.org/abs/2609.09855v1A Unifying Perspective on Probabilities as Model Predictions2026-09-09T08:07:09ZAlthough probabilistic statements are ubiquitous, foundational disagreements persist about their understanding, as exemplified by debates between Bayesians and frequentists; moreover, it is unclear when and why acting on them actually leads to desirable outcomes. Here, we argue that every probability is the output of a \emph{prediction method}, that is, it depends on both a particular way of constructing abstractions and a way of transforming them into predictions. Through this, we provide a unifying perspective on supposedly different kinds of probabilities and show that even supposedly objective ones are model-dependent. We demonstrate that when a finite calibration criterion is met, one can anticipate the distribution of utilities for a given policy and inform successful decision-making on finite sets of events. Based on the notion of prediction methods, inductive arguments, and the probability calculus, we explain the feasibility of the calibration criterion in many settings. Overall, we develop a coherent perspective on probabilities and their use, connecting key intuitions behind other interpretations along the way.2026-09-09T08:07:09ZBenedikt Höltgenhttp://arxiv.org/abs/2607.10747v4A Verifier-Centric Conceptual Model for Digital Credential Ecosystems2026-09-09T07:15:25ZDigital credential ecosystems increasingly combine multiple standards. Because implementations have evolved independently across jurisdictions and application domains, systems described under the common label ``digital credential'' often remain mutually non-interoperable. Conventional element-by-element comparisons of identifiers, data models, credential formats, protocols, and signature algorithms do not explain why interoperability fails even when stacks share a data model, nor do they identify what a verifier must obtain, and what it must trust, before accepting a credential. We present a verifier-centric conceptual model built on two decompositions. The first separates credential processing into signature verification (L1), semantic interpretation (L2), and validation (L3), and models the supporting materials through two orthogonal planes: Constitution, which captures ecosystem-level arrangements and trust declarations, and Logistics, which captures how verification materials are stored and delivered; the Shinken framework makes trust assumptions explicit across all five functions. The second characterizes where each function may be placed along three dimensions (placement, timing, and disclosure). From the condition of being verifiable, the model derives seven consequences, distinguished as definitional corollaries, operational implications, and design trade-offs. Applying the model to four learner-credential stacks and to existing ecosystems including authentication federations, we show that it explains interoperability failures, verifier-side burden, offline verifiability, privacy implications, and terminological ambiguities that element-wise comparison leaves unresolved.2026-07-12T12:59:26Z22 pages, 4 figures, 6 tables. A version of this paper has been accepted for publication in IEEE AccessShigeya SuzukiRyosuke Abehttp://arxiv.org/abs/2609.09687v1Chance, Persistent Advantage, and the Generative-AI Era in Open-Source Package Careers2026-09-09T04:03:18ZStudies of careers in science, film, music, and books report a common pattern. When a person's most successful work arrives is close to a random draw over the works they produce. How large their successes tend to be, in contrast, follows a stable, person-specific factor. We test whether this pattern holds for open-source software careers and whether it changed when generative AI coding tools arrived. From the complete public record of GitHub push events (2015-2025), we reconstruct 102.2M career works by 6.15M contributors, and for the 908k contributors whose repositories publish packages, we measure each work's impact by how many downstream packages come to depend on it. First, we find that the timing of a career's biggest hit is close to a lottery over their works, as in science and the arts, with a small, replicable lean toward early career that grows as careers get longer. Second, some coders reliably produce higher-impact work than others, but this lasting personal factor accounts for only part of why impact persists (about a fifth in our primary specification); the rest behaves like momentum, success feeding on itself for a period of time. Third, within the same contributors, this structure did not change after ChatGPT's release. The stable factor's weight grew by about as much as it grew for an earlier cohort that simply aged, and subtracting the effect of aging from the effect of generative AI puts the shift at +0.03 (95% CI [-0.22, +0.23]), indistinguishable from zero. The success pattern documented in science and the arts therefore describes open-source careers too, and it shows no detectable break across the arrival of generative AI. These results have implications for how track records on open platforms should be read and on what to expect from generative AI for the careers built on them.2026-09-09T04:03:18Z25 pages, 11 figuresHazem IbrahimYasir Zakihttp://arxiv.org/abs/2604.21938v2The Biggest Risk of Embodied AI is Governance Lag2026-09-09T02:30:24ZEmbodied AI is widely discussed as a job-displacement problem. The deeper risk, however, is governance lag: the time and capability gap between a measurable change in technology deployment and an institutional response able to address its consequences. Building on the established pacing problem and the Collingridge dilemma, this article argues that embodied AI intensifies that gap through scalable models and platforms, task-level reorganization, and the separation of upstream technological control from downstream social impact. We distinguish three mutually reinforcing forms of lag, observational, institutional, and distributive, and propose a compliance architecture based on deployment visibility, stack-level accountability, trigger-based adjustment, and automatic distributional response. The central policy challenge is not automation alone, but whether governance systems can become observable, responsive, and adaptive before disruption becomes entrenched.2026-04-07T03:56:14ZIEEE Computer Magazine, 2026Shaoshan Liuhttp://arxiv.org/abs/2609.09609v1Who You Are Adds Nothing Detectable to Where You Go Next: Sociodemographic Conditioning in LLM Next-Location Prediction2026-09-09T02:06:10ZLarge language models (LLMs) are increasingly used for individual next-location prediction, while sociodemographic conditioning is common in LLM-based travel simulation. Yet the incremental predictive value of sociodemographic attributes remains unclear. To directly test this contribution, sociodemographic records were linked with passively sensed mobility data from 5,000 Shenzhen residents to construct a closed-set benchmark in which models rank 100 candidate destinations. Each prediction instance is evaluated with and without age, gender, occupation and income, while holding mobility history, candidates and all other prompt content fixed. Results show that across four history lengths, the paired change in top-1 accuracy ranges from -0.8 to +0.5 percentage points, with no detectable gain from attributes. This result remains consistent when stay history is withheld, across alternative prediction times, in two additional LLMs and in a supervised reranker trained on the same benchmark. The null does not reflect a lack of model responsiveness to demographic information, as permuted attributes reduce LLM accuracy whereas correctly matched attributes do not improve it. A further asymmetry emerges in the reverse predictive direction, as pre-cut mobility trajectories recover income with an AUC of 0.708, while sociodemographic attributes contribute little to next-location prediction. Beyond demographic conditioning, candidate construction exerts a much larger influence on reported performance. Removing distance raises top-1 accuracy by 7.7 percentage points under proximity sampling but lowers it by 22.3 points under popularity sampling, with the reversal reproduced across all three LLMs. These results distinguish demographic association from incremental predictive usefulness and show that sampled next-location accuracy depends strongly on how candidate alternatives are constructed.2026-09-09T02:06:10Z17 pages, 4 figuresXin WangParaic CarrollKerry NiceSachith SeneviratneLi Zhanghttp://arxiv.org/abs/2609.09604v1Watermarks Without Verification: AI Text Watermarking After the EU AI Act2026-09-09T01:56:17ZOn August 2, 2026, the obligations of Article 50 of the EU AI Act took effect, requiring generative AI providers to mark the content their systems produce and ensure it can be detected as AI-generated. Days later, Anthropic disclosed that every Claude model released after that date embeds a watermark based on SynthID-Text in all generated text, enabled by default with no user opt-out; Google has deployed SynthID-Text in Gemini since 2024. Users objected that the watermark degrades quality, particularly for code, that it secretly encodes identifying information, and, in mutual contradiction, that it is easily removable and inescapable; the vendor answered with assurances of unchanged quality, no identifying information, and robustness to light editing. In this work, we argue that neither the objections nor the assurances can currently be verified and that this unverifiability, rather than watermarking itself, is the substantive governance failure. We sort the contested assertions by what it would take to settle each and evaluate the open-source SynthID-Text implementation on two open-weight models, because no public tool can test the deployed systems. On prose, the measured effect of the watermark does not exceed that of changing the sampling seed. On code, the cost is three points of correctness on one model and below measurement on the other, while detection remains near chance, a limitation of detectability rather than quality. The remaining gaps trace to withheld access or missing institutions and we map each to a requirement: release of matched outputs, configuration disclosure, accredited audits, a shared evaluation protocol, and interoperable detection.2026-09-09T01:56:17Z10 pages, 2 figuresAlexander NemecekVipin ChaudharyErman Aydayhttp://arxiv.org/abs/2609.09533v1Scalable Oversight for AI in Mental Health: Lessons from 350,000 AI Coaching Conversations between Therapy Sessions2026-09-08T23:27:05ZClinician review of every AI output is often proposed as a safeguard in mental healthcare, but vigilance research suggests this approach fails at scale and may paradoxically reduce safety. Drawing on our experience deploying an AI coaching tool across 350,000+ conversations between therapy sessions, we describe how we arrived at a three-layer human-on-the-loop oversight framework combining preventive design, real-time monitoring, and continuous clinician evaluation. We show how specific findings from clinical review drove iterative improvements, and offer practical recommendations for mental health professionals evaluating AI systems.2026-09-08T23:27:05Z2 figuresMatthew A. ScultJohn L. HavlikKevin RamotarEthan GohManoj Kanagarajhttp://arxiv.org/abs/2609.09496v1The Mutations of Machine Speech2026-09-08T22:23:38ZAlgorithmic outputs now populate the digital environments through which contemporary life is organized. The role of law in facilitating and constituting (rather than merely responding to) these processes is gaining increasing traction across scholarly accounts. This inquiry traces the evolution of algorithmic outputs attending to their legal underpinnings and social implications, surfacing the mutations of machine speech.
The first mutation redefined speech as data to be queried: search engines transformed the web from a space of information retrieval into an economic regime of algorithmic visibility. The second mutation reframed speech as engagement: social media platforms fused moderation with amplification, turning expression into a metric of attention, governed by corporate architectures. The third mutation emerges in conversational systems and interfaces, where generative text displaces information retrieval, bringing with it dense technolegal entanglements and profound epistemic consequences.
Scholars of freedom of expression, informational privacy, and communication studies have long grappled with these dynamics, yet their implications for broader legal thought have also become urgent. This piece seeks to organize and clarify the evolving debate around algorithmic speech, making this critical but often fragmented discourse more accessible to wider legal and interdisciplinary audiences. In doing so, it bridges the gap between observing technological transformation and critically assessing the constitutive role of law within it, offering a conceptual resource for researchers, students, policymakers, and practitioners navigating and contesting this evolving landscape.2026-09-08T22:23:38ZMauricio Figueroa10.1007/978-3-031-87993-7_188-1http://arxiv.org/abs/2502.03682v3Towards On-Device Evidence Gathering for Intimate Partner Infiltration: A Feasibility Study for Joint Identity-Action Detection2026-09-08T22:14:21ZIntimate Partner Infiltration (IPI) refers to phone-side privacy infiltration in intimate or close relationships, often enabled by physical access to a person's smartphone and discussed in technology-facilitated Intimate Partner Violence (IPV) contexts. Unlike conventional cyberattackers, IPI perpetrators leverage proximity and personal knowledge to circumvent standard protection, underscoring the need for targeted interventions, motivating device-side tools that surface such risk evidence for later review. While prior works have extensively studied IPV, and some have provided tailored and effective solutions such as security clinics, they are necessarily episodic and human-expert-intensive, and offer limited automated visibility into what happens on a smartphone between support sessions. Guided by a formative interview with experts (n=5), we take the first exploration into gathering IPI-risk evidence from a mobile system perspective and present AID, Automated IPI Detection, a data-driven system that continuously logs unauthorized access and suspicious behaviors on smartphones. In a controlled 27-participant study, AID achieves an end-to-end F1 score of 0.928 with a 7.0% false positive rate for Top-1 phone-side risk flagging; when preserving top-3 candidate action categories as report context, AID achieves an F1 score of 0.981 and a false positive rate of 1.6%. These findings demonstrate AID's potential as an evidence-support tool that complements current clinic-based interpretation and safety-planning.2025-02-06T00:07:08ZAccepted to ACM IMWUT 2026Weisi YangShinan LiuFeng XiaoNick FeamsterStephen Xiahttp://arxiv.org/abs/2609.09465v1Democracy Needs Reach: Political Equality, Online Speech, and Algorithmic Recommendation2026-09-08T21:31:48ZWithin democracies, the capacity to influence political outcomes through speech depends not only on the right to express oneself, but also on the opportunity to reach relevant audiences. In this paper, I argue that the unequal distribution of algorithmic reach on social media platforms undermines equality of opportunity for political influence (EOPI), which is a central democratic ideal. Drawing on Niko Kolodny's work, I contend that current recommendation algorithms create and perpetuate informal inequalities by concentrating attention among a small minority of already-amplified speakers while systematically marginalizing others. To address this problem, I propose recommendation floors as a mechanism for equalizing political speech. To help users achieve meaningful participation, each verified account would receive guaranteed minimum recommendation for up to a limited number of political posts per week. Although this measure represents one component of the structural reforms needed to move the digital public sphere closer to democratic ideals, it offers a feasible pathway to reducing informal inequalities in political influence online.2026-09-08T21:31:48ZPublished in Ethical Theory and Moral PracticeEtienne Brown10.1007/s10677-026-10535-1http://arxiv.org/abs/2609.03666v2WebXR and Commercial Game Engines for the Metaverse: A Socio-Technical Analysis of Openness, Interoperability, and Sustainability2026-09-08T21:24:35ZThe Metaverse is often framed as a persistent, interoperable, and embodied network of virtual and augmented environments. Yet, most contemporary XR applications are developed through commercial game engines and distributed through proprietary app stores, creating tensions between openness and platform dependency. This paper critically examines open WebXR technologies with conventional commercial game-engine pipelines, with particular attention to XR hardware, software architectures, developer workflows, governance, ethics, interoperability, and sustainability. We argue that WebXR may provide a viable route toward a more accessible, device-independent, and institutionally sustainable Metaverse, especially for education, research, cultural heritage, prototyping, and public-interest applications. At the same time, commercial engines remain advantageous for graphically intensive, low-latency, deeply integrated, and large-scale XR products. The paper concludes that the choice should not be framed as WebXR versus engines, but as a continuum: WebXR is preferable when accessibility, interoperability, low-friction deployment, and long-term maintainability are primary goals, whereas native engines remain preferable when performance, platform-specific hardware access, and production-grade tooling dominate.2026-09-03T11:02:17ZIEEE International Symposium on Emerging Metaverse 2026Luca TurchetMichel Buffahttp://arxiv.org/abs/2608.30107v2AtlasNLP: A Country-Aware Atlas of Dataset Representation in NLP2026-09-08T20:29:40ZUnderstanding which countries are represented in NLP datasets is essential for identifying gaps, targeting data collection, measuring progress, and informing AI policy. However, geographic metadata is very rarely available, and country-level representation is often hidden behind broad language-level claims. We introduce AtlasNLP, a country-aware atlas of over 13,000 NLP dataset records across normalized NLP task categories, tracking both the populations represented and where datasets are produced. AtlasNLP includes AtlasNLP-Gold, a human-curated reference set, and AtlasNLP-Core, an ACL-derived large-scale collection. Using this resource, we show that (1) dataset coverage is highly uneven across countries and tasks; (2) dataset production and representation are geographically asymmetric; and (3) language coverage does not imply geographic representation. These findings reveal blind spots in current dataset documentation practices and motivate more explicit geographic metadata for country-aware NLP evaluation.2026-08-31T00:48:30ZProceedings of the 2026 Conference on Empirical Methods in Natural Language ProcessingJoan NwatuTsedeniya Solomon AmareLongju BaiBontu Fufa BalchaZayd BashirAngana BorahZara BurzoYubin ChoiNaihao DengSamika GuptaMichel FaloughiClaude KwizeraZiqiao MaCynthia Yacel Fuertes PanizoEllie SeehornHui ShenJiayi TangZesen ZhaoBoyuan ZhengRada Mihalceahttp://arxiv.org/abs/2609.09429v1Applying foundation model embeddings towards urban livability evaluation2026-09-08T20:28:47ZWhile accurate measurement of socioeconomic indicators remains challenging in data-scarce regions, which limits policy interventions and resource allocation, high-resolution geospatial data is widely available and can contain information on various livability statistics. We investigate which physical features are encoded within foundation model embeddings, such as AlphaEarth, AnySat, and TerraMind, and provide a systematic framework for identifying the most predictive geospatial indicators. By analyzing how different types of geospatial data influence urban livability predictions, our approach enables researchers to prioritize the most informative features for their specific applications. Additionally, we demonstrate how to leverage foundation model embeddings to enhance prediction performance for these outcomes. This work contributes a principled methodology for extracting actionable information from satellite imagery while accounting for complex spatial dependencies, with applications in predicting urban livability in regions with limited observation data.2026-09-08T20:28:47Z11 pages, 5 figures, 6 tablesAyush KhotWen ZhouShaowen Wang10.1145/3841645.3843316http://arxiv.org/abs/2604.26964v2Dont Just Teach, Explain! A Gamified 20Q Recommender for Cybersecurity Education2026-09-08T19:53:22ZThe escalating complexity of modern cyber threats demands innovative approaches to security education that transcend traditional pedagogical methods. Conventional training paradigms often fail to engage learners meaningfully or develop the intuitive reasoning necessary for effective threat recognition. This paper introduces an interactive educational framework that reimagines cybersecurity awareness through the lens of a structured guessing game. Our approach integrates explainable artificial intelligence (XAI) principles with reinforcement learning to create a dynamic learning environment where users discover cybersecurity concepts through guided inquiry. The proposed system employs a policy-based reinforcement learning agent that assumes the role of a knowledgeable questioner, systematically narrowing down user-described security scenarios until it can both identify the underlying threat and provide transparent reasoning for its conclusion. By framing security education as an interactive dialogue, we transform passive knowledge acquisition into active discovery. We present the complete system architecture, detail the underlying algorithmic foundations, and demonstrate practical application through comprehensive case studies examining diverse attack vectors including the Cyber Kill Chain, phishing campaigns, ransomware outbreaks, and web application vulnerabilities. This work represents a significant departure from static security training methodologies, offering a personalized and game-based approach to cybersecurity education.2026-04-14T14:22:46Z11 pages, 5 figuresMary NusratSarfuddin BhuiyanGahangir Hossainhttp://arxiv.org/abs/2603.22188v2Generalized Sequential Monte Carlo Sampling for Redistricting Simulation2026-09-08T19:26:56ZSimulation methods have become important tools for quantifying partisan and racial bias in redistricting plans. We generalize the Sequential Monte Carlo (SMC) algorithm of McCartan and Imai (2023), one of the commonly used approaches. First, our generalized SMC (gSMC) algorithm can split off regions of arbitrary size, rather than a single district as in the original SMC framework, enabling the sampling of multi-member districts with a varying number of representatives. Second, the gSMC algorithm can operate over various sampling spaces, providing additional computational flexibility. Third, we derive optimal-variance incremental weights and show how to compute them efficiently for each sampling space, leading to more efficient sampling. Finally, we propose a hybrid gSMC-MCMC algorithm by incorporating Markov chain Monte Carlo (MCMC) steps to handle large-scale redistricting applications without changing the target distribution. We demonstrate the effectiveness of the proposed methodology through analyses of the Irish Parliament, which uses multi-member districts of varying sizes, and the Pennsylvania House of Representatives, which has more than 200 single-member districts.2026-03-23T16:48:43ZPhilip O'SullivanKosuke ImaiCory McCartan