https://arxiv.org/api/m5QSiqg2EtWcKePAUhLrZH+dSzU2026-09-11T17:47:03Z30324015http://arxiv.org/abs/2609.11814v1Don't Trust the Super-App: A Case Study of Russia's Max2026-09-10T17:08:00ZSuper-apps, an emerging mobile architecture, host third-party mini-apps inside a single app, allowing users to access diverse services. A decade of security research on the super-app ecosystem has all assumed super-apps to be a trusted intermediary. We argue this implicit trust is difficult to justify: China's WeChat is already shown to passively track its user's activity across mini-apps at extraordinary scale; Russia's MAX's parent company is reported to be deeply entangled with the state prosecution of online speech; and Iran's Bale was reported to be functioning in the world's longest internet shutdown due to its state-backed support.
In this paper, we show how malicious super-apps have undeniable capabilities to silently undermine the security and privacy of mini-apps and users without leaving any trace. Using MAX as an example, we show how it can capture mini-app UI, read and write mini-app local storage, inject arbitrary JavaScript into a mini-app's runtime, mediate mini-app network traffic, and control authentication context in ways that can enable silent user impersonation. Sadly, these capabilities manifest themselves in any super-app because of the architectural privileges granted to them by design. We argue that mobile OS and app store interventions are urgently needed to close this architectural blind spot before it is further exploited.2026-09-10T17:08:00ZRicha PriyankaAaron OrtweinJoel ReardonMichael SpecterPiyush Kumar SharmaRoya Ensafihttp://arxiv.org/abs/2508.05830v3"Mirror" Large Language Model Evaluations of Depression are Criterion Contaminated2026-09-10T16:39:44ZLarge Language Model (LLM) studies that use language responses elicited from depression assessments to predict scores on those same assessments often report near-perfect prediction of depression. We refer to these as "Mirror" evaluations and demonstrate an applied case of criterion contamination. N = 110 participants completed both structured diagnostic depression interviews (Mirror condition) and life history interviews ("Non-Mirror" condition). LLMs were prompted to predict depression scores in each condition. As expected, Mirror evaluations were near-perfect. However, Non-Mirror evaluations also displayed prediction sizes considered outstanding in psychology. Further, both Mirror and Non-Mirror predictions correlated with Patient Health Questionnaire-9 scores at similar sizes, suggesting the Mirror condition's advantage collapses when predicting an independent depression measurement. Topic modeling revealed differing depression-related themes across interview types. Mirror evaluations are better considered as reliability evaluations than as validity evaluations. Incorporating Non-Mirror approaches in LLM depression assessment may support more valid and clinically-relevant applications. Keywords: large language models, psychological assessment, psychopathology, depression, reliability, validity, criterion contamination2025-08-07T20:13:00Z48 pages, 10 figuresTong LiRasiq HussainMehak GuptaJoshua R. Oltmannshttp://arxiv.org/abs/2609.11728v1Reproducibility in the Age of Agentic AI: Context Engineering at the Timescale of a Codebase2026-09-10T15:43:34ZReproducible research practices are context engineering for AI coding agents. I argue that agents lower the cost of maintaining tests, commit histories, repository structure, instructions, and decision records while making their benefits immediate. Researchers remain responsible for verifying these artifacts and the scientific judgments they encode.2026-09-10T15:43:34Z10 pagesLorena A. Barbahttp://arxiv.org/abs/2609.11611v1Who Bears the Risk When Generative AI Enters Transport? A Distributional Sociotechnical Audit of Algorithmic Equity, Synthetic-Data Validity, and Public Trust2026-09-10T14:29:09ZGenerative artificial intelligence is entering transportation through traveler-facing advisories, synthetic crash-record generation, and policy decision support. Existing governance frameworks lack transport-specific statistical tools to measure distributional risks across heterogeneous populations. We develop a Distributional Sociotechnical Audit (DSA) that integrates algorithmic equity, synthetic-data validity, and public-attitude heterogeneity into one empirical pipeline. The audit analyzes 5,760 persona-controlled queries to four LLM families across 12 demographic cues and four transport topics, uses two cross-family judges and a Wasserstein-2 Equity Dispersion Index, tests three FARS crash-record generators with conditional projected maximum mean discrepancy (cpMMD), fits a Bayesian ordered-logit model to Pew American Trends Panel Wave 152 (N = 4,538), and combines the signals into a continuous Sociotechnical Risk Index. Congestion-pricing advice has the highest persona-based dispersion (mean EDI = 1.96; highest direct EDI = 2.20). CART synthetic crash records fail all conditional tests (p < 0.001), while the Gaussian copula has borderline conditional stress (p = 0.105) despite passing marginal checks. Attitudes to AI vary across demographic strata. Distributional audits and continuous risk indices with sensitivity reporting offer a more defensible basis for transport GenAI governance than categorical approval tiers, which show a 75% assignment flip rate under weight perturbation.2026-09-10T14:29:09ZAmir RafeSubasish Dashttp://arxiv.org/abs/2603.18677v4Cognitive Amplification vs Cognitive Delegation in Human-AI Systems: A Metric Framework2026-09-10T13:53:34ZArtificial intelligence is increasingly embedded in human decision-making, yet distinguishing systems that genuinely amplify human cognition from those promoting excessive dependence remains underdefined. This paper introduces a framework to distinguish cognitive amplification (improving hybrid performance without degrading human capability) from cognitive delegation (outsourcing reasoning to the AI).
We define four metrics: the Cognitive Amplification Index (CAI*), Dependency Ratio (D), Human Reliance Index (HRI), and Human Cognitive Drift Rate (HCDR). We test this framework in an agent-based NetLogo simulation across three reliance regimes and multiple dependency-atrophy configurations, performing constrained optimizations and parameter sweeps to determine if positive collaborative gain is recoverable. Finally, we introduce an extension with an explicit human-AI interaction term.
Our metrics effectively distinguish degenerate AI-dominated delegation, capability-preserving but weakly competitive interaction, and structurally dependent boundary regimes. Across all baseline configurations, no regime achieves positive collaborative gain relative to the best standalone baseline, even when reducing capability atrophy to zero. This limitation proves structural rather than merely parametric. Positive collaborative gain (CAI* > 0) becomes attainable only after introducing an explicit interaction term allowing retained human capability to contribute directly to the assisted output.
This framework provides a basis for evaluating whether human-AI systems remain cognitively sustainable. The results suggest that preventing capability erosion alone is insufficient for genuine amplification if the architecture remains delegation-oriented. Amplification requires both preserved human capability and a coupling mechanism through which it contributes productively to the hybrid outcome.2026-03-19T09:39:24Z25 pages, 2 figures. Under review at SpringerEduardo Di SantiCarla Floridahttp://arxiv.org/abs/2509.13359v4Generative AI performance in core undergraduate mathematics: a curriculum-level case study2026-09-10T13:49:29ZGenerative artificial intelligence (GenAI) tools such as OpenAI's ChatGPT are transforming the educational landscape, prompting reconsideration of traditional assessment practices. In parallel, universities are exploring alternatives to in-person, closed-book examinations, raising concerns about academic integrity and pedagogical alignment in uninvigilated settings. This study systematically investigates the performance of GenAI on typical mathematics questions from across a first-year mathematics curriculum. Adopting an empirical approach and utilising current examination questions as a proxy for course content, we generate, transcribe, and blind-mark GenAI submissions to eight undergraduate mathematics assessments, spanning the entirety of the first-year curriculum. By combining independent GenAI responses to individual questions, we enable a meaningful evaluation of GenAI performance, both at the level of modules and across the first-year curriculum. We find that GenAI attainment is at the level of a first-class degree, though current performance can vary between modules. Further, we find that GenAI performance is remarkably consistent when viewed across the entire curriculum, significantly more so than that of students in invigilated examinations. Our findings evidence the pressing need for redesigning assessments in mathematics in the era of generative artificial intelligence.2025-09-15T10:34:31ZBenjamin J. WalkerNikoleta KalaydzhievaBeatriz Navarro LamedaRuth A. Reynoldshttp://arxiv.org/abs/2609.11391v1Buyer Artificial Intelligence-Enabled Environmental Governance and Supplier Environmental Controversies: An Organizational Information Processing and Signaling2026-09-10T11:27:27ZEnvironmental controversies in global supply chains pose significant risks for global buyers. This study examines whether overseas suppliers' exposure to buyers' artificial intelligence (AI)-enabled environmental governance reduces supplier environmental controversies. Drawing on organizational information processing theory and signaling theory, we investigate how suppliers' exposure to AI-enabled governance influences their environmental controversies and the institutional contingencies under which this effect varies. Using text analysis to measure buyer AI-enabled environmental governance, we analyze panel data on 2,505 suppliers of U.S.-listed firms across 41 countries from 2020 to 2024 with multidimensional fixed-effects models. We find that suppliers' exposure to buyer AI-enabled environmental governance is negatively associated with supplier environmental controversies in the following year. This negative relationship is stronger in supplier countries with higher AI readiness and regulatory quality. The study contributes to research on AI-enabled sustainability governance and sustainable supply chain risk management.2026-09-10T11:27:27ZYongchao Martin MaXinya Guanhttp://arxiv.org/abs/2606.28335v3LLM-Ideoplasticity: Measuring Ideological Plasticity in the Political Behavior of LLMs as a Context-Conditioned Distribution2026-09-10T11:17:47ZWe argue, with systematic empirical evidence, that a large language model's political ideology is not a fixed point, but a conditional distribution $\mathbb{P}($position$\mid$context$)$ over a real political space. We evaluate nine current LLMs using a unified measurement framework anchored by VAA-CHES projection models, which map responses onto three validated dimensions (lrgen, lrecon, galtan) across six contextual axes. Our findings reveal high sensitivity to context: persuasive framing and under-represented languages displace coordinates by up to 0.57 and 0.52 units, respectively, while chain-of-thought reasoning often amplifies rather than dampens paraphrase instability. Despite this local plasticity, the model cohort occupies a remarkably narrow Overton envelope overall, occupying roughly one-third the spread of major European parties. Supported by a multi-trait multi-method (MTMM) analysis, we conclude that a single point cannot summarize LLM political behavior; it must be characterized as a shape. Our code and data are publicly available at https://github.com/sakhadib/LLM-Ideoplasticity.2026-05-26T17:01:07ZAccepted in Proceedings of the 15th International Joint Conference on Natural Language Processing and the 5th Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (IJCNLP-AACL 2026), 43 pages, 18 figures, 17 tablesAdib SakhawatSyed Rifat RaiyanTahsin IslamTakia FarhinHasan MahmudMd Kamrul Hasanhttp://arxiv.org/abs/2609.11373v1Characterizing Bluesky Content Moderation Service: From Automation of Service to Landscape of Harms2026-09-10T11:08:48ZEmpirical research on content moderation is fundamentally constrained by the opaque deployment of moderation systems on major social media platforms. To this end, the recent emergence of decentralized platforms with transparent, public moderation logs presents an unprecedented opportunity for independent audits. In this work, we leverage this architectural transparency to conduct the first large-scale audit of the default moderation system on Bluesky, the Bluesky Moderation Service (BMS). Analyzing its 10.6M moderation labels from 2025, we investigate three foundational aspects: (i) its mechanism (the degree of automation versus human oversight), (ii) its efficacy (accuracy in detecting harms), and (iii) its purpose (the landscape of harms it identifies). Our findings reveal a human-AI collaborative system where labels for sexual and graphic content are applied automatically in seconds, while nuanced and high stakes labels require more human oversight, taking hours or days. Through a manual annotation study, we find the BMS operates with high precision (0.837), but struggles with low recall (0.222), with our annotators identifying 4.5$\times$ more harmful content than the moderation system in a random sample. Finally, unsupervised clustering of the most frequently applied labeled posts uncovers detected harms ranging from hostility in discourse toward protected groups to the spread of sexually explicit and other graphic content. Our work offers a look into the operational realities of a deployed moderation system, providing a concrete data-driven foundation for designing more effective and transparent moderation systems.2026-09-10T11:08:48Z18 pages, 7 figures, 12 tables. Accepted for publication at ICWSM 2027Pushpdeep SinghSayeh JarollahiAyan MajumdarVabuk PahariAbhijnan ChakrabortyKrishna P. GummadiIngmar WeberAbhisek Dashhttp://arxiv.org/abs/2607.16221v2Students' Perceptions of Peer Grading2026-09-10T10:21:11ZPeer grading is widely used in education, yet it elicits mixed reactions from educators and students. Although many studies have examined students' views of peer grading, their findings are scattered, and no clear overall picture has emerged. To address this gap, we conducted a mixed-source thematic analysis of literature and student discussions on Reddit. To scale our analysis of the Reddit data, we fine-tuned a Gemini 2.5 text-classification model to classify an initial dataset of 659 posts and 6,607 comments by relevance. Manual review of the items classified as relevant by the model yielded a final dataset of 114 posts and 300 comments. Drawing on evidence from 107 papers and the Reddit dataset, we found that students perceive peer grading as both beneficial and problematic. Positive perceptions included learning and understanding benefits, skill development, engagement, and collaboration, while negative perceptions centered on unreliable grading, unfairness, weak feedback quality, emotional stress, and workload. Reddit discussions also suggested an emerging concern that remains underexplored in the literature: AI use in peer grading may weaken students' trust in the accuracy and authenticity of the process. We further identified eight mitigation strategies and mapped them to the negative perceptions they help address. Among these, instructor oversight and training played the most central role.2026-06-13T02:46:09ZExpanded author version of a paper accepted for presentation at the European Conference on Technology Enhanced Learning (ECTEL 2026)Uchswas PaulJash ShahKeira McArthurAref BabaeiNiranjan RajendranParvez RashidEdward Gehringerhttp://arxiv.org/abs/2609.11320v1A global mobile network coverage raster product at 1km resolution, 1999--20302026-09-10T09:48:42ZWhere a mobile signal is available shapes who can work, learn, bank, seek health care and respond to crises in the digital age, yet no globally consistent, sub-national record of mobile network coverage exists. We present such a record: annual 1km maps of the probability of 2G, 3G and 4G coverage for 214 countries and territories for the years 1999 to 2030. The maps are produced by three independent models: a calibrated machine-learning model, a techno-economic simulator of network build-out, and a spatial deep-learning model. The three estimates are then combined, per country and technology and in proportion to their measured accuracy, into a single best estimate with per-pixel 90% uncertainty bands; all four layers are released as part of the dataset. Because mobile roll-out closely follows a country's socio-economic conditions (population distribution, electrification, physical infrastructure), the models are grounded in existing geospatial data and tuned on 2,409 quality-screened operator-reported coverage maps, which are available up to 2020. For 2021--2024 the maps are predicted from recent geospatial data alone; for 2025--2030 they are extrapolated from demographic and infrastructure projections. On countries held out during training, the machine-learning model attains AUC 0.89--0.92. Baseline comparisons and the combined product's external validation are reported in Technical Validation. The dataset supports mapping the global digital divide, linking connectivity to household-survey outcomes, and humanitarian and infrastructure planning.2026-09-10T09:48:42ZDataset linked to that paper: https://zenodo.org/records/21594337Till KoebeTheophilus AidooAli El ChamiAli KansoAkansh MauryaPurushottam SharmaIngmar WeberRidhi Kashyaphttp://arxiv.org/abs/2609.11261v1INDRA: A New AI Tool for Exploring Tobacco, Fossil Fuel, and Chemical Industry Archives2026-09-10T08:55:00ZFive decades of litigation have disgorged hundreds of millions of pages of formerly secret business records from the tobacco industry, along with documents from the makers of drugs, chemicals, food, firearms, and fossil fuels. Yet these archives have been effectively inaccessible to general-purpose large language models (LLMs) because they have never been compiled into an LLM-readable corpus. Chatbots may be familiar with some of the materials contained in such archives but, with no direct access to the documents, they are vulnerable to hallucination and other defects. Here we introduce INDRA, a research platform designed to remedy such failures by embedding the conventions of archival historiography into a system-level protocol governing every output. The platform federates UCSF's Industry Documents Library, Columbia and CUNY's ToxicDocs, Stanford's SRITA, and other heretofore siloed collections, and provides three interlinked safeguards: (1) a closed evidentiary sandbox confines the model to a user-selected corpus, blocking retrieval from external sources that could introduce bias; (2) real-time provenance tagging marks the boundary between archival evidence and parametric inference; and (3) a system-level protocol enforced by deterministic scripts guides the structure of every output. Together these safeguards prevent the model from conflating "the documents say X" with "I think X" or "I learned X from prior training." The result is an LLM-powered research partner enabling massive multi-archival investigations, a tool whose outputs are designed to be checked rather than trusted, and whose architecture makes the conditions of knowledge production visible and auditable. Three case studies demonstrate the method's analytical value and limitations, including what we call the Heraclitus effect, the steppingstone dilemma, and the gullibility (or mafia) problem.2026-09-10T08:55:00Z35 pages, 6 figures, Appendices available at https://indra.stanford.edu/methods/appendicesDaniel AkselradRobert N. Proctorhttp://arxiv.org/abs/2609.11258v1SoulAuth: An Actor-native Identity Architecture and Rust Reference Implementation for Humans and Long-lived AI Actors2026-09-10T08:54:02ZAs AI systems move from transient model invocations toward long-lived actors that persist across credentials, clients, sessions, and runtime instances, identity infrastructure must answer a basic question: where should the canonical continuity boundary be placed? This paper introduces Actor-native Identity and presents SoulAuth, an open-source Rust reference implementation for Humans and long-lived AIActors. We argue that any subject that must persist under its own identity and remain independently attributable should have an ActorIdentity that is not replaced by an Account, Credential, Client, AuthSession, IdentityBinding, or runtime instance. SoulAuth therefore treats Humans and long-lived AIActors as first-class identity subjects while keeping authentication distinct from downstream authority. Methodologically, we use a Philosophical Engineering approach that translates conceptual analysis of subjecthood into identity objects, invariants, lifecycle semantics, system responsibilities, implementation boundaries, and inspectable conformance evidence. Evaluation against the fixed SoulAuth v0.1.0 artifact shows that the implementation realizes core boundaries including Human/AIActor first-class identity status, Client/Actor separation, and Authentication/Authority separation, while gaps remain in unified Credential modeling and historical attribution anchored to ActorIdentity. We therefore report partial, not full, architecture conformance.2026-09-10T08:54:02Z34 pages, 8 figures. Preprint v1.0. Open-source Rust reference implementation and fixed v0.1.0 software artifact: https://github.com/TrantorLabs/SoulAuthKun YuanHarold WangEcho LiEgusi GuiKiki HuLucas LuoMagnus Huhttp://arxiv.org/abs/2609.11198v1(Whose defaults?) Is artificial intelligence reorienting archaeological methods?2026-09-10T08:08:15ZGenerative AI and the practice of "vibe coding" are changing how archaeologists carry out computational research, but their effects on the discipline's range of methods is still understudied. In this paper, we evaluate whether large language models (LLMs) are narrowing the variety of methods archaeologists use. We first analysed approximately 119,000 archaeology abstracts from Scopus, covering publications from 2010 to 2025. Using a locally run LLM, we identified the computational methods reported in each abstract and organised them into 25 broad categories (L2) and 241 finer clusters (L3). A Bayesian Dirichlet-multinomial model of method composition within sub-disciplines found a small but credible shift in method use after 2023. However, this shift was smaller than the variation already present across the full study period. No individual technique showed a significant change, and overall methodological diversity increased rather than declined. We then ran a controlled experiment to see whether LLMs recommend a narrower set of methods than archaeologists have used in practice. Two different open-weight models were asked to suggest methods for 28 archaeological research problems, with prompts providing three levels of methodological guidance: novice, intermediate, and expert. Recommendation diversity was much lower than in the published literature, particularly without methodological guidance. The models also tended to favour methods that were widely used before 2023, and their recommendations more closely resembled the post-2023 literature. Taken together, these results are consistent with LLMs pushing methodological choice towards convergence, although our study cannot establish a causal effect. They raise a broader question: how can archaeology retain methodological diversity as LLMs become more involved in research?2026-09-10T08:08:15ZLorenzo CardarelliRoberto Ragnohttp://arxiv.org/abs/2605.23234v4Assessing Predictive Models for Fairness Based on Activity-Space Patterns2026-09-10T07:16:55ZAssessing the spatial fairness of predictive models involves establishing whether they are statistically penalizing (favoring) individuals associated with certain geographical locations. Literature on this topic makes the fundamental assumption that each individual is assigned to a single geographical location (e.g., place of residence). However, fairness with respect to the set of regions where one regularly spends time, i.e., the individual's activity space, also matters when fairness is considered. Consequently, we argue that it is necessary to generalize the notion of spatial fairness to also account for such activity-space patterns, leading to the novel problem of assessing predictive models for fairness relative to the movements of individuals. To deal with this problem, we propose an approach that first associates individuals with geographic regions relevant to their activity spaces, considering multiple spatial partitions with different resolutions and alignments, and then employs a suitable spatial scan statistic to assess whether a predictive model is fair based on activity-space patterns. In the experimental evaluation, we study the performance of our approach over thousands of synthetic unfair datasets, showing that it is effective at detecting this new type of unfairness and at retrieving the set of objects treated unfairly, while localization performance exhibits a consistent multi-resolution trade-off.2026-05-22T04:53:34Z35 pages, 10 figures, 7 tablesFrancesco LettichMario A. NascimentoChiara PuglieseChiara Renso