https://arxiv.org/api/cpF51MBK/tJgdZplG4WIxqN5F0w 2026-09-11T19:59:49Z 30324 30 15 http://arxiv.org/abs/2609.10738v1 What Makes Creation Human? Authorship, Reasons, and Meaningful Human Control in Generative AI 2026-09-09T18:32:42Z Generative artificial intelligence (GenAI) significantly expands creators' productive capacity, but this does not necessarily entail a corresponding increase in creative agency or authorship. This paper distinguishes creativity at the level of the work from creative agency at the level of the creator, and argues that human authorship cannot be determined solely by manual intervention, degree of automation, the origin of an initial idea, or final selection authority. Rather, authorship depends on whether human judgment and reasons genuinely shape the development of the work. To articulate this requirement, the paper introduces Meaningful Human Control (MHC) into generative creation and identifies a limitation of its classical tracking condition. Creative reasons are not always fully specified prior to interaction with AI; they may emerge, change, or be abandoned as the creative process unfolds. The paper therefore proposes dynamic-reflexive tracking (DRT), which requires that a creator's evolving reasons undergo reflective uptake, exert genuine influence on the subsequent trajectory of creation, and remain capable of rejecting and redirecting the system's default direction. DRT consists of four conditions: diachronic reason formation, reflective uptake, trajectory efficacy, and contestability and redirection, together with a minimal tracing requirement. The paper argues that human authorship under generative AI depends not on how many steps a person personally performs, but on whether that person's reasons continuously, reflectively, and effectively shape what the work becomes. 2026-09-09T18:32:42Z 27 pages Yuxi Cao http://arxiv.org/abs/2609.06769v2 Ordinary, Reasonable Chatbots: Do AI Models Track Human Legal Judgments? 2026-09-09T18:21:26Z As people increasingly rely on artificial intelligence (AI) for guidance in their own lives, scholars, lawyers, and even judges have begun to consider the role of AI in legal decision-making. As "silicon sampling" -- the use of generative AI models in social science research -- is now impacting academia, "silicon jurors" could make an appearance in courtrooms. This study joins an emerging line of research on generative AI models' ability to simulate human legal judgments. In particular, we study how large language model (LLM)-powered chatbots respond to series of questions about legal reasonableness. When the law needs to judge the appropriateness of a behavior, it most often asks whether the behavior was "reasonable." Yet despite the ubiquity of reasonableness judgments, they are the site of constant vexation for lawyers, judges, and lay people. Reasonableness seems inherently vague and unpredictable, since it relies on variable context and implicit conceptual schemas. Moreover, many scholars caution that reasonableness judgments may vary along demographic lines. We compare the answers of human participants to those of twenty-six LLMs across twenty-five different legally relevant reasonableness judgments. Overall, our findings suggest that chatbot responses generally track those of human participants. Nonetheless, we find some suggestive -- and potentially concerning -- results. Compared to humans, LLMs generate more homogeneous responses and occasionally treat a variable standard as an invariant rule. And, compared to humans, LLMs tend to generate answers that are more favorable to the government and to corporations. Finally, our results indicate that LLMs' responses tend to align more closely with those of respondents who are white, male, older, and more educated. More systematic research is needed to confirm or reject these initial findings. 2026-09-06T18:20:15Z Accepted to the Ninth AAAI/ACM Conference on AI, Ethics, and Society Nirav Patel Emily Wenger Christopher Buccafusco http://arxiv.org/abs/2609.10421v1 Emergency Department Revisit Quality Review Screening: Exploring Human Decision-Making and Artificial Intelligence Support 2026-09-09T16:37:35Z Background: Emergency Department (ED) return visits are commonly reviewed for quality assurance, but are often limited (e.g., to revisits within 48-72 hours) to increase actionable finding yield while minimizing chart review burden. Those limitations may lead to missed quality improvement opportunities. Methods: We conducted an exploratory, retrospective study of randomly selected ED visits to a multihospital health system having an ED revisit within 1-14 days to the same health system. Given only each visit's primary diagnosis, raters (2-3 clinicians and GPT-4 large language model [LLM]) assessed characteristics of the diagnosis pairs, including the "target": whether a pair warranted further assessment. Informed by rater response analyses, an algorithm leveraging an LLM-populated knowledge graph ("KGA") was created to automatically screen for potentially concerning pairs, then preliminarily assessed. Results: 99 diagnosis pairs were included. GPT-4 responses poorly correlated to clinician raters, rating nearly all (94%) pairs as warranting follow-up (4.4-13.3 times more than clinicians). However, prompt engineering was minimal. Among clinician raters, revisit medical gravity was consistently significantly associated with the target, while a differential diagnosis/complication composite was significantly associated on unadjusted, but not adjusted (though less powered) analysis. The KGA achieved 83-100% positive predictive value for at least one clinician rater determining further assessment was warranted based on the diagnosis pair. Conclusion: These results can inform next steps for improving screening with LLMs like ChatGPT. Further research is warranted to validate this preliminary work's finding that the KGA may enable enhancing the scope and yield of screening without substantially increasing reviewer workload. 2026-09-09T16:37:35Z 12 pages, 1 figure Jonathan A. Handler Marlene I. Robles-Granda Jacob E. Mefford Jeremy S. McGarvey Gregory S. Podolej Colleen J. Klein Matthew D. Dalstrom William F. Bond http://arxiv.org/abs/2609.10410v1 Can Foundation Models Moderate Online Content? Evaluating Instruction- vs. Example-Driven Policy Operationalization 2026-09-09T16:30:15Z The growing complexity of content moderation policies presents a critical challenge for their consistent operationalization. While foundation models possess the basic capabilities needed to confront this challenge, whether they can reliably moderate online content remains an unanswered question. In this paper, we systematically compare two competing paradigms for Vision-Language Model (VLM) guidance: an instruction-driven approach where models reason from policy precepts, and an example-driven approach where they generalize from prior precedents. We ground this investigation in ModerationBench, a new benchmark of 4,000 manually annotated, in-the-wild posts from the Bluesky platform. Our experiments reveal that foundation models can substantially outperform Bluesky's deployed moderation system, nearly tripling its $F_1$ score (0.60 vs. 0.22) on Random Posts in the benchmark, with both instruction- and example-driven paradigms achieving comparable peak effectiveness. Our findings thus chart a path toward reliable and adaptable policy operationalization at scale. 2026-09-09T16:30:15Z 33 pages, 28 figures, 8 tables Ayan Majumdar Shounak Paul Pushpdeep Singh Ines Abdelaziz Sayeh Jarollahi Seungeon Lee Krishna P. Gummadi Ingmar Weber Abhisek Dash http://arxiv.org/abs/2609.10350v1 Cyber-Financial Contagion: Modeling the Propagation of an AI Vendor Compromise Through the Banking System 2026-09-09T15:46:16Z The banking system now depends on a small set of shared artificial intelligence vendors for fraud screening, credit decisioning, anti-money-laundering triage, customer analytics, and internal decision support. This paper studies how a compromise inside one of those vendors can propagate along a chain of operational, informational, and financial linkages until it triggers losses that look, from the outside, like a classical banking crisis. We build a four-layer heterogeneous network that couples AI vendors, financial institutions, interbank exposures, and customer accounts, and we propose CFC-Prop, a stochastic epidemic-and-clearing model that runs on that network. On a synthetic dataset with 60 vendors, 220 banks, roughly 2,500 vendor-bank service edges, and 1,400 interbank exposures, CFC-Prop reproduces the heavy-tailed loss distributions and the sharp dependence on patch latency that are consistent with prior cyber-financial evidence. We also train an early-warning model, CFC-GNN, that uses vendor-side incident telemetry and graph structure to flag high-cascade-risk vendors before impact. Across four baselines the proposed model reaches AUROC 0.82 and AUPRC 0.60 while keeping calibration errors bounded. We release the full code, synthetic data, and reproducible scripts. The results argue that cyber concentration among AI vendors is a first-order financial-stability problem and give supervisors a concrete quantitative tool for reasoning about it. 2026-09-09T15:46:16Z 11 fig and 10 tables Alex Leytes http://arxiv.org/abs/2609.10280v1 Total Simulated Survey Error: Designing and Diagnosing Survey Responses from Large Language Models 2026-09-09T14:59:07Z Large Language models (LLMs), having been trained on vast amounts of human-generated data, may encode the attitudes and behaviors of these humans. As such, LLMs show promise in mimicking human-like patterns that facilitate their use in simulating people in a wide variety of contexts. One such context is using LLMs as 'silicon samples', i.e., proxies of people in answering survey questions to establish public opinion, design policies, or use as (social) scientific data. However, several critical questions of social biases, generalization, and technical limitations remain, further complicated by a vast design space open to simulation designers. Multiverse analyses might help us make sense of the impact of different design choices, however, we lack a systematic understanding of the design space of LLM-generated surveys as well as how these decisions interplay with inherent LLM limitations. Therefore, how do we systematically identify, trace, and document limitations in LLM-generated survey responses? Building on traditions in the quantitative social sciences, specifically survey methodology and measurement theory, we investigate threats to the validity of LLM-generated survey responses. To do so, we design a framework that enumerates conceptual errors and systematic biases that can occur at different stages of the survey simulation lifecycle. Our framework, called the Total Simulated Survey Error (TS2E) Framework, provides a unified and end-to-end perspective on LLM-generated survey data. The framework, illustrated through a theoretical and empirical case study, enables survey simulation designers to systematically identify and reflect on errors in LLM-generated surveys. 2026-09-09T14:59:07Z Preprint Indira Sen Georg Ahnert Leah von der Heyde Jana Lasser Bernd Weiß Markus Strohmaier http://arxiv.org/abs/2609.10271v1 Senseful Consense: Towards Simplified Cookie Banners using Plain Language 2026-09-09T14:54:50Z While the GDPR and ePrivacy Directive mandate that consent information must be clear and accessible, most modern cookie banners remain obscured by technical jargon, vague phrasing, and frequent content overload or underload. This feasibility study investigates the impact of applying plain language (Einfache Sprache) to cookie banners within the IAB Transparency & Consent Framework (TCF). In our study, we analysed cookie banner texts from 200 websites, using AI-based mapping to categorise extracted content into standardised processing purposes. By substituting complex legal terms with simplified descriptions, we successfully demonstrated that the comprehension barrier can be lowered from a college-graduate level to a 7th-grade level. However, the effectiveness of plain language is inherently constrained by the informativeness of the original content; it cannot compensate for banners that omit legally required details. We conclude that while plain language is a vital tool for digital accessibility, it must be paired with standardised implementation guidelines to ensure that cookie banners are both readable and informative. 2026-09-09T14:54:50Z 14 pages, 3 figures Minela Bećirović Ha Dao Mannat Kaur Martin Johns Alexandra Dirksen http://arxiv.org/abs/2510.05124v3 MADS: Multi-Agent Dialogue Simulation for Diverse Persuasion Data Generation 2026-09-09T13:48:01Z We propose MADS (Multi-Agent Dialogue Simulation), a scalable framework for generating persuasive multi-turn dialogues via agent self-play. MADS employs three coordinated agents: User Agents designed to simulate diverse persona-driven behaviors by leveraging personality signifiers such as Zodiac Signs and MBTI types, a Dialog Agent executing task-oriented persuasion strategies and an Optimization Agent evaluating and refining dialogue outcomes. We further validate its effectiveness through users' Chain-of-Attitude (CoA) modeling and dedicated LLMs' persuasion assessment. This approach enables low-cost generation of training data without human annotation, addressing key industry challenges such as lack of user data, cold-start evaluation difficulties, and prompt inefficiency. Applied to a real-world marketing scenario, MADS significantly improved the persuasion capacity of small LLMs, increasing the organic traffic conversion rate by 22.4% (from 1.83% to 2.24%) , demonstrating clear business value. 2025-09-30T06:55:39Z Accepted to EMNLP 2025 Industry Track (https://aclanthology.org/2025.emnlp-industry.26.pdf) Mingjin Li Yu Liu Huayi Liu Xiang Ye Chao Jiang Hongguang Zhang Yu Ruan http://arxiv.org/abs/2609.08027v2 Delusions and Harms Associated with AI Chatbot Use: Early Evidence from 185 Real-World Reports 2026-09-09T13:06:11Z Importance: Reports have raised concerns that AI chatbots may validate or elaborate delusional beliefs, respond inappropriately to suicidal ideation, and contribute to mental health harms, but real-world data on reported harms remain limited. Objective: To characterize psychopathological features, chatbot behaviors, timing, and outcomes in first- and second-hand accounts of mental health harm linked with AI chatbot use. Design: Cross-sectional secondary analysis of deidentified online survey responses gathered between August 7, 2025, and February 2, 2026. Main Outcomes and Measures: The primary quantitative outcome was the presence of delusional beliefs, coded by paired raters with relevant clinical experience. Additional variables included reason for chatbot use, current episode features, delusional content, chatbot validation of beliefs, harms, social and occupational consequences, healthcare use, and timing. Results: 95 first-hand and 90 second-hand accounts were analyzed. Median age was 35.0 (IQR 27.0 - 45.0). Raters coded descriptions consistent with delusional beliefs in 102 reports (55.1%), with chatbot validation of beliefs in 50/102 (49.0%). Common outcomes included isolation, relationship breakdown, hospital admission, job loss, and financial loss. Four second-hand reports described death by suicide. Conclusions: In this self-selected convenience sample, AI-chatbot-associated harms were frequently described in relation to delusional beliefs, perceived chatbot validation, intensive use, and substantial social, occupational, and clinical consequences. Because reports were retrospective, unverified, and collected from individuals seeking to report harm, our findings should be interpreted as preliminary signal detection rather than as suggesting prevalence or providing evidence of causality. Prospective surveillance and trajectory-based safety evaluations are needed. 2026-09-07T22:25:37Z 18 pages, 3 figures; improved HTML compatibility; article text unchanged Hamilton Morrin Vinitha Soundararajan Thomas Cheliotis-James Boris Warszawski Joshua Fakulujo Zeqi Jia Etienne Brisson Thomas A. Pollak http://arxiv.org/abs/2601.17036v2 LLM-Generated or Human-Written? Comparing Review and Non-Review Papers on ArXiv 2026-09-09T13:04:23Z ArXiv recently prohibited the upload of unpublished review papers to its servers in the Computer Science domain, citing a high prevalence of LLM-generated content in these categories. However, this decision was not accompanied by quantitative evidence. In this work, we investigate this claim by measuring the proportion of LLM-generated content in review vs. non-review research papers in recent years. Using two high-quality detection methods, we find a substantial increase in LLM-generated content across both review and non-review papers, with a higher prevalence in review papers. However, when considering the number of LLM-generated papers published in each category, the estimates of non-review LLM-generated papers are almost six times higher. Furthermore, we find that this policy will affect papers in certain domains far more than others, with the CS subdiscipline Computers & Society potentially facing cuts of 50%. Our analysis provides an evidence-based framework for evaluating such policy decisions, and we release our code to facilitate future investigations at: https://github.com/yanaiela/llm-review-arxiv. 2026-01-19T21:21:42Z Accepted to Findings of EMNLP 2026 Yanai Elazar Maria Antoniak http://arxiv.org/abs/2609.10105v1 Beyond Training: A Feasibility Taxonomy for Inference-Time AI Governance 2026-09-09T12:37:17Z Compute governance today is a governance of training: the thresholds, reporting requirements, and frontier-AI regimes now in force attach to training compute and treat the trained model as the regulatory unit. That picture is incomplete: capability increasingly migrates to the deployment stage through inference-time scaling, agentic scaffolding, and compression onto consumer hardware. This paper asks which mechanisms are available once the regulatory object shifts from the training run to the inference call. We develop a feasibility taxonomy of twenty inference-time mechanisms across monitoring, verification, and enforcement, each rated on a four-point readiness scale against a documented four-vendor evidence base. We then stress the taxonomy against a two-dimensional adversary model (three capability tiers crossed with four adversary roles) and map each mechanism to four governance scenarios (domestic regulation, bilateral or multilateral coordination, industry self-regulation, and compute-marketplace governance). Fifteen of the twenty mechanisms have commercial technical substrates in production today, although governance-grade assurance and adversarial robustness vary substantially. The adversary analysis shows that this readiness holds only against a cooperative deployer and a low-to-medium-capability user: no mechanism rates adequate against a high-capability state-level deployer, and fine-tuning removes the model-internal components of the enforcement cluster, although platform-external controls can persist. A substitution analysis connects the taxonomy to a companion hardware paper as a conditional substitution principle describing when inference-stage and hardware-stage mechanisms provide comparable regulatory coverage under stated conditions. A second-rater reliability check on a random subset of the readiness ratings returned a quadratic-weighted Cohen's kappa of 0.74. 2026-09-09T12:37:17Z Samar Ansari http://arxiv.org/abs/2609.06784v2 GreenPassport: Request-Level Carbon Accounting for Cross-Border AI Inference 2026-09-09T12:05:56Z AI inference often crosses regional boundaries as prompts travel to remote data centers and generated tokens return to users. Regional averages cannot represent the resulting differences in serving hardware, electricity, and network delivery. Request-level accounting needs a common boundary for the service, serving site, route, local comparator, uncertainty, and data provenance. GreenPassport Carbon Accounting (GPCA) associates these inputs with each request. It estimates serving and route carbon, then selects a reporting level from the available documentation. Our public-data implementation covers data-center instances, accelerators, model families, electricity mixes, routes, and cloud-region carbon intensity. Against six accounting baselines and four energy-prediction baselines, GPCA reduced median absolute percentage error by 56.3\% and median absolute error by 15.5\% relative to EcoLogits under the aligned accelerator-energy boundary. It produced zero rule overstatement in the deterministic conformance tests. In the buyer case, the clean-electricity CN-West scenario produced $0.0148$ gCO$_2$e /request, 88\% below the local service at $0.1220$ gCO$_2$e /request. 2026-09-06T19:00:20Z Rui Lu http://arxiv.org/abs/2609.09899v1 Strangers to Themselves: What Language Models Say About Themselves Is Generic 2026-09-09T08:56:02Z Language models can fluently describe how they would behave: whether they would cave to pushback, misuse a tool, or lie under pressure. Is that description actually about the model speaking? We turn self-knowledge into a prediction test. Across nine behavioral evaluations, we measure how a model behaves under different conditions, ask it to predict those rates, and compare its predictions with controls that remove the self from the question. We find that: (i) Direct self-report is weak (r = +0.04), and even showing the model the exact items only raises prediction to +0.24. Crucially, the same item-informed question about "capable AI agents in general" does just as well (+0.28), while other models' answers about themselves predict the target model at least as well as its own. (ii) Frontier scale does not detectably change this pattern: any gains in prediction are not self-specific, and are consistent with a better theory of how AI assistants behave rather than better self-knowledge. (iii) First-person framing does have one robust effect: it shifts reports in the flattering direction, understating harmful behavior relative to the same question about a generic agent. (iv) Finetuning on a model's own behavioral record can teach narrow self-predictions, but it also changes the behavior being predicted and the gains do not transfer broadly. The practical implication is simple: asking a model what it would do mostly reveals a theory of AI assistants in general, plus a favorable bias, rather than privileged knowledge of that model. 2026-09-09T08:56:02Z Phil Blandfort Urja Pawar http://arxiv.org/abs/2609.09887v1 When Does Defendant Statement Matter? A Study of Bias and Persuasion in LLM-Simulated Jurors 2026-09-09T08:43:34Z LLMs have been used to simulate human decision-making in professional settings, yet their behaviors in common-law jury trials remain unexplored. We study when and how a defendant's courtroom statement affects LLM-simulated jurors, focusing on persuasion, ideological bias, and background-based affinity. To support the analysis, we introduce JuryBench, a benchmark containing controversial criminal cases in U.S. criminal law. In each case, a defendant can claim various plausible justifications to support acquittal or reduced liability. We fix the base case and design defendants of different backgrounds, who give courtroom statements with varying emotional appeal or rebuttal. Jurors with diverse ideological profiles across the spectrum are simulated. We examine 20 frontier LLMs, resulting in a total of 432K decisions and rationales, and quantify changes in verdict severity. Our findings show that LLM-jury simulation echoes many human-jury findings. First, emotional persuasion can be detrimental, since jurors may perceive it as evidence of guilt or inconsistency. Next, we show that background fit between jurors and defendants is a stronger and significant factor than other isolated factors, and that jurors are in general harsher toward opposite-background defendants and lenient toward same-background ones. Finally, we find that juror ideology also strongly shapes severity judgments. These findings highlight both the promise and risks of using LLMs to model jury reasoning and call for careful evaluation. The data and code are available at https://github.com/choyingw/JuryBench 2026-09-09T08:43:34Z Accepted to EMNLP 2026 Cho-Ying Wu http://arxiv.org/abs/2609.09856v1 With a Thermomix You Lose the Ability to Cook: A Kitchen Machine Analogy for Applications of Generative AI in Education 2026-09-09T08:10:45Z The rapid adoption of generative AI tools such as ChatGPT has sparked intense debate about their risks and opportunities for education, as well as the ways researchers should investigate them. In this paper, we approach these discussions through an analogy with the Thermomix, a smart kitchen appliance that has similarly provoked both enthusiasm and critique. By mapping Thermomix use cases onto examples of learning with generative AI, and situating them within the ICAP and SAMR frameworks, we show how different modes of tool use can either support or undermine meaningful engagement and learning. The Thermomix metaphor underscores that the central question is not whether learners employ AI, but how such use shapes their learning processes. In doing so, we provide a conceptual lens for researchers and practitioners to critically examine - and more effectively guide - the integration of generative AI into educational practice. 2026-09-09T08:10:45Z First two listed authors have shared first authorship Nikol Rummel Valentina Nachtigall Ernesto Panadero