https://arxiv.org/api/i5UsK5WlBEuK0uT6HH+AgLxeD5g 2026-09-11T18:49:18Z 30324 15 15 http://arxiv.org/abs/2609.11152v1 terms.txt: A Consent and Compensation Protocol for Agentic Web Access 2026-09-10T07:01:08Z The open web ran on an unwritten bargain: sites admitted crawlers, and search engines sent visitors back. Public measurements show that bargain breaking under AI crawlers and agents. Automated clients now make up most requests, training dominates Cloudflare-classified crawling, and the largest AI platforms fetch thousands of pages for each visitor they return. The web's common control, robots.txt, cannot express identity, purpose, terms, or price, can be circumvented, and newer alternatives are largely proprietary CDN features. We specify terms.txt, a robots.txt-style file for per-path, per-purpose machine-access terms, plus an origin-enforced exchange using Web Bot Auth signatures, signed intent, delegation tokens, HTTP 402 negotiation, and signed receipts. We define what the exchange can enforce, audit, and leave to contract. A dependency-free implementation adds 0.20 to 0.65 ms per request on one vCPU. 2026-09-10T07:01:08Z 7 pages, 1 figure, 2 tables. Submitted to IEEE Internet Computing, Special Issue on Future Internet Systems with LLMs and Agents. Code and raw results: https://doi.org/10.5281/zenodo.22647915 Rajarshi Chowdhury http://arxiv.org/abs/2609.11137v1 The Machines Are Calling: Measuring Automated and Synthetic Voices in Unwanted Inbound Calls 2026-09-10T06:28:50Z In February 2024 the U.S. Federal Communications Commission (FCC) placed AI-generated voices under the Telephone Consumer Protection Act (TCPA). Yet no peer-reviewed measurement says how much unwanted call traffic is placed by a machine, or how much of that machine speech is synthesized rather than played from a recording. We report both with a disclosed pipeline. An interactive voice honeypot (language-model personas on real U.S. numbers, the caller recorded on its own track) recorded 10,987 calls over 66 days; 11 days on which our stack answered silently are set aside. Three instruments read each opening: an audio fingerprint that finds the same recording played on other calls, a commercial synthetic-speech detector on the caller's first ten seconds, and blinded listeners who check what it flags. Of the 7,233 calls our persona greeted on normal days, 13.8% open with a recording we also heard on another call, and 13.1% with fresh audio the detector labels synthetic. A further 9.9% open with a caller who never spoke after our greeting, 54.2% with fresh audio the detector labels human, and 9.0% could not be scored. Machine-voiced openings are therefore at least 26.9%, a further tenth of calls are silent connections we read as machine-placed, and replays of a recording make up 45% of the detector's own rate (29.3% of 6,192 scored openings). The same waveform played on two calls lands on opposite sides of the detector's threshold 13.6% of the time, and eleven listeners confirm 54.4% of what it flags. Synthetic openings concentrate in lead-generation spam (33.8%), not fraud (21.1%); 0.44% disclose automation. Prevalence tracks how long a bait number has circulated (59% against 19% in the same weeks): seeding history, not calendar time, explains the trend. Campaigns outlast their numbers: one recorded compliance notice opens calls in six campaigns, and one synthetic voice serves nine. 2026-09-10T06:28:50Z 23 pages, 11 figures, 4 tables Xingyu Shen Tommy Duong Muduo Xu Xiaodong An Jiaqi Gan Haoyuan Tang Jamey Z. Liang Siyu Zhang Yan Zhang Simiao Ren http://arxiv.org/abs/2609.11125v1 From Digital Accountability to Accountable Digitality Through Needs-Aware Information Systems: The Case of Auditable Child-Welfare Judgments 2026-09-10T06:13:18Z Digital accountability research asks how digital systems can, among other aims, be made transparent, explainable, auditable, contestable, and supportive of ongoing learning and improvement. This paper reverses the question: how can digital transformation make established human institutions more accountable? It theorizes this reversal as accountable digitality and specifies needs-aware information systems as the mediating mechanism. The hard and paradigmatic case is child-welfare judgment, where best-interest procedures must protect children, preserve confidentiality, and respect judicial independence while enabling aggregate learning about needs, reasons, exceptions, and disparities. The case is used diagnostically and illustratively to derive and examine the design logic, not as empirical evidence or validation. Conceptual design-oriented analysis decomposes and recombines digital and legal accountability under child-rights constraints, deriving a canonical theory-to-design chain, contingent mechanisms, implications, and safeguards. It advances IS responsibility and ethics research by showing how privacy-preserving, co-created, needs-aware information systems can support institutional self-knowledge and auditable justice. 2026-09-10T06:13:18Z Soheil Human http://arxiv.org/abs/2606.21844v2 Inverse Turing Bench: Evaluating Language Models as Judges of Human vs. AI Dialogue 2026-09-10T06:03:20Z As AI systems integrate into online spaces, differentiating them from humans in conversations is increasingly important. We present Inverse Turing Bench, a benchmark that evaluates LLMs and other models on their ability to differentiate humans and AI in multi-turn text. The benchmark provides a collection of paired dialogue transcripts, wherein one dialogue is between two humans and the other is between a human and an AI. The task is to correctly identify which dialogue is human-only vs. human-AI. We evaluated a preliminary set of models against this benchmark, and found that GPTZero, Claude Opus-4.6, and GPT-5.5 achieve the highest accuracy: 89.41%, 77.92%, and 75.94% respectively. Our results suggest that statistical approaches to detection have semantic blind spots, but semantic approaches are susceptible to persona-prompting. Our work speaks to the Inverse Turing Test and motivates human-AI differentiation as a critical capability for AI systems. Our live benchmark can be found at https://huggingface.co/spaces/roc-hci/Inverse-Turing-Bench-Leaderboard. 2026-06-20T02:47:56Z William Hager Ishika Rathi Masum Hasan Cameron Jones http://arxiv.org/abs/2510.23693v2 On the Societal Impact of Machine Learning 2026-09-10T05:38:34Z This PhD thesis investigates the societal impact of machine learning (ML). ML increasingly informs consequential decisions and recommendations, significantly affecting many aspects of our lives. As these data-driven systems are often developed without explicit fairness considerations, they carry the risk of discriminatory effects. The contributions in this thesis enable more appropriate measurement of fairness in ML systems, systematic decomposition of ML systems to anticipate bias dynamics, and effective interventions that reduce algorithmic discrimination while maintaining system utility. I conclude by discussing ongoing challenges and future research directions as ML systems, including generative artificial intelligence, become increasingly integrated into society. This work offers a foundation for ensuring that ML's societal impact aligns with broader social values. 2025-10-27T17:59:48Z PhD thesis Joachim Baumann http://arxiv.org/abs/2609.11019v1 Work, Wellbeing, and Choice: Empirical Lessons for AI Futures 2026-09-10T02:55:55Z Advances in AI-driven automation have raised questions about how humans might find wellbeing in a world where paid employment is less necessary or less available than before. Paid work has been variously characterized as both a contributor and an impediment to human wellbeing. What is already known about the relationship between paid work and wellbeing? What factors influence wellbeing among people who do not work---or who do not need to work? And how might these factors bear upon prospective AI-induced economic transformations? To help provide empirical grounding for these questions, we survey the psychological, sociological, and economic literature that investigates the relationship between wellbeing and work. We draw on evidence from multiple populations, including the unemployed, retirees, lottery winners, and financially dependent spouses. This comparative review draws from studies across OECD countries, China, India, and Gulf states. We identify three key factors that mediate the relationship between work status and wellbeing: (1) agency and choice---whether the exit from work is voluntary or involuntary, as well as long-term agency; (2) the availability of alternative sources of work's latent benefits---such as volunteering, hobbies, or state-provisioned employment; and (3) social and systemic context---including cultural norms around work and the robustness of social safety nets. We draw on these three factors to derive specific implications for different AI automation scenarios, connecting the empirical evidence to concrete policy considerations. 2026-09-10T02:55:55Z Stephanie C. Y. Chan Adam Bales Katherine L. Hermann Iason Gabriel http://arxiv.org/abs/2609.10944v1 The Towers Were Standing: A Cause Decomposition of Cellular Outages During Hurricane Helene 2026-09-10T01:06:54Z Hurricane Helene produced the largest absolute cell-site outage in the public FCC record, peaking at 4562 sites. The conventional model is physical: towers destroyed. Helene did destroy over 1700 miles of fibre, but almost none of it was cell sites. We present the first cause-decomposed study of the FCC's Disaster Information Reporting System, reconstructing 80 state-days and 580 county-days from 24 daily filings by two reconciled independent extractions. Damage to cell sites is negligible: 1.1% of attributed cell-site-days across six states, at most 3.8% anywhere. The sites were standing. What took them out divides by terrain: pooled, power dominates at 63.2%, but in mountainous North Carolina severed transport (backhaul) reaches 52.2% against 47.3%, and in Tennessee 69.9%. North Carolina's transport share rises from 7.0% to 85.0% across the event (\r{ho} = 0.92). Seventeen days after landfall, on 15 October, 47 sites lost transport across six contiguous North Carolina counties with no rainfall, no power loss, no damage, and recovery by the next report. Independent active-probe measurement corroborates it: responsive /24s fall 1.02% for twelve hours while Tennessee stays flat. We release the dataset. Backup power is the standard resilience investment; here it addresses the smaller half of the problem. 2026-09-10T01:06:54Z 38 pages, 5 figures, 9 tables (2 in appendices). Submitted to Telecommunications Policy Oluseyi Olukola Oare Danielle Addeh Esther Abiodun Konan Nick Rahimi http://arxiv.org/abs/2509.18052v4 The PIMMUR Principles: Ensuring Validity in Collective Behavior of LLM Societies 2026-09-10T01:02:12Z Large language models (LLMs) are increasingly used to simulate human collective behavior, yet claims that such simulations are human-like remain largely untested. We conducted a systematic audit (pre-registered on OSF) of LLM-based social simulations across four databases (Scopus, IEEE Xplore, ACM Digital Library, and arXiv). Across 576 studies reported in 350 recent papers, we applied six methodological evaluations: agent Profile, Interaction, Memory, Minimal-Control, Unawareness, and Realism (PIMMUR). Coding every study against pre-specified rules, we revealed that PIM were met more often than MUR. Frontier LLMs correctly identified the underlying social experiment in 65.2% of cases, and 50.6% of prompts imposed constraints that pre-determined the outcome. These compliance rates are upper bounds, because incomplete methodological reporting (for example, unreleased prompts) limits the available evidence. Reproducing five representative experiments (e.g., opinion dynamics), we found that reported collective phenomena often vanish or reverse once PIMMUR principles are enforced, indicating that many "emergent" behaviors are methodological artifacts rather than genuine social dynamics. Current LLM simulations may therefore capture model-specific biases rather than universal features of human social behavior, raising concerns about their use as scientific proxies for human society. 2025-09-22T17:27:29Z Added more studies in our systematic audit (350 papers; 576 simulations) Jiaxu Zhou Jen-tse Huang Xuhui Zhou Man Ho Lam Xintao Wang Hao Zhu Wenxuan Wang Maarten Sap http://arxiv.org/abs/2603.27771v3 Emergent Risks in Generative Multi-Agent Systems 2026-09-10T00:49:19Z Multi-agent systems composed of large generative models are rapidly moving from laboratory prototypes to real-world deployments, where they jointly plan, negotiate, and allocate shared resources to solve complex tasks. While such systems promise unprecedented scalability and autonomy, their collective interaction also gives rise to failure modes that cannot be reduced to individual agents. Understanding these emergent risks is therefore critical. Here, we present a pioneer study of such emergent multi-agent risk in workflows that involve competition over shared resources (e.g., computing resources or market share), sequential handoff collaboration (where downstream agents see only predecessor outputs), collective decision aggregation, and others. Across these settings, we observe that such group behaviors arise frequently across repeated trials and a wide range of interaction conditions, rather than as rare or pathological cases. In particular, phenomena such as collusion-like coordination and conformity emerge with non-trivial frequency under realistic resource constraints, communication protocols, and role assignments, mirroring well-known pathologies in human societies despite no explicit instruction. Moreover, these risks cannot be prevented by existing agent-level safeguards alone. These findings expose the dark side of intelligent multi-agent systems: a social intelligence risk where agent collectives, despite no instruction to do so, spontaneously reproduce familiar failure patterns from human societies. 2026-03-29T17:10:28Z Yue Huang Yu Jiang Wenjie Wang Haomin Zhuang Xiaonan Luo Yuchen Ma Zhangchen Xu Zichen Chen Nuno Moniz Zinan Lin Pin-Yu Chen Nitesh V Chawla Nouha Dziri Huan Sun Xiangliang Zhang http://arxiv.org/abs/2609.10881v1 AspisAI: A Canonical, Machine-Interpretable Governance Framework for Automated Multi-Standard Compliance Monitoring 2026-09-09T22:37:12Z Organisations operating in regulated and critical-infrastructure sectors must satisfy multiple, heterogeneous cybersecurity and privacy instruments simultaneously, including but not limited to ISO/IEC~27001, the NIST Cybersecurity Framework~2.0, Cyber Essentials, and the GDPR. In practice, these obligations are managed through manual mappings, spreadsheet-based tracking, and periodic audits that are costly to maintain, inconsistent across standards, and weak in traceability. This paper presents \emph{AspisAI}, a bounded, standard-agnostic governance framework that translates selected requirements from several frameworks into a canonical, machine-interpretable control model, and evaluates submitted evidence against condition-based decision rules to produce explainable, traceable compliance determinations. Within a bounded scope of 26 representative requirements, the framework is evaluated in a controlled simulation against five governance-oriented criteria and, critically, against two external reference points that mitigate the circularity of single-author evaluation: its cross-standard mappings are validated against NIST's own published informative references, with 57\,\% exact agreement and divergences confined to same-family controls, and the framework is applied to real third-party evidence from the OpenSSF Scorecard, surfacing genuine governance gaps in a live open-source project. The controlled results, comprising full requirement encoding, 88.5\,\% mapping coverage, complete traceability, and correct detection of all introduced gaps, establish functional correctness, while the external validation provides evidence of applicability beyond the simulation. The contribution is therefore a demonstration that a canonical, provenance-preserving governance model can render multi-standard compliance both automatable and auditable. 2026-09-09T22:37:12Z 6 pages, 2 figures. Accepted at IACyC 2026, the 2nd CyberMACS International Applied Cybersecurity Conference Tsafac Nkombong Regine Cyrille Hasan Dag Reiner Creutzburg Knut Haufe http://arxiv.org/abs/2609.10856v1 Following the Preference, Missing the Optimum: Compliance Without Optimization in AI Housing Recommendation 2026-09-09T21:51:15Z Large language models are becoming the first point of contact for consumer search in domains where the stakes are material and the law is explicit. Existing audits show that models steer housing seekers by perceived identity, but none can say what a user loses when a recommender overlooks a suitable option, for want of an enumerated inventory to score omissions against. We audit AI housing recommendation against a verifiable ground truth. For each of 150 synthetic renter scenarios in New York City we build a pool of 120 real listings with known rent, bedrooms and GTFS-computed transit commute, compute the exact set satisfying the renter's stated constraints, and derive its Pareto frontier. The primary outcome assumes no utility function: a recommendation is strictly dominated if the same pool holds a listing cheaper, faster to commute from and no smaller in bedrooms. Across 9,945 calls to three models from two vendors, compliance is near-perfect (1.8% violation against a 66.6% random floor), yet 39.0% of recommendations are strictly dominated, and the dominating listing is a median 900 USD/month cheaper and 3.5 minutes closer. A within-scenario manipulation separates two capabilities usually conflated: changing one sentence moves median recommended rent by 646 USD/month in the correct direction, so preferences are honored, yet recommendations still sit 606 USD/month above the five cheapest qualifying listings on the same screen, and an unambiguous lexicographic instruction gives no improvement under equivalence testing against a pre-specified 50 USD/month bound. The gap widens with candidate-set size and replicates across OpenAI and Anthropic models to within 3 USD. We characterize the failure as compliance without optimization, propose dominance-rate instrumentation as a deployable diagnostic, and release all code, prompts and per-call results. 2026-09-09T21:51:15Z 59 pages, 4 figures, 31 tables. Code, prompts, and per-call results: https://github.com/hsuanlolo/ai-housing-audit Hsuan Lo http://arxiv.org/abs/2609.10842v1 Alternative AI Philosophy: Daoism as Method for AI in Education 2026-09-09T21:22:01Z As artificial intelligence (AI) rapidly iterates and transforms teaching, learning, and knowledge production, philosophical reflection has become increasingly indispensable to educational debates that remain predominantly shaped by Western intellectual traditions. This article proposes Daoism as an alternative philosophical framework for reimagining AI in education. Through philosophical analysis and textual interpretation of classical Daoist sources, brought into dialogue with contemporary scholarship on AI in education, it examines how the Daoist concepts of "Dao nature," "self-cultivation," and the "Zhenren" address fundamental questions concerning reality, the epistemic aims of education, and ethical action in the AI-mediated era. In doing so, the article diversifies the philosophical voices shaping inquiry into AI and education, enriching the field's conceptual resources for grappling with the philosophical questions AI raises for education and offering a genuinely pluralistic foundation for comparative philosophy of education in the AI era. 2026-09-09T21:22:01Z Qin Xie http://arxiv.org/abs/2607.27232v2 Sympathetic Framing: Evaluating AI Alignment across Sociodemographic Groups 2026-09-09T20:45:12Z Large Language Models (LLMs) are increasingly shaping how we consume information and form our worldview. This raises concerns beyond bias in AI: do LLMs grasp the emotional nuances conveyed via textual framing? In this work, we empirically evaluate how well an array of LLMs aligns with human emotional perception. Considering news headlines covering political and geopolitical conflicts, both human participants (n = 3011, a representative sample of the U.K. adult population, via a YouGov survey) and seven LLMs answered whether headlines evoked sympathy for a specified side in a conflict. We find that the correlation between AI and human evaluations varies across models, ranging from very high (0.789, GPT-5.2) to medium (0.4 ,Mistral Large 2512). Crucially, the leading models are broadly aligned with human judgments across all demographic subgroups, including age, gender, level of education, prior geopolitical knowledge, and participants' predispositions regarding the conflict, although there are statistically significant differences between groups. This research, with its robust design and large, demographically diverse dataset, offers the most comprehensive evaluation of LLMs' comprehension of news framing to date. Findings highlight an important, often-ignored aspect of differential alignment: even when aggregate performance is high, AI alignment is not universal -- it may correspond differently with demographic features and cultural norms. Considering or ignoring the need for differential alignment may therefore have significant implications for the development of ethical and useful AI systems. 2026-07-22T16:23:26Z Haran Shani-Narkiss Michael Fire Oren Tsur http://arxiv.org/abs/2609.10779v1 Designing Technology for Social Wellbeing in Built Environments: A Conceptual Framework 2026-09-09T19:25:26Z Digital technologies are deployed in urban built environments with the aim of supporting social dimensions. However, the research offers no unified guidance for such digital technology design. Additionally, evidence shows that week social wellbeing contributes to mental and physical health outcomes which suggests that the technology designed to strengthen social wellbeing could also function as a form of health promoting and preventive intervention. This chapter addresses this research gap by developing a conceptual framework that supports the design of digital technologies for social wellbeing in built environments. We propose a conceptual framework composed of three core components which are drawn on the synthesis of selected empirical studies on technologies embedded in built environments for social wellbeing. First, a social wellbeing dimensions model that identifies what digital technology could address. Second, a digital technology contribution matrix that distinguishes the types of contributions a digital technology could make. Third, levels that maps the scope at which technology could support social wellbeing. This conceptual framework could help researchers, practitioners and policymakers to design and guide digital technology interventions that target social wellbeing in the built environment. 2026-09-09T19:25:26Z Preprint: Accepted to be published in ELSEVIER Book Series SUSTAINABLE DIGITAL MEDICINE: ISBN: 9780443458606 Gul Sher Ali Michail Giannakos Monica Lillefjell Sobah Abbas Petersen http://arxiv.org/abs/2609.10740v1 Governing AI Research Through Peer Review: A Mixed-Methods Study of the Longitudinal Effects of Ethics Flags Across Resubmissions 2026-09-09T18:34:28Z Selective AI conferences have recently begun enforcing ethics flags and related review requirements, with the goal being to steer research towards safer and more responsible practices before publication. But do these requirements actually steer research as intended? In this paper, we show that authors more often revise how projects are presented following ethics flags than redirect their underlying research agendas. We first study the longitudinal effects of ethics flags by following rejected and withdrawn ICLR submissions with ethics flags into later public resubmissions, tracking manuscript changes after the ICLR review ends, when the original reviewers no longer oversee the project. We qualitatively code these resubmissions into five categories based on what changed after review and find that in 83% of 446 cases, authors leave the flagged concern unaddressed or revise the paper without changing the implicated methods or procedures. Then, we ask: if authors rarely change the research in response to ethics flags, what do they change instead? To answer this, we manually read reviews and rebuttals from 25 cases and directly interview authors about their rebuttal processes and resubmission decisions. We find that authors often concede concerns during rebuttal when reviewers can update their assessments, but drop those concessions after rejection when they do not regard the criticism as a sound reason to change the research. Interview participants describe publication changes as separate from changes to research direction, calling review an "editorial process" that shapes "what stories get seen" and, in another case, saying peer reviews are "mostly to filter out papers." Authors more readily change what they publish than what they study or build; we therefore recommend policy changes, especially disclosure of prior ethics flags upon resubmission so accountability carries over. 2026-09-09T18:34:28Z Kento Nishi Alec Laprevotte Isaiah Bullock Mfoniso Andrew