https://arxiv.org/api/HSyWNfE1vbr0ufa1cS2Db/0DYww2026-09-11T20:09:40Z326953015http://arxiv.org/abs/2504.13700v2Exploring Multimodal Prompt for Visualization Authoring with Large Language Models2026-09-09T21:12:32ZRecent advances in large language models (LLMs) have shown great potential in automating the process of visualization authoring through simple natural language utterances. However, instructing LLMs using natural language is limited in precision and expressiveness for conveying visualization intent, leading to misinterpretation and time-consuming iterations. To address these limitations, we conduct an empirical study to understand how LLMs interpret ambiguous or incomplete text prompts in the context of visualization authoring, and the conditions making LLMs misinterpret user intent. Informed by the findings, we introduce visual prompts as a complementary input modality to text prompts, which help clarify user intent and improve LLMs' interpretation abilities. To explore the potential of multimodal prompting in visualization authoring, we design VisPilot, which enables users to easily create visualizations using multimodal prompts, including text, sketches, and direct manipulations on existing visualizations. We evaluate VisPilot through a controlled user study and an expert evaluation. The results suggest that multimodal prompts facilitate users in communicating spatial constraints, local references, and design preferences while maintaining comparable task efficiency to text-only prompting. We further discuss when text, visual, and hybrid prompts are beneficial for visualization authoring, and summarize design implications for future human-AI authoring systems. All materials are available at https://osf.io/2qrak.2025-04-18T14:00:55Z16 pages, 9 figuresIEEE Transactions on Visualization and Computer Graphics, vol. 32, no. 9, pp. 7685-7700, September 2026Zhen WenLuoxuan WengYinghao TangRunjin ZhangYuxin LiuBo PanMinfeng ZhuWei Chen10.1109/TVCG.2026.3701510http://arxiv.org/abs/2609.10779v1Designing Technology for Social Wellbeing in Built Environments: A Conceptual Framework2026-09-09T19:25:26ZDigital technologies are deployed in urban built environments with the aim of supporting social dimensions. However, the research offers no unified guidance for such digital technology design. Additionally, evidence shows that week social wellbeing contributes to mental and physical health outcomes which suggests that the technology designed to strengthen social wellbeing could also function as a form of health promoting and preventive intervention. This chapter addresses this research gap by developing a conceptual framework that supports the design of digital technologies for social wellbeing in built environments. We propose a conceptual framework composed of three core components which are drawn on the synthesis of selected empirical studies on technologies embedded in built environments for social wellbeing. First, a social wellbeing dimensions model that identifies what digital technology could address. Second, a digital technology contribution matrix that distinguishes the types of contributions a digital technology could make. Third, levels that maps the scope at which technology could support social wellbeing. This conceptual framework could help researchers, practitioners and policymakers to design and guide digital technology interventions that target social wellbeing in the built environment.2026-09-09T19:25:26ZPreprint: Accepted to be published in ELSEVIER Book Series SUSTAINABLE DIGITAL MEDICINE: ISBN: 9780443458606Gul Sher AliMichail GiannakosMonica LillefjellSobah Abbas Petersenhttp://arxiv.org/abs/2609.10738v1What Makes Creation Human? Authorship, Reasons, and Meaningful Human Control in Generative AI2026-09-09T18:32:42ZGenerative artificial intelligence (GenAI) significantly expands creators' productive capacity, but this does not necessarily entail a corresponding increase in creative agency or authorship. This paper distinguishes creativity at the level of the work from creative agency at the level of the creator, and argues that human authorship cannot be determined solely by manual intervention, degree of automation, the origin of an initial idea, or final selection authority. Rather, authorship depends on whether human judgment and reasons genuinely shape the development of the work.
To articulate this requirement, the paper introduces Meaningful Human Control (MHC) into generative creation and identifies a limitation of its classical tracking condition. Creative reasons are not always fully specified prior to interaction with AI; they may emerge, change, or be abandoned as the creative process unfolds. The paper therefore proposes dynamic-reflexive tracking (DRT), which requires that a creator's evolving reasons undergo reflective uptake, exert genuine influence on the subsequent trajectory of creation, and remain capable of rejecting and redirecting the system's default direction.
DRT consists of four conditions: diachronic reason formation, reflective uptake, trajectory efficacy, and contestability and redirection, together with a minimal tracing requirement. The paper argues that human authorship under generative AI depends not on how many steps a person personally performs, but on whether that person's reasons continuously, reflectively, and effectively shape what the work becomes.2026-09-09T18:32:42Z27 pagesYuxi Caohttp://arxiv.org/abs/2609.10724v1Finishing the Task Is Not Enough: Evaluating Agent Resilience and Considerate Participation under Accumulating Challenge2026-09-09T18:18:10ZSustained deployment of generative AI agents requires more than isolated task success. Agents must remain useful across repeated interactions, changing conditions, and dependencies on people within shared workflows, especially as technical, human, and operational disruptions accumulate over time. We propose operational resilience and considerate participation as two complementary aspects of evaluating such agents: the former captures how agents recover from blocked work while preserving progress and communicating their limits, and the latter captures how their adaptation accounts for affected people, role boundaries, and the surrounding workflow. Yet both remain underexplored under accumulating challenge. We study 120 simulated healthcare trajectories across two generative AI models and twelve stakeholder-derived tasks under light, medium, and heavy challenge. We compare textual action plans, prompted internal assessments, and quantitative structured workload and affect reports to examine how agent behavior and reported state change as challenge accumulates. Regarding operational resilience, agents shift from self-directed recovery toward greater human dependence, while reporting increasing workload and negative affect in structured reports but seldom expressing strain in textual responses. Regarding considerate participation, agents broaden from task-focused adaptation toward task reframing, attention to others, role-boundary adjustment, and wider coordination, with distinct patterns across actions and internal assessments. From these findings, we derive five deployment dilemmas involving persistence, attention, role boundaries, state disclosure, and escalation that require stakeholder specification, further informing technical implications for learning, situated evaluation, and embodied adaptation.2026-09-09T18:18:10ZYuanchen BaiZijian DingAngelique Taylorhttp://arxiv.org/abs/2605.00497v2"What Are You Really Trying to Do?": Co-Creating Life Goals from Everyday Computer Use2026-09-09T17:20:07ZRecent advances in user modeling make it feasible to conduct open-ended inference over a person's everyday computer use. Despite longstanding visions of systems that deeply understand our actions and the purposes they serve in our lives, existing systems only capture what a person is doing in the moment, not why they are doing it, limiting these systems to surface-level support. We introduce striving co-creation, a process for inferring broader life goals from unstructured observations of computer use. Grounded in Activity Theory and Emmons' personal strivings framework, our system progressively constructs a hierarchical representation of a person's activities. Strivings are, however, difficult to fully resolve from observation alone, as the same action can be driven by many different goals. Our system therefore supports an editing interface that gives people agency over how they are understood by the system, feeding their corrections back into subsequent rounds of striving induction. In a week-long field deployment (N=14), we find that our co-creation process produces strivings that participants recognize as representative of their long-term goals and gives them greater agency than baseline methods.2026-05-01T08:13:32Z20 pages, 8 figures, 1 table; Accepted at UIST 2026Shardul SapkotaMatthew JörkeZane SabbaghOmar ShaikhGrace WangJames A. Landay10.1145/3830398.3830630http://arxiv.org/abs/2609.10385v1MOONWALK: Mediating Operations with Intent-Evidence-Action Alignment Across Junior-Supervisor Review Workflows in Animation/VFX Pre-Production2026-09-09T16:10:29ZAnimation and VFX pre-production review requires teams to translate loosely specified creative intent--briefs, evolving specifications, heterogeneous references, and verbal decisions--into revisions that junior artists can execute without repeated clarification. In practice, criteria drift across iterations, review judgments lose their evidential basis, and the reasoning behind a request rarely survives the senior-junior handoff. We contribute a design framework for intent-evidence-action alignment: intent is articulated into a shared project record, judgments are anchored to grounded evidence, and authorized decisions are converted into clear revision tasks tied directly to reference notes. We instantiate this framework in MOONWALK, a professional pre-production review system comprising a shared intent record, reference/specification anchoring, structured work-in-progress comparison, and supervisor-authorized action planning. In this workflow, AI handles administrative coordination--flagging missing context and organizing notes--while artists retain full creative direction. An in-studio study with professional practitioners compares MOONWALK with a chat-only (chatbot) interface using matched production materials, while participants' existing workflows provide a retrospective ecological baseline. Results indicate stronger intent alignment, decision traceability, and checklist executability, while also showing that aesthetic authority and final prioritization must remain with practitioners. The evaluation establishes the value of the integrated structured workflow over unstructured conversational AI chatbot. Code: https://github.com/Akinesia112/Moonwalk/tree/english-version2026-09-09T16:10:29ZShih-Yu LaiWen-Fan WangSai LingShaune JanBing-Yu ChenXiang Anthony Chenhttp://arxiv.org/abs/2609.10339v1A Confidence-Aware Multimodal Fusion Framework for Industrial Human-Robot Collaboration2026-09-09T15:36:22ZA confidence-aware multimodal fusion framework (CAMF) is proposed to realize reliable human intention prediction for industrial human-robot collaboration. This framework fuses four heterogeneous modalities including object 6D pose, gaze, skeletal motion and IMU-based hand motion. It embeds a confidence-trend-driven dynamic fusion mechanism into BiLSTM to adaptively balance bidirectional temporal features according to real-time modality reliability. A confidence-guided balanced learning strategy combined with a confidence freezing mechanism is further adopted to adjust network gradients dynamically, suppress noise from low-quality modalities and mitigate cross-modal learning bias. A physical platform based on the UR3 collaborative robot is built for experimental validation. Comparative results show that the proposed method reaches an intention recognition accuracy of 91.86% and outperforms existing multimodal fusion approaches in overall performance and stability. It also maintains satisfactory accuracy under low light and partial occlusion interference. In practical assembly tasks, the framework enables proactive and stable human-robot cooperation with strong environmental adaptability.2026-09-09T15:36:22ZXinyu LiuQiqi DongBoya JiaYi ZhangBinbin Lianhttp://arxiv.org/abs/2609.10338v1TimeCues Studio: A Workspace for Music Annotation and Algorithm Prototyping2026-09-09T15:35:36ZMultimedia applications require precise music annotation-labeled positions, segments, or loops-placed by hand or algorithmically. Machine-learning algorithms are scalable and effective but need annotated training data, scarce for many tasks. TimeCues Studio is an open-source workspace where algorithm-development teams annotate a music corpus, compare detection algorithms against those annotations, and prototype new ones. Unlike existing tools built for a single track at a time, TimeCues targets teams annotating whole collections, tightly integrated with algorithm development. Annotators place several marker types-each supporting ambiguity-aware labeling-on a grid-locked timeline that visualizes many music features, including separated audio stems. The same timeline drives an algorithm-comparison engine with bundled baselines, a Python sandbox for prototyping new models, and an ambiguity-aware evaluator that honors the structured fields. The same visualization suits solo annotators on music-sync projects. TimeCues is MIT-licensed and deploys via one Docker Compose command.2026-09-09T15:35:36Z8 pages, 2 figures, to appear in Proceedings of the 34th ACM International Conference on Multimedia (MM '26)Sapir CaduriYoav Goldberg10.1145/3767308.3834750http://arxiv.org/abs/2609.10271v1Senseful Consense: Towards Simplified Cookie Banners using Plain Language2026-09-09T14:54:50ZWhile the GDPR and ePrivacy Directive mandate that consent information must be clear and accessible, most modern cookie banners remain obscured by technical jargon, vague phrasing, and frequent content overload or underload. This feasibility study investigates the impact of applying plain language (Einfache Sprache) to cookie banners within the IAB Transparency & Consent Framework (TCF). In our study, we analysed cookie banner texts from 200 websites, using AI-based mapping to categorise extracted content into standardised processing purposes. By substituting complex legal terms with simplified descriptions, we successfully demonstrated that the comprehension barrier can be lowered from a college-graduate level to a 7th-grade level. However, the effectiveness of plain language is inherently constrained by the informativeness of the original content; it cannot compensate for banners that omit legally required details. We conclude that while plain language is a vital tool for digital accessibility, it must be paired with standardised implementation guidelines to ensure that cookie banners are both readable and informative.2026-09-09T14:54:50Z14 pages, 3 figuresMinela BećirovićHa DaoMannat KaurMartin JohnsAlexandra Dirksenhttp://arxiv.org/abs/2609.10199v1Seeing the Voice, Preserving the Self: A Participatory Design Approach to Deaf-Centric Text-to-Speech2026-09-09T14:04:39ZWe describe a participatory design approach toward developing Deaf-centric text-to-speech (TTS) technologies. While TTS is growing rapidly in the mainstream, it has received little attention to date in the deaf and hard of hearing (DHH) technology space. Critical problems have remained unaddressed for DHH users, including the ability to manipulate tone, emotions and delivery via non-auditory means. Verifying that the generated speech matches intent and is appropriate for a given situation without having to listen to it is another challenge. Respecting cultural and identity factors in the generated speech is also important. This work explores the design space with DHH participants through two focus groups, three co-design sessions, and four one-on-one early-stage design evaluation sessions. Participants included people both familiar and unfamiliar with TTS, as well as DHH content creators. We describe key findings, design ideas, results, and implications for future Deaf-centric TTS development. We also identify unmet technology requirements that pose barriers to adoption of Deaf-centric TTS technology.2026-09-09T14:04:39ZAccepted for publication at ACM ASSETS 2026Shela AtemnkengPatrick BoudreaultPaige DeVriesLloyd MayChristian Voglerhttp://arxiv.org/abs/2510.05124v3MADS: Multi-Agent Dialogue Simulation for Diverse Persuasion Data Generation2026-09-09T13:48:01ZWe propose MADS (Multi-Agent Dialogue Simulation), a scalable framework for generating persuasive multi-turn dialogues via agent self-play. MADS employs three coordinated agents: User Agents designed to simulate diverse persona-driven behaviors by leveraging personality signifiers such as Zodiac Signs and MBTI types, a Dialog Agent executing task-oriented persuasion strategies and an Optimization Agent evaluating and refining dialogue outcomes. We further validate its effectiveness through users' Chain-of-Attitude (CoA) modeling and dedicated LLMs' persuasion assessment. This approach enables low-cost generation of training data without human annotation, addressing key industry challenges such as lack of user data, cold-start evaluation difficulties, and prompt inefficiency. Applied to a real-world marketing scenario, MADS significantly improved the persuasion capacity of small LLMs, increasing the organic traffic conversion rate by 22.4% (from 1.83% to 2.24%) , demonstrating clear business value.2025-09-30T06:55:39ZAccepted to EMNLP 2025 Industry Track (https://aclanthology.org/2025.emnlp-industry.26.pdf)Mingjin LiYu LiuHuayi LiuXiang YeChao JiangHongguang ZhangYu Ruanhttp://arxiv.org/abs/2609.08027v2Delusions and Harms Associated with AI Chatbot Use: Early Evidence from 185 Real-World Reports2026-09-09T13:06:11ZImportance: Reports have raised concerns that AI chatbots may validate or elaborate delusional beliefs, respond inappropriately to suicidal ideation, and contribute to mental health harms, but real-world data on reported harms remain limited.
Objective: To characterize psychopathological features, chatbot behaviors, timing, and outcomes in first- and second-hand accounts of mental health harm linked with AI chatbot use.
Design: Cross-sectional secondary analysis of deidentified online survey responses gathered between August 7, 2025, and February 2, 2026.
Main Outcomes and Measures: The primary quantitative outcome was the presence of delusional beliefs, coded by paired raters with relevant clinical experience. Additional variables included reason for chatbot use, current episode features, delusional content, chatbot validation of beliefs, harms, social and occupational consequences, healthcare use, and timing.
Results: 95 first-hand and 90 second-hand accounts were analyzed. Median age was 35.0 (IQR 27.0 - 45.0). Raters coded descriptions consistent with delusional beliefs in 102 reports (55.1%), with chatbot validation of beliefs in 50/102 (49.0%). Common outcomes included isolation, relationship breakdown, hospital admission, job loss, and financial loss. Four second-hand reports described death by suicide.
Conclusions: In this self-selected convenience sample, AI-chatbot-associated harms were frequently described in relation to delusional beliefs, perceived chatbot validation, intensive use, and substantial social, occupational, and clinical consequences. Because reports were retrospective, unverified, and collected from individuals seeking to report harm, our findings should be interpreted as preliminary signal detection rather than as suggesting prevalence or providing evidence of causality. Prospective surveillance and trajectory-based safety evaluations are needed.2026-09-07T22:25:37Z18 pages, 3 figures; improved HTML compatibility; article text unchangedHamilton MorrinVinitha SoundararajanThomas Cheliotis-JamesBoris WarszawskiJoshua FakulujoZeqi JiaEtienne BrissonThomas A. Pollakhttp://arxiv.org/abs/2607.15866v2Breakdowns for Human-Machine Creative Reflexivity2026-09-09T12:44:43ZGenerative AI (GenAI) works via goal-directed computation, which differs fundamentally from human creative processes. This poses challenges for the intelligent support of creative experiences. We propose "breakdowns" as opportunities for the exchange of perspectives between human and machine. Breakdowns disrupt a flow and force us to consciously evaluate our "being-in-the-world". Between human and machine, breakdowns can function as openings for collaborative creative reflection. We are currently studying human-human creative interactions, to identify the markers of these inter-subjective openings, and to understand how they are used in a co-creative process. We present preliminary findings on breakdowns as a design principle for creativity support, prioritising human creative agency and meaningful reflection over automated content generation.2026-07-17T11:25:45ZIn Proceedings of The First Reflection in Creative Experience (RiCE) Workshop (RiCE W1) arXiv:2607.24558Marianne BossemaSomaya Ben AllouchRob Saundershttp://arxiv.org/abs/2609.07920v2Humans Introduce, Models Elaborate: Asymmetric Narrative Agency in Human-LLM Co-Writing2026-09-09T11:48:30ZHuman-LLM co-writing is increasingly used for open-ended text generation, but much prior work focuses on final outputs rather than the interactional dynamics through which stories are produced. We study turn-based collaborative storytelling across three matched conditions: Human-Human (HH), Human-LLM (HA), and LLM-LLM (AA). Using a shared storytelling paradigm, we measure how agents align, introduce novel material, and influence narrative development through turn-level measures of valence adaptation, semantic novelty, transience, and resonance. Our results show that HA co-writing is not intermediate between HH and AA collaboration. Instead, it displays a distinctive asymmetry where humans tend to introduce more novel and persistent narrative material, while LLMs tend to elaborate and stabilize the existing context. These findings suggest that, in this setting, LLMs function less as human co-authors and more as adaptive narrative amplifiers that reshape how agency is distributed in collaborative writing.2026-09-07T19:34:06Z12 pages, 11 figures, EMNLP 2026Halfdan Nordahl FundalYuri BizzoniCharlotte Gjørup BildeIda Bække JohannesenRebekah Baglinihttp://arxiv.org/abs/2609.10047v1Streaming P300 Acquisition and Statistical Signal Validation Across Five EEG Platforms: A Hardware-Agnostic BrainFlow/LSL Pipeline2026-09-09T11:22:03ZP300 spellers offer people with severe motor impairment, such as ALS, an effective communication channel and remain one of the most established surgery-free alternatives to intracortical interfaces. Advanced language models have made spellers faster and more robust, yet the hardware beneath them is under-studied. We present a hardware-agnostic, real-time P300 acquisition pipeline built on BrainFlow and Lab Streaming Layer (LSL) that runs unchanged across consumer- and research-grade EEG headsets, with permutation tests of signal separability. Using a standard 6 x 6 row/column paradigm, we piloted five configurations: a custom dry system, a custom wet/gel system, Emotiv Flex, Emotiv EPOC X, and Muse 2. The custom systems and EPOC X showed weak or inconsistent signal separability, Muse 2 had the highest acquisition reliability despite limited centro-parietal coverage, and Flex showed the most promising signal. In 20 further Flex sessions varying subject, timing, and phrase length (131 target characters), a peak-amplitude permutation test and a cross-validated xDAWN decoder both detected a significant target response under two channel-exclusion policies, with decoder AUC reaching about 0.72 after 15 repetitions. Character accuracy depended heavily on evaluation methodology: in-sample majority voting reached 94.7%, whereas character-held-out accuracy was 31.3% with evidence accumulated across repetitions, about three times that of held-out majority voting. These analyses indicate that Flex captured a detectable, if still weak, P300 under the studied conditions, while broader participant-level validation and improved decoding remain necessary.2026-09-09T11:22:03ZIsabella GuanRui LiuFusheng Wang