https://arxiv.org/api/PEpiPd/FXyKhM8TULTogIACtcoo2026-07-21T14:04:38Z236657515http://arxiv.org/abs/2607.09771v1The backbone of science: analysis of citation networks between papers and their sources2026-07-07T14:26:52ZThe bibliography of scientific papers lists items with variable degree of relevance for the contents of the paper itself. If we could identify the sources, i.e., the works that actually inspired the paper, their citations can help us uncover the genesis of scientific projects and would be more representative of the actual importance of papers and authors than the standard citation counts, when all references are considered. Here we present an analysis of the \textit{backbone of science}, i.e., the network of citations between papers and their sources. The latter are extracted from the full body of papers via Large Language Models (LLMs), which are currently very capable of correctly identifying the context in which a paper is cited. Using two different but related prompts, we find that the LLMs select only a small set of references, not taken at random, and that the resulting backbone networks are quite similar to each other with respect to their in-degree distributions, modularity, transitivity, and degree correlations. Backbone networks have higher heterogeneity in their in-degree distributions, compared to the full network, but the most cited papers are usually the same, with some important exceptions. Citation rankings among authors are also remarkably stable. We conclude that the full citation network, despite its redundancy with respect to the backbones, presents a reliable picture of the relative citation impact of papers and authors.2026-07-07T14:26:52Z26 pages, 8 figures, and 21 tablesWonhee JeongDimitri MarinelliSatyaki SikdarGaurang Singh YadavSanto Fortunatohttp://arxiv.org/abs/2607.06260v1Large language models create an uneven informational layer over cities2026-07-07T13:26:38ZLarge language models (LLMs) are emerging as a new informational layer over cities, shaping which places people discover, consider, and ultimately visit. Yet little is known about which places they surface, which they ignore, and whether these patterns vary across communities and users and translate into real-world economic consequences. Here, we audit restaurant recommendations from three major LLMs across 304 neighborhoods in five U.S. cities using 320 synthetic user profiles spanning income, age, sex, and residential status. We find that LLMs both fabricate venues and systematically overlook real ones. Fabrication is concentrated in neighborhoods with weaker digital and physical footprints and disappears when models are provided with verified venue lists. In contrast, invisibility persists: even when choosing from a fixed set of real venues, 47.5% of establishments are never recommended, and 31.9% of these blind spots are shared across all three model families, indicating that uneven visibility reflects not only missing knowledge but also stable patterns of selective attention rooted in shared patterns of visibility rather than model-specific errors. The same selectivity extends to users. Within identical venue pools, higher-income users receive more expensive and less popular venues, while tourists are directed toward costlier but more socially diverse establishments than local residents. Simulating the resulting shifts in consumer demand suggests that widespread reliance on LLM recommendations would redirect visits and revenue away from chain and quick-service restaurants toward independent and full-service dining. Together, our findings show that LLMs act as a selective layer of urban information that unevenly distributes visibility across places and people, with potential consequences for local economies and urban inequality.2026-07-07T13:26:38ZLin ChenGuangyuan WengEsteban Morohttp://arxiv.org/abs/2607.06080v1From Blueprint to Reality: Modeling and Applying Putnam's Social Capital Theory with LLM-based Multi-agent Simulations2026-07-07T09:49:23ZPutnam's Social Capital Theory is a foundational framework for collective action and community prosperity. However, traditional empirical methods face practical limits on control and replication. Meanwhile, LLM-based social simulations are typically behavior-driven and lack theory-aligned environments for modeling Putnam's core propositions. To address these gaps, we introduce SocaSim, an LLM-based multi-agent simulation framework to study Putnam's Social Capital Theory from theoretical blueprint to simulated reality. Specifically, we build an environment integrating social network evolution, trust dynamics, and norm propagation, where agents engage in repeated collective-action experiments, and then apply the three dimensions to analyze adaptation challenges in smart elderly care. Our simulations reproduce Putnam's macro-level patterns and exhibit strong human-agent alignment at the group level. Unlike traditional methods, SocaSim traces micro-level causal pathways of social network, trust, and norms via round-by-round simulations and counterfactual interventions, enabling process-level interpretability. Taken together, these capabilities establish a research paradigm that leverages LLM agents to bridge social science and computer science.2026-07-07T09:49:23Z23 pages, 13 figures, 11 tablesShiyi LingZhi ZhengHui ZhengWenjun XueFeng YeTong Xuhttp://arxiv.org/abs/2607.05952v1Signed-Graph Recommendation as Structural Consistency Maximization2026-07-07T07:53:23ZWhile signed social recommendation has shown great potential by modeling both trust and distrust relations, its effectiveness is often hindered by structural noise and data sparsity. In this work, we first identify a fundamental inconsistency across the structural, propagation, and semantic layers of existing models, which leads to biased representations learned from sparse or noisy datasets. Furthermore, we observe that most existing methods treat the observed graph as fixed, failing to bridge the gap between noisy topologies and reliable social semantics. To address these issues, we propose a unified framework named SSC-Loop that treats signed social recommendation as the maximization of structural consistency. SSC-Loop includes three dedicated modules: ESA-DA for structural consistency, a P/N/O propagation mechanism for propagation consistency, and a contrastive learning objective for semantic consistency. Experiments on Epinions demonstrate that SSC-Loop achieves strong performance on explicit signed social rating prediction, while auxiliary results on Slashdot under a derived link-existence setting further suggest its ability to exploit signed social structures. Source code is available at https://github.com/Refrainwww/SSC-Loop.2026-07-07T07:53:23ZZifan WangSiyu ChenWenzhuo Songhttp://arxiv.org/abs/2605.07056v2The University AI Didn't Replace -- Rethinking Universities in the AI Era2026-07-07T04:54:47ZGenerative artificial intelligence (AI) is reshaping higher education, yet many universities remain in early stages of adoption where AI innovation occurs informally and without institutional recognition. This paper presents a framework describing four levels of AI adoption in universities and illustrates these dynamics through a case study of AI-enabled curriculum initiatives in several units. We contend that the key institutional challenge is moving from isolated innovation to strategic integration, where universities redesign learning around AI-supported reasoning and align policies, workload models, and recognition systems to support educational transformation.2026-05-08T00:07:55Z8 pages, 1 figure. Position paper on Generative AI and the transition from isolated educational innovation to institutionally supported adoption in higher educationKarol P. BinkowskiAndrew Hopkinshttp://arxiv.org/abs/2604.21457v2Context-Aware Displacement Estimation from Mobile Phone Data: A Methodological Framework2026-07-07T04:53:18ZTimely population displacement estimates are critical for humanitarian response during disasters, but traditional surveys and field assessments are slow. Mobile phone data enables near real-time tracking, yet existing approaches apply uniform displacement definitions regardless of individual mobility patterns, misclassifying regular commuters as displaced. We present a methodological framework addressing this through three innovations: (1) mobility profile classification distinguishing local residents from commuter types, (2) context-aware between-municipality displacement detection accounting for expected location by user type and day of week, and (3) operational uncertainty bounds derived from baseline coefficient of variation with a disaster adjustment factor, intended for humanitarian decision support rather than formal statistical inference. The framework produces three complementary metrics scaled to population with uncertainty bounds: displacement rates, origin-destination flows, and return dynamics. An Aparri case study following Super Typhoon Nando (2025, Philippines) applies the framework to vendor-provided daily locations from Globe Telecom. Context-aware detection reduced estimated between-municipality displacement by 1.6-2.7 percentage points on weekdays versus naive methods, attributable to the commuter exception but not independently validated. The method captures between-municipality displacement only. Within-municipality evacuation falls outside scope. The single-case demonstration establishes proof of concept. External validity requires application across multiple events and locations. The framework provides humanitarian actors with operational displacement information while preserving individual privacy through aggregation.2026-04-23T09:14:11Z24 pages, 4 figures, 14 tables. Case study: Super Typhoon Nando, Philippines (2025)Rajius IdzalikaMuhammad Rheza MuztahidRadityo Eko Prasojohttp://arxiv.org/abs/2607.05689v1UCSC NLP at SemEval-2026 Task 10: Boundary-Aware Span Extraction and RoBERTa Classification for Conspiracy Detection2026-07-06T23:15:52ZWe present our systems for SemEval-2026 Task 10 (PsyCoMark), addressing conspiracy marker extraction (Subtask 1) and document-level conspiracy detection (Subtask 2). For marker extraction, we formulate the task as multi-label span classification over enumerated candidate spans, using IoU >= 0.95 positive labeling, hard-negative sampling, and containment-based non-maximum suppression (NMS) with boundary-aware span representations. Document classification is modeled independently using a sequence classifier with label smoothing and a stratified train-validation split. Analysis shows that entity-like roles (Actor, Victim) are detected robustly, while abstract roles (Action, Effect, Evidence) remain sensitive to boundary criteria. On the official test set, our systems rank 7th in Subtask 1 (0.2251 macro F1) and 11th in Subtask 2 (0.7694 weighted F1).2026-07-06T23:15:52Z6 pages, 2 tables. System description paper for SemEval-2026 Task 10 (PsyCoMark: Psycholinguistic Conspiracy Marker Extraction and Detection)Dom MarhoeferMilos SuvakovicGlenn Grant-RichardsAidan PineroRyan Kinghttp://arxiv.org/abs/2607.01148v2Emergence of Preferential Attachment and Glass-Ceiling Effects in Autonomous Networks of LLMs2026-07-06T21:34:41ZWe investigate the emergence of structural disparities in networks of collaborating large language model (LLM) agents. When LLM agents autonomously choose collaborators, the resulting communication network exhibits preferential-attachment dynamics: agents that are already prominent become increasingly likely to attract additional connections. In some cases, weaker LLM agents (agents with smaller base model or older version) can disproportionately occupy central and influential network positions relative to stronger LLM agents. We interpret this as a type-dependent glass-ceiling effect (GCE). We model the network of LLM agents as a time-evolving sequence of directed weighted graphs, where the vector-valued edge weights represent cumulative tokens exchanged, number of interaction rounds, and reasoning effort. Using a contraction mapping argument on the mean-field dynamics, we prove that the importance (centrality) of each agent type converges to a unique stable equilibrium. To ground the model in LLM decision mechanisms, we introduce a cross-attention-inspired utility for collaborator selection. This utility specifies the local connection dynamics and, together with the mean-field model, yields a predictive characterization of the limiting network structure and its type-dependent centrality gaps. To validate the theory, we develop an experimental testbed with 100 LLM agents. Our experiments show that autonomous network formation can generate persistent centrality disparities, with their magnitude and direction depending on model family, model size, system-prompt design, and task context. They further show that the effect of preferential attachment depends on its alignment with model capability: reinforcing it improves collective performance when stronger agents become central, whereas weakening it improves performance when network dynamics instead favor weaker agents.2026-07-01T16:28:48ZYiming ZhangVikram Krishnamurthyhttp://arxiv.org/abs/2607.05203v1Finfluencers on TikTok: A Longitudinal Analysis of Content, Engagement, and Disclaimer Practices2026-07-06T15:17:18ZThe rise of social media financial influencers (finfluencers) has transformed how financial information is disseminated to broad and often inexperienced audiences. While these creators may contribute to financial literacy, concerns remain regarding the reliability of their content and the adequacy of risk disclosures.
Using data collected through TikTok's Research API, we analyze UK finfluencer content, engagement dynamics, disclaimer practices, audience sentiment, and network structure. The primary dataset comprises 13,215 videos and 104,097 comments posted by 71 UK-based finfluencers between April and September 2024, while a follow-up dataset covering October 2025 to March 2026 enables longitudinal analysis of disclaimer practices, engagement trends, and hashtag usage.
Using topic modeling, we identify four dominant themes: Entrepreneurship \& Side Hustles, Property Investing, Active Trading, and Saving \& Budgeting. Sentiment analysis of audience comments reveals predominantly neutral-to-positive responses, while engagement analysis shows only a negligible association between video duration and engagement rate. Social network analysis indicates a collaborative ecosystem in which mid-tier finfluencers frequently act as bridges between creator groups. Explicit disclaimers and risk-related language remain relatively uncommon overall and are concentrated primarily in trading-related content.
The findings highlight challenges related to financial transparency and disclosure practices within short-form financial content ecosystems. We discuss implications for consumer protection and the design of clearer and more standardized financial risk disclosures on social media platforms.2026-07-06T15:17:18ZAccepted for presentation at MSNDS 2026, an associated workshop of ASONAM 2026Essam GhadafiPanagiotis Andriotishttp://arxiv.org/abs/2606.29596v2Boundary Degree as a Node-level Feature for Epidemic Scenario Identification in Agent-based Cascade Simulations2026-07-06T10:54:24ZCharacterizing the scenario underlying an epidemic from its disease cascade is an important task in simulation analytics. We propose boundary degree, the count of an infected node's contacts in the underlying contact network that were not infected, as a per-node cascade feature for this task. Through systematic ablation on realistic social contact networks of Tennessee and Virginia, we show that boundary degree alone improves scenario identification accuracy by 19%. Edge features, whose importance was observed empirically by prior work, consistently improve accuracy across all settings; we provide theoretical grounding for this observation. These effects are complementary. We prove that certain epidemic scenarios are indistinguishable without boundary or edge information. Prior feature engineering approaches included aggregate boundary statistics, but these were not among the top-ranked feature groups; the per-node representation we propose reveals their importance clearly. Our results suggest that contact tracing applications should track contacts with non-infected individuals, not only transmissions.2026-06-28T20:42:01Z28 pages, 10 figures, preliminary version; not finalAmro Alabsi AljundiGalen HarrisonJiangzhuo ChenAbhijin AdigaAnil Kumar VullikantiMadhav V. Marathehttp://arxiv.org/abs/2607.05469v1Breaking Structural Isolation: Scalable Graph Clustering via Community-Aware Sampling and Structural Entropy2026-07-06T07:55:53ZUnsupervised graph clustering is a fundamental technique for uncovering underlying semantic patterns in large-scale networks. Although Graph Contrastive Learning has demonstrated promising performance, existing methods often suffer from the "structural isolation" issue during mini-batch training, making it challenging to capture cohesive community structures that characterize the global topological distribution. To address these challenges, we propose SCISE, a Scalable unsupervised graph Clustering framework that preserves structural Integrity by synergizing community-aware sampling with constrained Structural Entropy. Specifically, we first introduce the Structural Entropy Community Constraint operator (SECC), which optimizes structural information within a constrained solution space to mitigate community fragmentation and enhance partition cohesion. Second, to prevent global information loss during batch training, we design a Community-Aware Sampling Expansion (CSampE) mechanism that incorporates the community context of target nodes into sampling batches, effectively breaking structural barriers and preserving topological integrity. Finally, we devise a Structural Contrastive Learning (StructCL) module that refines edge weights based on intra-batch structural similarity, guiding the encoder to learn representations in a higher-order structural space. Extensive experiments on six mainstream benchmark datasets demonstrate that SCISE significantly outperforms state-of-the-art algorithms, with ablation studies and robustness analyses further validating its effectiveness and reliability for real-world large-scale graphs.2026-07-06T07:55:53ZAccepted to the Proceedings of the VLDB Endowment (VLDB 2026). 18 pages, 15 figures, 15 tablesJingyun ZhangHao PengJianxin LiAngsheng LiPhilip S. Yuhttp://arxiv.org/abs/2607.04220v1The Politics Attention Makes: Platform Media Logic and the Mediatization of Politics2026-07-05T10:24:42ZEmpirical research on social media and politics has primarily treated platforms as distributive systems that expose users to particular messages. The mediatization literature, however, suggests shifting attention upstream: from circulation to production. Under intense competition for platform attention, political actors who depend on visibility face pressure to learn from recurrent differences in reach and engagement - shaping politics around platform media logic. This paper examines that production-side dimension of platforms political impact by introducing attention price analysis: an exploratory method for estimating the differentiated attention returns associated with forms of expression. Using RoBERTa reward models trained on residualized engagement across X/Twitter, Bluesky, and Mastodon, the analysis compares how platform environments reward rhetorical, emotional, epistemic, and relational features of public communication. The attention signal differs sharply across platforms and engagement actions. X/Twitter sharing rewards antagonism while penalizing respect and nuance; Bluesky reposting favors neutral, lower-emotion language; and Mastodon boosts reward reasoning, nuance, compassion, and collective expression. Toxicity is rewarded across platforms, but in bounded and nonlinear ways. The findings suggest that moving from X/Twitter to less engagement-optimized alternatives such as Bluesky and Mastodon does not eliminate attention pressures, but it may reward less antagonistic and more deliberative forms of politics. The paper contributes a production-side approach to social media and politics by making one dimension of platform media logic empirically visible.2026-07-05T10:24:42ZPetter Törnberghttp://arxiv.org/abs/2607.04178v1Dynamic Interest Rate Discovery in Decentralized Finance: A Reverse Kelly Automated Market Maker for Risk-Adjusted Lending2026-07-05T08:51:10ZDecentralized Finance (DeFi) lending protocols currently rely on heuristic, utilization-based bonding curves that mandate severe over-collateralization, systematically excluding under-collateralized assets like corporate invoices. This paper introduces a mathematically optimal pricing mechanism for decentralized credit: the Reverse Kelly Automated Market Maker (rkAMM), the core engine of our proposed lending framework. By inverting the Kelly Criterion, traditionally used for optimal bet sizing, we construct a dynamic interest rate discovery protocol that explicitly prices individual loan risk. The rkAMM ingests real-time Probability of Default (PD) streams from an off-chain Explainable AI oracle and dynamically calculates the exact interest rate required to sustain target liquidity provider (LP) yields. We mathematically derive the Reverse Kelly pricing function ($r = \frac{y + PD}{1 - PD}$), proving its strictly convex superiority over Aave and Compound's static utilization curves in managing capital efficiency. Furthermore, we deploy the rkAMM architecture via Solidity smart contracts, optimizing for gas-efficient 1e18 (WAD) floating-point arithmetic. To ensure decentralized transparency, our simulation infrastructure leverages MLflow for tracking yield hyperparameters, Data Version Control (DVC) linked to DagsHub for versioning Real-World Asset (RWA) data arrays, and localized edge-inference via Ollama (Llama-3) and Hugging Face (FinBERT) for zero-cost predictive modeling. Monte Carlo simulations across 10,000 macroeconomic stress scenarios confirm that the rkAMM maintains protocol solvency and stabilizes LP yields at 12-15\% net of expected credit losses. This work provides the foundational financial engineering required to bridge the \$2 trillion global supply chain finance gap using permissionless blockchain infrastructure.2026-07-05T08:51:10ZSai Srikanth MadugulaPeplluis Esteva de la RosaDaya Shankarhttp://arxiv.org/abs/2607.04042v1An Exploratory Study of Malicious Link Posting on Social Media Applications2026-07-04T22:04:59ZSocial network platforms are now widely used as a mode of communication globally due to their popularity and their ease of use. Among the various content-sharing capabilities made available via these applications, link-sharing is a common activity among social media users. While this feature provides a desired functionality for the platform users, link sharing enables attackers to exploit vulnerabilities and compromise users' devices. Attackers can exploit this content-sharing feature by posting malicious/harmful URLs or deceptive posts and messages which are intended to hide a dangerous link. However, it is not clear how the most common social media applications monitor and/or filter when their users share malicious URLs or links through their platforms. To investigate this security vulnerability, we designed an exploratory study to examine the top five android social media applications' performance when it comes to malicious link sharing. The aim was to determine if the selected applications had any filtering or defenses against malicious URL sharing. Our results show that most of the selected social media applications did not have an effective defense against the posting and spreading of malicious URLs. While our results are exploratory, we believe our study demonstrates the presence of a vital security vulnerability that malicious attackers or unaware users can use to spread harmful links. In addition, our findings can be used to improve our understanding of link-based attacks as well as the design of security measures that usability into account.2026-07-04T22:04:59ZMuhammad HassanMahnoor JameelMasooda Bashir10.14722/usec.2023.234399http://arxiv.org/abs/2604.25596v2Volition-Guarded Multiagent Atomic Transactions: Describing People and their Machines2026-07-04T10:24:56ZFormal models for concurrent and distributed systems describe machines; the people who operate them are either ignored or treated as external environment. Yet, key distributed systems -- notably grassroots platforms -- include people operating their personal machines (smartphones), and their faithful description must include the states of both people and machines and how they jointly effect system behaviour.
Here, we propose volition-guarded multiagent atomic transactions -- executed atomically by machines and guarded by their people's volitions -- as a novel mathematical foundation for specifying systems consisting of people operating machines. Each agent's state consists of a volitional state and machine state; a transaction is enabled when the machine precondition holds and the guarding persons are willing. For example, befriending two people is guarded by both; unfriending, by either; voluntary swap of coins and bonds is guarded by both parties, while a payment is guarded by the payer.
We develop the mathematical machinery to express safety and liveness of platforms specified in this framework, to implement one platform by another, and for an implementation to be resilient to faults; and provide example specifications of two grassroots platforms: social networks, and coins and bonds. These specifications are then used by AI to derive working implementations.
We employ here a novel and simpler definition of `grassroots' that better captures the informal notion -- multiple instances can form and operate independently, yet may coalesce -- and show that the platforms specified here are grassroots under the new definition. We further introduce \emph{volitionally grassroots} protocols, in which two groups can become connected only by mutual consent -- the first transaction coupling them must be willed by a member of each -- and show that both platforms are volitionally grassroots.2026-04-28T13:02:49ZAndy Lewis-PyeEhud Shapiro