https://arxiv.org/api/VHaK7fGCjQbTk4nBEapOBu4sB4I 2026-07-21T17:08:25Z 23665 90 15 http://arxiv.org/abs/2606.28120v2 The Reciprocal Impact of Science and Software: A Cross-Corpus Analysis of How Research Shapes Software and Software Enables Research 2026-07-03T15:07:29Z Software and scientific knowledge co-evolve, yet they are catalogued in separate corpora that rarely speak to one another. We bridge them at global scale by linking World of Code (a near-complete mirror of public version-control history) to Semantic Scholar and OpenAlex through a typed cross-corpus graph of 69.8M edges over eight relation types (paper-to-software mentions, software-to-paper citations, software dependencies, authorship, affiliation, and identity bridges). Anchoring on 18,247 curated science repositories, we ask two reciprocal questions: what is the impact of science on software, and of software on science? To test whether this Science-Software Supply Chain (S3C) view is feasible, we run basic investigations rather than claim a definitive measurement. The two directions appear to illuminate different, complementary strata: the literature's reach into software is dominated by a reproducibility and packaging layer (nf-core, Nextflow, Bioconda) and sequence-analysis tools, whereas software's reach back into science is proxied by a largely invisible machine-learning and data-science infrastructure tier (PyTorch, seaborn, NLTK). The direct paper-names-software channel is too sparse to rank: a human-curated gold benchmark links none of its 65 in-scope cases. Dependency reuse stands in as a proxy and is at most weakly coupled to citation count and to stars (Spearman rho=0.36). Our most cautionary finding is about measurement itself: the reuse-citation coupling flips sign and confidence across two reasonable ways of pairing a repository with a citation count, through papers that name it (n=137, rho=0.05, CI straddling zero) versus DOIs a repository declares for itself (n=1,067, rho=0.13, CI [0.07,0.19]). With linkage this sparse, the sign of a headline correlation depends on which gap one tolerates, so we report both and refrain from a strong decoupling claim. 2026-06-26T14:20:33Z Audris Mockus http://arxiv.org/abs/2211.15590v2 A Bayesian Approach for the Network Reconstruction of Interdependent Critical Infrastructure Systems from Cascading Failures 2026-07-03T13:55:56Z Analyzing the behavior of complex interdependent networks requires complete information about the network topology and the interdependent links across networks. For many applications such as critical infrastructure systems, understanding network interdependencies is crucial to anticipate cascading failures and plan for disruptions. However, data on the topology of individual networks are often publicly unavailable due to privacy and security concerns. Additionally, interdependent links are often only revealed in the aftermath of a disruption as a result of cascading failures. We propose a scalable nonparametric Bayesian approach to reconstruct the topology of interdependent infrastructure networks from observations of cascading failures. Metropolis-Hastings algorithm coupled with the infrastructure-dependent proposal are employed to increase the efficiency of sampling possible graphs. Results of reconstructing a synthetic system of interdependent infrastructure networks demonstrate that the proposed approach outperforms existing methods in both accuracy and computational time. We further apply this approach to reconstruct the topology of one synthetic and two real-world systems of interdependent infrastructure networks, including gas-power-water networks in Shelby County, TN, USA, and an interdependent system of power-water networks in Italy, to demonstrate the general applicability of the approach. 2022-11-28T17:45:41Z Accepted for publication in Physical Review E. Source code available at: https://github.com/MirSaleh/PRE-Paper-Source-Code MirSaleh Bahavarnia Hiba Baroud Yu Wang Jin-Zhu Yu 10.1103/vswp-h4hx http://arxiv.org/abs/2605.03796v4 Capability centrality: the next step from scale-free property 2026-07-03T11:49:24Z In this article we present a new centrality measure called ksi-centrality. We show that ksi-centrality distinguishes real networks from random ones, similar to degree centrality: the ksi-centrality distribution is right-skewed for real networks and centered for random Erdos-Renyi networks, and has linear pattern with a heavy tail on a log plot. Furthermore, the ksi-centrality distribution is centered for models simulating real networks: Barabasi-Albert, Watts-Strogatz, and Boccaletti-Hwang-Latora. Thus, this centrality distribution is an additional and independent property with respect to scale-freeness. We also introduce a normalized version of ksi-centrality and show that it is related to algebraic connectivity and the Chegeer's value of a network. Moreover, the average value of this normalized centrality is in bijective correspondence with the relative number of edges that a new node connects to others in the Barabasi-Albert preferential attachment model, thus answering the question of how to choose the parameter $m$ to model a given real-world network. 2026-05-05T14:25:15Z Mikhail Tuzhilin http://arxiv.org/abs/2607.03233v1 Agentic and Generative AI for Open-Source Intelligence and Cyber Investigations: Taxonomy, Evaluation, Challenges, and Future Directions 2026-07-03T11:42:29Z The rapid growth of publicly available digital information has rendered manual open-source intelligence (OSINT) analysis insufficient for modern intelligence, cybersecurity, and cyber investigation. Large language models (LLMs) and agentic AI systems, capable of tool use, multi-step reasoning, and iterative intelligence generation, have emerged as promising solutions, yet evaluation frameworks have not kept pace with reported capabilities. This survey systematically reviews 74 studies and makes four contributions. First, it establishes agentic AI as a distinct analytical category rather than an extension of LLM prompting, organising the literature through an 11-category taxonomy covering LLM foundations, agentic architectures, retrieval-augmented generation (RAG), knowledge graphs, prompt engineering, domain adaptation, evaluation benchmarks, and risk. Second, it identifies the hallucination-validation gap as a corpus-level finding: although hallucination is recognised as a major reliability concern in over twenty studies, end-to-end hallucination is empirically measured in only one OSINT-specific RAG-based system, non-reproducible conditions, while related reasoning and factual-correction studies evaluate general-domain question answering rather than OSINT. Third, it maps existing research to the OSINT lifecycle, showing strong support for collection and analysis but limited coverage of verification, reporting, dissemination, and decision support. Fourth, it derives a ten-point research agenda addressing evaluation, benchmarking, hallucination measurement, adversarial robustness, dark-web coverage, multimodal intelligence, and governance. It concludes that a human-AI co-pilot model, where LLMs assist collection and triage while analysts retain responsibility for verification and decision-making, represents the most defensible near-term deployment architecture. 2026-07-03T11:42:29Z 36 Pages Eduardo Almeida Palmieri Mohamed Chahine Ghanem Dipo Dunsin Zubair Baig Ed de Quincey Kim-Kwang Raymond Choo http://arxiv.org/abs/1908.07092v11 Linear stability analysis for large dynamical systems on directed random graphs 2026-07-03T08:46:31Z We present a linear stability analysis of stationary states (or fixed points) in large dynamical systems defined on random directed graphs with a prescribed distribution of indegrees and outdegrees. We obtain two remarkable results for such dynamical systems: First, infinitely large systems on directed graphs can be stable even when the degree distribution has unbounded support; this result is surprising since their counterparts on nondirected graphs are unstable when system size is large enough. Second, we show that the phase transition between the stable and unstable phase is universal in the sense that it depends only on a few parameters, such as, the mean degree and a degree correlation coefficient. In addition, in the unstable regime we characterize the nature of the destabilizing mode, which also exhibits universal features. These results follow from an exact theory for the leading eigenvalue of infinitely large graphs that are locally tree-like and oriented, as well as, for the right and left eigenvectors associated with the leading eigenvalue. We corroborate analytical results for infinitely large graphs with numerical experiments on random graphs of finite size. We discuss how the presented theory can be extended to graphs with diagonal disorder and to graphs that contain nondirected links. Finally, we discuss the influence of small cycles and how they can destabilize large dynamical systems when they induce strong enough feedback loops. 2019-08-19T22:47:49Z Typo's have been corrected in the equations (C3), (C8) and(C9), and the caption of Figure 5. The original manuscript also interchanged sOut for sIn, and conversely, which affects Eqs.(56), (60), (D10), (H1), and (H4) Phys. Rev. Research 2, 033313 (2020) Izaak Neri Fernando Lucas Metz 10.1103/PhysRevResearch.2.033313 http://arxiv.org/abs/2504.08152v2 Dynamics of collective minds in online communities 2026-07-03T08:01:37Z Collective discourse and action are driven by collective minds. These shared semantic representations and related processes shape societal responses to critical societal challenges such as climate change and political upheavals. In online communities, collective minds are susceptible to the influences of editorial practices and community dynamics, making them vulnerable to manipulation. However, understanding these influences is difficult because of the limits of experimenting with and predicting complex social systems. Here, we develop a computational model of collective minds, calibrated and validated with data from 400 million comments across five U.S. online news platforms and a survey. Our model enables us to quantitatively describe and experiment with different editorial agenda-setting practices and aspects of community dynamics to understand how they shape the collective mind. We find that some editorial influences can be reversed relatively rapidly, but others, such as amplification and reframing of certain topics, as well as community influences such as trolling and counterspeech, tend to persist and durably change the collective mind. These findings illuminate ways collective minds can avoid manipulation and pathways for communities to maintain healthy and authentic collective discourse amid ongoing societal challenges. 2025-04-10T22:22:40Z 80 pages,21 figures Seungwoong Ha Henrik Olsson Kresimir Jaksic Max Pellert Mirta Galesic http://arxiv.org/abs/2607.02900v1 Angry but Accurate: Detecting and Profiling the Counter-Misinformation Ecosystem on Twitter 2026-07-03T02:52:50Z On social media, many users actively push back against false claims. Understanding who pushes back and how they do so matters, as this corrective activity is central to how misinformation is contested. We study this counter-misinformation ecosystem at scale: applying a domain-specific NLI model from our prior work to a large corpus of COVID-19 tweets, we classify 264,737 posts as supporting or opposing false claims and compare 23 user- and text-level features across the two groups. Contrary to the dominant assumption that negative emotion is a signature of falsehood, we find that anti-misinformation posts are more emotionally negative than pro-misinformation posts, with higher levels of anger, disgust, and sadness. These differences are modest in magnitude but consistent in direction across the negative emotions. We also find that posts opposing misinformation tend to come from more established users, i.e., older accounts, more followers, and higher listed counts. 2026-07-03T02:52:50Z 7 pages, 6 figures. Accepted to ACM Hypertext 2026 Eun Cheol Choi Emilio Ferrara http://arxiv.org/abs/2607.02844v1 Overload-Based Cascades in Multiplex Flow Networks with Partial Functionality 2026-07-03T00:47:18Z Cascading failures driven by load or flow redistribution arise in networked systems such as power grids, supply chains, and cloud computing centers. Most flow-network models assume that a node either functions or fails as a whole. In many real systems, however, a node supports several distinct flows that share node-level resources, and failure in one of them does not necessarily imply failure in the others. We study this setting through multiplex flow networks with partial functionality, where a node can remain operational in some functionalities while failing in others. A heavy load on one functionality reduces the capacity available to the others, as quantified by cross-layer influence factors. When a node fails in one layer, its load is redistributed among surviving nodes in that layer, while the node may continue to operate in the others. Using mean-field analysis, we derive recursive equations for the final system sizes, namely the fraction of surviving nodes in each layer after the cascade stops. We validate the analysis through simulations for several load-capacity distributions. We then examine key features of the cascade dynamics, including non-monotone robustness curves, different cascade-outcome regimes, and their relation with cross-layer influence. We map the outcomes to distinct steady-state regimes, including single-layer survival phases absent in joint-functionality models, and show that partial functionality can increase robustness relative to the joint-functionality case. Finally, we study robustness maximization under a fixed total capacity budget by comparing several capacity allocation strategies. We propose a strategy that combines cross-layer influence with local neighborhood information on load and degree, and show that it gives the strongest robustness performance across the configurations considered. 2026-07-03T00:47:18Z Orkun İrsoy Osman Yağan http://arxiv.org/abs/2607.02760v1 Gendered Pixels: Exploring Gender Differences in Computer-Mediated Self-Presentation among Douyu Live Streamers 2026-07-02T20:54:05Z Live streaming platforms, as computer-mediated communication (CMC) systems, provide streamers with a range of tools, such as webcams, beauty filters, and stream titles, to shape their online personas in ways that either conform to or deviate from viewers' expectations. Drawing on gender role and CMC theories, this study examines how streamers leverage CMC self-presentation tools to fulfill gender role expectations and their associated live streaming outcomes. Analyzing data collected from 867 streamers and 94,227 streams on Douyu, a popular Chinese live streaming platform, we find that although both female and male streamers make extensive use of CMC tools, female streamers are more likely to employ visual strategies, such as webcams and beauty filters, than male streamers. We further find that although different tools have varying associations on streamers' earnings and audience engagement, the benefits of webcam use are weaker for female than for male streamers. These findings underscore the complex interplay between gender roles and technology use in the live streaming domain. 2026-07-02T20:54:05Z Mengxiao Zhu Xiyue Wang Chunke Su Han Zhao Cuihua Shen http://arxiv.org/abs/2504.07337v2 FLASH: Flexible Learning of Adaptive Sampling from History in Temporal Graph Neural Networks 2026-07-02T20:22:59Z Aggregating temporal signals from historic interactions is a key step in future link prediction on dynamic graphs. However, incorporating long histories is resource-intensive. Hence, temporal graph neural networks (TGNNs) often rely on historical neighbors sampling heuristics such as uniform sampling or recent neighbors selection. These heuristics are static and fail to adapt to the underlying graph structure. We introduce FLASH, a learnable and graph-adaptive neighborhood selection mechanism that generalizes existing heuristics. FLASH integrates seamlessly into TGNNs and is trained end-to-end using a self-supervised ranking loss. We provide theoretical evidence that commonly used heuristics hinder TGNNs performance, motivating our design. Extensive experiments across multiple benchmarks demonstrate consistent and significant performance improvements for TGNNs equipped with FLASH. 2025-04-09T23:35:09Z IJCAI 2026, 24 pages, 5 figures, 14 tables Or Feldman Krishna Sri Ipsit Mantri Carola-Bibiane Schönlieb Chaim Baskin Moshe Eliasof http://arxiv.org/abs/2607.02163v1 Visual Analytics of Neighborhood Attribute Profiles for Exploring Structural Equivalence 2026-07-02T13:34:43Z Exploring similar nodes in attributed networks represents a key challenge in data mining. While recent representation learning methods embed networks into low-dimensional vectors, they often implicitly assume a uniform and continuous feature space. This paper proposes a visual analytics approach using dimensionality reduction to help clarify the true topological structure of high-dimensional feature spaces formed by nodes' neighborhood attribute profiles. Analyzing inter-firm transaction networks indicates that structural roles can form complex, non-linear manifolds with density biases. Comparing this feature space with industry classifications suggested: (1) supply chain hierarchies transition continuously; (2) categories treated identically under general semantics can be clearly separated by actual transaction networks; and (3) a single industry label may fragment into multiple regions. These findings suggest potential limitations in assuming identical semantics imply similar structural roles and highlight the possible need for new similarity metrics aligned with manifold topology. 2026-07-02T13:34:43Z 5 pages, 3 figures. Accepted as a Short Paper at IEEE VIS 2026 Kohei Arimoto Masahiko Itoh http://arxiv.org/abs/2607.02627v1 A large-scale dataset of Android applications and their SDK dependencies 2026-07-02T13:19:42Z Mobile applications (apps) increasingly rely on third-party Software Development Kits (SDKs) to provide services such as advertising, analytics, authentication, crash reporting, and location services. These components form an important, but often hidden, layer of the mobile ecosystem. Here, we present a large-scale dataset linking Android apps to the third-party SDKs they integrate. The dataset was constructed by combining app package files (APKs) and metadata from AndroZoo with SDK detection rules provided by Exodus Privacy. We implemented a reproducible, static-analysis pipeline that downloads APKs, inspects their compiled code, and detects SDKs through code-level signature matching: the released dataset contains 334,719 unique app-version observations (associated with 99,722 Android apps) and 246 third-party SDKs - including app-level metadata such as categories, comments, file size, download counts, ratings and SDK-level metadata such as categories, code roots, code versions, names. The dataset supports the construction of a large-scale app-SDK dependency network and its projection onto both layers; in addition, SDKs are mapped to their corresponding operating companies, enabling provider-level analysis of technological concentration and upstream control in the mobile ecosystem. To support reproducible research, the pipeline and dataset are publicly released on GitHub and Zenodo; more broadly, the project provides a reusable research infrastructure for studying third-party technological dependencies, data-collection capabilities, and privacy infrastructures across Android apps. 2026-07-02T13:19:42Z Aurora Gori Savellini Tiziano Squartini Massimo Riccaboni http://arxiv.org/abs/2509.15860v2 PoliTok-DE: A Multimodal Dataset of Political TikToks and Deletions From Germany 2026-07-02T11:59:35Z We present PoliTok-DE, a large-scale multimodal dataset (video, audio, images, text) of TikTok posts from two German elections: the 2024 Saxony state election and the 2025 German federal election. The corpus contains over 930,000 posts, of which over 330,000 were later deleted from the platform (18.7% of Saxony posts, 39.7% of federal posts). In the federal-election collection, about two thirds of the deletions were creator withdrawals, and the platform-deletion rate we computed was 13.0% of all posts, more than an order of magnitude (14-19x) above the platform-wide rate TikTok reported. Posts were identified via the TikTok research API and complemented with web scraping to retrieve full multimodal media and metadata. PoliTok-DE supports social science research across substantive and methodological agendas: substantive work on intolerance and political communication, and methodological work on platform policies around deleted content and qualitative-quantitative multimodal research. To illustrate, we report a case study on intolerance and entertainment in an annotated subset of deleted posts: about one in five posts conveyed intolerance and a majority conveyed humor. 2025-09-19T10:53:58Z Tomas Ruiz Andreas Nanz Ursula Kristin Schmid Carsten Schwemmer Yannis Theocharis Diana Rieger http://arxiv.org/abs/2503.02887v3 Mapping the Intellectual Landscape of Digital Social Networks: A Bibliometric and Citation Network Analysis 2026-07-02T09:39:17Z Network science and digital social network research span sociology, communication, and computational modeling, yet the field's intellectual structure and cross-paradigm connectivity remain insufficiently characterized. Using records retrieved from the Web of Science Core Collection (n = 1,859; queried on 21 November 2024), we conduct a bibliometric, keyword co-occurrence, and citation-network analysis with CiteSpace, VOSviewer, Gephi, and NetworkX. The citation landscape exhibits a strongly centralized core (a largest connected component of 1,293 papers) surrounded by numerous small, weakly connected components, indicating a pronounced core-periphery organization. Keyword clustering and citation-burst analysis further reveal a thematic shift from classic sociological mechanisms (e.g., homophily and tie formation) toward algorithmically mediated communication and misinformation-related topics. We also highlight a set of highly central bridging works that connect otherwise separated thematic communities, suggesting that interdisciplinary exchange is concentrated around a limited number of canonical references. 2025-02-12T05:52:10Z Soc. Netw. Anal. Min. (2026) Pengjia Cui 10.1007/s13278-026-01617-0 http://arxiv.org/abs/2605.02800v2 The Activist's Guide to the Decentralized Social Universe: A Framework for Exploring How Decentralized Social Networks Can Support Collective Action 2026-07-02T00:31:02Z The overreaches of mainstream social media platforms have been extensively reported and studied. For activist communities, these platforms pose risks of surveillance, censorship, or erasure. Decentralized social networks (DSNs) serve as alternative online spaces that appear to prioritize values such as user privacy, free speech, and community control. However, the decentralized ecosystem is vast and complex, making it difficult for communities to understand how to best use these platforms for their organizing aims. We address this gap by proposing a conceptual framework for navigating the DSN landscape that defines core activist community needs -- minimal overhead, community building and reach, on- and offline safety, and operational sustainability -- and links them to concrete platform affordances such as resource efficiency, interoperability, and data ownership. We apply the framework to (1) evaluate and compare the sociotechnical tradeoffs of two contemporary DSNs (Mastodon and Bluesky), (2) understand broader community configurations that emerge across different DSN infrastructures and their implications for collective action, and (3) explore how two distinct activist communities facing infrastructural and political constraints might use the framework to find platforms that align with their needs. We conclude by reflecting on the theoretical promises of DSNs and the structural conditions that shape and constrain participation across them. 2026-05-04T16:42:08Z 28 pages, 1 figure, 2 tables Sybille Légitime Harini Suresh