https://arxiv.org/api/KeFKZPajB/M1cjs7ur2Ad/v1bik 2026-07-21T21:29:43Z 23665 120 15 http://arxiv.org/abs/2606.30801v1 Using AI Agents to Automate Black-Box Audits of Personalization Algorithms at Scale 2026-06-29T18:25:09Z Personalization algorithms determine what content users encounter on online platforms. Auditing these systems is difficult because independent auditors have only black-box access to the algorithms, while personalization depends on users' attributes, behavior, and evolving interaction histories. Existing auditing methods face a tradeoff: studies with real users capture realistic behavior but are costly and hard to control, whereas sock-puppet audits scale more easily but often rely on scripted behavior that limits realism. Beyond this, both approaches struggle to decouple user attributes from user behavior, limiting our ability to causally understand personalization. To address this gap, we introduce a framework for black-box audits of personalization algorithms using generative AI agents as behavioral engines for synthetic accounts. Each agent is instantiated with a fixed persona, grounded in demographic and political survey data, and interacts with a platform's content by reasoning about it and choosing actions. Because behavior is fixed within each persona while platform-visible signals such as age, gender, or location can be experimentally perturbed, our design enables counterfactual auditing of how platforms respond to user attributes. As a case study, we deploy 1,120 agents on X shortly after the 2024 U.S. election, spanning 14 personas and three counterfactual conditions, collecting over 200,000 content exposures. We find that X's algorithmic feed amplifies toxic, polarizing, political, and right-leaning content relative to the chronological feed, with amplification varying sharply by user ideology. Counterfactual analyses show that demographic signals affect content delivery in persona-dependent ways: pooled effects are largely null, while subgroup-level effects vary in direction and magnitude. Our work establishes GenAI-based agents as a new tool for algorithmic auditing. 2026-06-29T18:25:09Z 43 pages, 10 figures Alessandro Morosini Sarah H. Cen Andrew Ilyas Hedi Driss Aleksander Mądry Chara Podimata http://arxiv.org/abs/2606.30395v1 Uncovering Salience-Driven Dynamics in Consumer Confidence with Generative Social Simulation 2026-06-29T14:46:22Z Consumer confidence is typically modeled as a persistent macroeconomic index, yet its movements arise from households that interpret economic information through heterogeneous constraints, exposures, prior beliefs, and attention. We introduce ConsumerSim, a generative Human--Environment response framework that reconstructs Consumer Confidence Index (CCI) dynamics from a microdata-calibrated synthetic population, time-stamped macroeconomic, financial, policy, and news signals, survey-like response generation, post-stratified belief expansion, and behavioral inertia alignment. Across U.S., EU27, and Japanese official CCI target series, ConsumerSim ranks first among persistence, time-series, regression, and information-augmented baselines on the reported reconstruction metrics, with clear gains around high-salience shocks. Its reconstructed signal also improves short-horizon prediction of real activity, most consistently for housing outcomes. Mechanism analyses show that CCI movements concentrate around salient events; subgroup trajectories often align in direction while differing in magnitude; and signal sensitivity varies across income, homeownership, education, and political-alignment groups. Population-expansion and ablation results indicate that representative aggregation, situational signals, persona heterogeneity, and inertia are necessary for both accuracy and diagnosis. The findings support a behavioral view of consumer confidence as an interpretable Human--Environment response process rather than a purely aggregate time series. 2026-06-29T14:46:22Z Yixu Huang Yunlu Yin Jiayu Lin Xinnong Zhang Jia Wang Siyuan Wang Xuanjing Huang Liyin Jin Zhongyu Wei http://arxiv.org/abs/2606.30394v1 When Editors Revolt: Characterizing Journal Declarations of Independence 2026-06-29T14:45:48Z When editorial boards resign from their journals and publishers and declare their independence, two competing journals can result: the original journal under a new editorial board (a "zombie" journal), and a new journal established by the departing editors (a "breakaway"). The bibliometric community saw such an event when the board of Journal of Informetrics left Elsevier to found Quantitative Science Studies. We analyzed 39 breakaway-zombie journal pairs that have formed since 1989 and their declarations of independence to understand why and how they happen. Results show that declarations of independence were motivated by concerns related to governance and business model and overwhelmingly happened at journals owned by the Big Five publishers. Breakaway editors tended to found new journals at smaller publishers and adopt diamond publishing models. These findings suggest that dissatisfaction with commercial publishing models is growing, and that community-led alternatives can motivate change. 2026-06-29T14:45:48Z 17 pages, 6 figures, submission to STI-ENID 2026 Saskia van Walsum Lisa Matthias Juan Pablo Alperin Stefanie Haustein http://arxiv.org/abs/2606.30346v1 Detector-Output Instability near the Kesten-Stigum Boundary: Separating Hard Readout, Relaxation, and Fixed-Point Dispersion 2026-06-29T14:21:07Z Community-detection algorithms usually return a single partition, even when independent initializations or small data perturbations yield several plausible outputs. We probe this output distribution through three paired observables: hard-partition variation of information (VI), a residual-gated fixed-point VI, and a cutoff-free Jensen-Shannon distance between belief-propagation (BP) marginal fields. For the symmetric sparse stochastic block model, linearizing BP around the uninformative fixed point gives the Kesten-Stigum onset at $\mathrm{snr}=(c_{\rm in}-c_{\rm out})/(q\sqrt{c})=1$. The hard VI maximum is instead a finite-size, readout-dependent detector curve on the detectable side, typically $\mathrm{snr}^\star \simeq 1.05\text{-}1.10$; moving the polarization cutoff from 0.001 to 0.1 shifts it across 1.047-1.128. The nontrivial-readout activation obeys $\mathrm{snr}_{50}(τ)-1 = 0.0086 + 0.522\,τ$ ($R^2=0.996$). Long-budget residual gating separates readout and critical slowing from fixed-point dispersion: at $\mathrm{snr}=1.05$ and 1.10 the hard VI is 1.49 and 1.58 bits but the gated subsets have zero VI, whereas from 1.15 to 1.30 nearly all runs pass the gate and retain VI 1.31 down to 1.24 bits. A high-replication audit through $N=100000$ disfavors a zero-asymptote power law and finds a small plateau $\mathrm{snr}^\star-1 \simeq 0.024$ (graph-bootstrap 90% interval [0.0227, 0.0316]). On real networks, a label-free Bethe-Hessian modularity margin with a Chung-Lu null gate is run on political blogs and six SNAP graphs: the measurement stays label-free, while heterogeneous networks can retain null-significant structure even after strong edge subsampling. The result is a detector-output decomposition near the Kesten-Stigum boundary, reporting hard readout, relaxation dynamics, and fixed-point-field dispersion separately. 2026-06-29T14:21:07Z 16 pages, 5 figures. Ancillary files include all source code and the raw numerical data behind every figure Faruk Alpay Baris Basaran http://arxiv.org/abs/2509.02098v7 Maximum entropy temporal networks 2026-06-29T12:01:08Z Temporal networks consist of timestamped directed interactions that may appear continuously in time, yet few studies have directly tackled the continuous-time modeling of networks. Here, we introduce a maximum-entropy approach to temporal networks and with basic assumptions on constraints, the corresponding network ensembles admit a modular and interpretable representation: a set of global time processes and a static maximum-entropy edge, e.g. node pair, probability. This time-edge labels factorization yields closed-form log-likelihoods, degree, clustering and motif expectations, and yields a whole class of effective generative models. We provide the maximum-entropy derivation for the non-homogeneous Poisson Process (NHPP) intensities governing the probability of directed edges in temporal networks via the functional optimization over path entropy, connecting NHPP modeling to maximum-entropy network ensembles. NHPPs consistently improve log-likelihood over generic Poisson processes, while the maximum-entropy edge labels recover strength constraints and reproduce expected unique-degree curves. We discuss the limitations of this framework and how it can be integrated with multivariate Hawkes calibration procedures, renewal theory, and neural kernel estimation in graph neural networks. 2025-09-02T08:54:10Z 21 pages, 32 figures Paolo Barucca 10.1103/78vv-hs72 http://arxiv.org/abs/2606.30142v1 Minimizing cumulative infections in SIS epidemic models over networks via an edge deletion algorithm 2026-06-29T11:21:32Z In this paper, we investigate the discrete SIS (Susceptible-Infected-Susceptible) models. We focus on minimizing epidemic spreading over networks by extending an existing edge deletion algorithm to the SIS model. To achieve this, we employ the mean-field approximation to linearize the network dynamics into a deterministic SIS model. We analytically demonstrate that the total number of infections is upper-bounded by a super-modular function, thereby ensuring the efficiency of the edge-deletion approach. To evaluate the proposed method, we conduct experiments on synthetic Erdos-Renyi networks and the real-world dataset collected from BBC Pandemic Haslemere app. Numerical simulations validate our theoretical results, confirming that both configurations converge to the stable, disease-free equilibrium. 2026-06-29T11:21:32Z Phi Dung Hoang Khanh Ly Duong http://arxiv.org/abs/2606.30069v1 Phase Boundary of a Stochastic Watts-Threshold SIS Model on Random Networks 2026-06-29T10:01:42Z Complex contagion models, in which adoption requires reinforcement from multiple neighbors, have been extensively studied in the monotone (no-recovery) setting, but the phase diagram of threshold models with SIS-like recovery on networks remains unmapped. We study a stochastic Watts-threshold SIS model on Erdos-Renyi and Barabasi-Albert networks and reconstruct its extinction-persistence phase boundary in the joint parameter space of transmission rate $β$, adoption threshold $θ$, and infectious duration $d$. Using adaptive Delaunay-based sampling and weighted logistic regression on over 180,000 Monte Carlo trials, we find that: (i) the boundary is well described by a six-parameter interaction model whose structure is invariant across both topologies; (ii) the transition is sharp, with the 10-90\% extinction-probability band spanning only $Δθ\approx 0.005$-$0.008$; and (iii) the adoption threshold is the dominant parameter governing epidemic feasibility, with transmission rate and infectious duration playing secondary and asymmetric roles. The characterization provides a quantitative reference for the complex-contagion analogue of the classical SIS epidemic threshold. 2026-06-29T10:01:42Z Yasmine Beji Heger Arfaoui Slimane BenMiled http://arxiv.org/abs/2606.29826v1 Rethinking Collaborative Trust for Verifiably Decentralized Blockchain Systems 2026-06-29T06:12:52Z Despite the promise of decentralization, measurement studies have identified a conspicuous lack of decentralization in blockchains. Centralization has been observed in almost all layers of the blockchain, in decentralized applications, and in decentralized autonomous organizations. In many cases, it is practically impossible to definitively determine the extent of centralization in the system. While multiple works have proposed methods to decrease centralization, by and large blockchains continue to be significantly centralized. In this paper, we develop a general framework for building verifiably decentralized blockchain systems. Our framework is motivated by the core observation that the richness and diversity of collaborative interactions between users -- rather than resource uniformity -- captures the essence and extent of decentralization in a blockchain system. Existing blockchains do not have any incentive mechanisms to encourage inter-coalition collaboration, which directly contributes to centralization. We propose a novel reward design that incentivizes users to collaborate with other users without forming isolated coalitions. Technically, our method uses a Sybil-resistant asymmetric Shapley value for reward attribution within a collaboration group, and the theory of expander graphs for measuring and enforcing decentralization. Our framework is general and can be adapted to alleviate centralization in any layer, application, or decentralized organization. It also has important implications beyond the topic of centralization. For example, we show that our solution can naturally address the blockchain scalability problem. We also identify a new class of decentralized collaborative applications that have hitherto been unexplored in blockchains. 2026-06-29T06:12:52Z Yunqi Zhang Shaileshh Bojja Venkatakrishnan http://arxiv.org/abs/2606.29777v1 The Longevity of Innovation 2026-06-29T04:43:22Z Modern science is organized around specialization in training and teamwork. Scientists develop deep expertise within a field and combine complementary knowledge through collaboration to solve complex problems. Yet whether specialization is the most effective path to sustained innovation remains unclear. Here we introduce a quantitative framework that distinguishes generalists from specialists based on scaling patterns of disciplinary mobility while remaining independent of career age and productivity. Applying this framework to 49 million publications produced by 3 million scientists between 1900 and 2020, we examine how research style relates to innovation, learning, collaboration, and productivity. We find that scientists who move across fields are more likely to sustain innovative contributions throughout their careers, whereas those who remain within narrow fields exhibit the age-related decline in innovation. Generalists are less anchored to the literature of their training. They are more likely to pursue research independently, and, when they collaborate, they preferentially partner with other generalists. Teams with a greater share of generalists produce more innovative research, even after accounting for differences in knowledge diversity. Despite these advantages, generalists publish fewer papers on average and have become less common over time. These findings reveal a tension between the longevity of scientific careers and the longevity of scientific innovation. 2026-06-29T04:43:22Z Yiling Lin Zak Risha Erin Leahey Lingfei Wu http://arxiv.org/abs/2606.29722v1 Attraction, Not Adaptation: How AI Agent Communities Develop Distinct Linguistic Identities 2026-06-29T02:59:50Z When tens of thousands of autonomous AI agents interact in topical online forums, do they develop distinct community-specific linguistic identities? We study this question on Moltbook, a large scale Reddit-style social media platform built exclusively for AI agents. Using the public Moltbook Observatory Archive dataset with over 3.1 million posts and 1.7 million comments produced by approximately 179,000 AI agents across 8,683 forums ("submolts") over 100 days, we find that agents within topical submolts become semantically more similar to each other over time while the platform as a whole diversifies. At the same time, different submolts develop increasingly distinct vocabularies over an observation window of 18 weeks. Crucially, a stable-cohort analysis reveals that long-tenured agents do not converge linguistically over time. Instead, community-level linguistic differentiation operates through selective attraction - newcomers arrive already linguistically compatible with their chosen community - and differential retention - conforming agents remain active longer. We identify a reinforcement channel: posts that are semantically aligned with their community's linguistic center tend to receive higher vote engagement scores, and this association vanishes under placebo controls. Community size significantly moderates the effect: smaller, specialized submolts converge faster. Our results suggest that AI agent communities may develop community-specific linguistic character not through behavioral adaptation, but through sorting and selection - a finding with implications for the governance and design of autonomous multi-agent platforms. 2026-06-29T02:59:50Z 14 pages, 11 figures Daming Li Simeng Han Can Meng Wanyu Lei Jialu Zhang http://arxiv.org/abs/2506.12078v2 Modeling Earth-Scale Human-Like Societies with One Billion Agents 2026-06-28T16:58:52Z Understanding the dynamic evolution of complex social phenomena requires both high-fidelity modeling of human behavior and large-scale simulations. Traditional agent-based models (ABMs) have been employed to study these dynamics, but are constrained by simplified agent behaviors. Recent advances in large language models (LLMs) enable agents to exhibit sophisticated social behaviors, yet face significant scaling challenges. We present Light Society, an agent-based simulation framework that advances both fronts. Light Society formalizes social processes as structured transitions of agent and environment states, governed by a set of LLM-powered simulation operations. Joint algorithmic and system optimizations, particularly a mixture-of-models engine that combines full LLMs with distilled surrogates, enable Light Society to efficiently simulate societies with over one billion agents. Grounded in real-world demographic profiles from the World Values Survey, simulations of Trust Games and opinion diffusion at up to one billion agents demonstrate Light Society's high fidelity and efficiency in modeling diverse social phenomena, providing researchers with a practical foundation for hypothesis testing and the study of emergent collective behaviors at planetary scale. 2025-06-07T09:14:12Z Haoxiang Guan Jiyan He Liyang Fan Zhenzhen Ren Shaobin He Xin Yu Yuan Chen Xueyin Xu Shuxin Zheng Yan Gao Enhong Chen Tie-Yan Liu Zhen Liu http://arxiv.org/abs/2606.29488v1 Should children follow their parents' research paths? Intergenerational research continuity and divergence in academic families 2026-06-28T16:39:06Z How academic advantages are transmitted within families is usually studied as occupational inheritance, but it is not clear whether scholarly research orientations persist across generations and if it is an advantage when it does. To address this, we link Wikidata kinship records with OpenAlex bibliometric profiles to study 3,229 documented parent-child scholar pairs and 488,659 publications. Field-level research similarity was evident but not universal: whilst the median similarity was 0.546, 25.3% of parent-child pairs had no Field overlap (i.e., similarity 0). These pairs were substantially more similar than publication-period-matched comparison pairs (median 0.098). Direct academic interaction was uncommon: 10.4% of parent-child pairs had co-authored, 9.8% of children had cited their parents, and 6.9% of parents had cited their children. Nevertheless, each 0.1 increase in Field similarity was associated with 38-39% higher adjusted odds of co-authorship and cross-citation. There was also intergenerational continuity in academic achievement and recognition. Parents' publication volume and field-normalized citation impact were positively associated with those of their children. Children of national academy members had approximately twice the odds of becoming national academy members themselves (Odds Ratio = 2.04), while children of prizewinning parents had 46% higher odds of winning prizes (Odds Ratio = 1.46). However, children of national academy members showed lower research similarity to their parents. Greater research differentiation was associated with higher field-normalized citation impact among children, but not with publication output or higher odds of academic recognition. Academic families therefore appear to transmit resources and advantages with the sole exception that diverging from parental fields seems to confer a citation advantage. 2026-06-28T16:39:06Z Er-Te Zheng Xiaorui Jiang Zhichao Fang Mike Thelwall http://arxiv.org/abs/2606.29240v1 Blackknife: Hard-Label Query-Limited Black-Box Attacks on Heterogeneous Graph Neural Networks 2026-06-28T07:11:05Z Heterogeneous graph neural networks (HGNNs) have achieved strong performance in modeling complex graph-structured data with multiple node and relation types. However, their robustness under realistic black-box adversarial settings remains insufficiently explored. Existing attacks on HGNNs usually assume access to model gradients, soft prediction scores, or the complete graph structure, which is often unavailable when HGNN-based services are deployed as closed systems. In this paper, we propose Blackknife, a hard-label, query-limited, and structure-limited black-box evasion attack framework for heterogeneous graph neural networks. Blackknife assumes no access to the victim model architecture, parameters, gradients, logits, confidence scores, or the full graph structure. Instead, it only relies on locally observable one-hop heterogeneous structures and a small number of hard-label queries. To generate effective perturbations under these strict constraints, Blackknife first constructs a local relation-aware surrogate model from observable heterogeneous neighborhoods. It then relaxes discrete edge addition and deletion operations into continuous soft weights and optimizes them through projected gradient descent. Finally, the optimized perturbations are discretized into relation-preserving structural rewiring operations and verified using limited hard-label feedback from the victim model. Extensive experiments on three benchmark heterogeneous graph datasets, including ACM, DBLP, and IMDB, demonstrate that Blackknife consistently achieves strong attack success rates against representative HGNN models. The results further show that Blackknife remains effective under topology-based defense strategies, revealing the vulnerability of HGNNs to local structure-limited black-box attacks. 2026-06-28T07:11:05Z Honglin Gao Junhao Ren Lan Zhao Yue Yang Jindong Chang Gaoxi Xiao http://arxiv.org/abs/2504.07480v2 Quantifying and Mitigating Consensus Disparity in Social and Information Networks 2026-06-27T22:34:25Z We introduce a computational framework to measure and optimize disparity, which corresponds to the difference in consensus outcomes attributable to distinct social groups, under classical models of opinion dynamics. We study this problem in the Friedkin-Johnsen setting under uncertainty about group structure and characterize its algorithmic complexity. For the structural analysis, we demonstrate that disparity can be arbitrarily larger than polarization in well-connected networks that nonetheless carry an identifiable group structure. For the mitigation problem, we derive robust formulations and active set optimization procedures to minimize worst-case disparity via recommendation reweighing and opinion seeding. Our methods provide provable guarantees and are validated on multiple real-world social networks. The results bridge opinion dynamics and network optimization, offering computational tools for analyzing and reducing polarization in social networks. 2025-04-10T06:18:27Z Marios Papachristou Jon Kleinberg http://arxiv.org/abs/2606.29098v1 Connectivity Estimation using Stochastic Graph Heat Modelling 2026-06-27T21:55:02Z A growing number of techniques leverage the spatial structures that underlie many real-world datasets. Despite these advances, the complementary task of estimating spatial structures and understanding their role within these techniques has often been overlooked. In neurophysiological data analysis specifically, numerous methods exist to estimate brain connectivity, but most are not explicitly model-based, dynamic, multivariate, or directed. To address these limitations, we previously introduced noise-driven heat modelling on graphs for neurophysiological connectivity estimation. In this study, we extend this framework by relaxing earlier noise assumptions and adding regularisation to improve robustness. We also develop a simulation procedure to characterise and evaluate our technique in a controlled setting. Finally, we demonstrate that the technique is able to capture meaningful spatial structure across two experiments, each using two real-world datasets. The explicit model formulation of our connectivity estimator has the potential to improve the interpretability of graph-based techniques across a wide range of applications. The code implementing our method is available at https://github.com/sgoerttler/Heat_Connectivity. 2026-06-27T21:55:02Z 14 pages, 11 figures. Includes supplemental material Stephan Goerttler Min Wu Fei He