https://arxiv.org/api/Ts6crr6+axV4c6HA1nlcie8Dhj42026-09-11T17:47:30Z23900015http://arxiv.org/abs/2609.11736v1Learning structural balance of graphs from quantum spectral features2026-09-10T15:52:28ZWe develop a quantum approach to spectral feature extraction from the density of states (DOS) of a problem-dependent Hamiltonian, and apply it to machine learning on signed graphs. We propose to embed a signed graph as an Ising model instance with positive and negative interactions, and use the standardized moments of the Ising DOS as features for learning. We show that these moments count signed closed walks, are switching-invariant, and are size-free by construction. As a benchmark, we target learning the frustration index, an NP-hard measure of structural balance that can be labeled exactly at moderate size. At zero field, the models can be sampled classically, allowing the quantum extraction procedure to be certified against exact ground truth. We propose DOS-QPE, a phase estimation on a purified maximally mixed probe, which samples the spectral density with orders of magnitude fewer shots than Hadamard test-based trace sampling and feeds the resulting features directly into classically trained models. On $1.4\times10^5$ labeled graphs the exact DOS determines the frustration index, and five moments recover it with a mean error of 0.4, well below one sign flip. Beyond zero field, the underlying trace-estimation problem is DQC1-complete, providing access to spectral features for which no efficient classical sampling method is known. Our work opens routes towards quantum applications in social network balance analysis, spin-glass studies, correlation clustering, and protein-interaction networks.2026-09-10T15:52:28Z12 pages, 6 figuresStefano ScaliOleksandr Kyriienkohttp://arxiv.org/abs/2609.02990v2Toward Collective-Centric Evaluation of Preference Inference for Participatory Democracy2026-09-10T11:28:12ZTo scale up collective decision-making, participatory democracy platforms such as Polis and Remesh enable online deliberation among thousands of participants. However, at this scale, participants cannot review every opinion submitted by others, producing highly sparse voting data that misrepresent patterns of consensus, conflict, and minority support. Platforms therefore increasingly rely on Preference Inference (PI) models to predict missing votes. Yet this automation is not neutral: inferred preferences can artificially amplify, suppress, or reorder existing patterns of support, ultimately reshaping how the outcomes of a deliberation are interpreted. More generally, we lack a systematic understanding of how existing PI methods affect the collective preference landscape. To address this gap, we benchmark several existing PI approaches in this context. Moving beyond conventional user-centric evaluations centered on the accuracy of individual predictions, we introduce a collective-centric evaluation framework that measures whether inferred votes preserve salient properties of the broader preference landscape. We further contribute the largest multilingual dataset of its kind: four consultations spanning over 90k participants, 1M votes, and 22 languages. Our experiments show that models with comparable predictive accuracy can differ substantially in the degree to which they preserve the collective structure. These results demonstrate that accuracy alone is insufficient for evaluating PI in democratic settings. By contributing a novel comprehensive and collective-centric evaluation benchmark for the task of PI, this work aims to support the development of AI systems that scale deliberation without compromising the integrity of its democratic outcomes.2026-09-02T15:23:00Z7 pages of contentPierre-Antoine LequeuSalim HafidPaul LernerNazanin ShafiabadiLaurène CaveDavid MasJean-Philippe CointetBenjamin PiwowarskiFrançois Yvonhttp://arxiv.org/abs/2609.11320v1A global mobile network coverage raster product at 1km resolution, 1999--20302026-09-10T09:48:42ZWhere a mobile signal is available shapes who can work, learn, bank, seek health care and respond to crises in the digital age, yet no globally consistent, sub-national record of mobile network coverage exists. We present such a record: annual 1km maps of the probability of 2G, 3G and 4G coverage for 214 countries and territories for the years 1999 to 2030. The maps are produced by three independent models: a calibrated machine-learning model, a techno-economic simulator of network build-out, and a spatial deep-learning model. The three estimates are then combined, per country and technology and in proportion to their measured accuracy, into a single best estimate with per-pixel 90% uncertainty bands; all four layers are released as part of the dataset. Because mobile roll-out closely follows a country's socio-economic conditions (population distribution, electrification, physical infrastructure), the models are grounded in existing geospatial data and tuned on 2,409 quality-screened operator-reported coverage maps, which are available up to 2020. For 2021--2024 the maps are predicted from recent geospatial data alone; for 2025--2030 they are extrapolated from demographic and infrastructure projections. On countries held out during training, the machine-learning model attains AUC 0.89--0.92. Baseline comparisons and the combined product's external validation are reported in Technical Validation. The dataset supports mapping the global digital divide, linking connectivity to household-survey outcomes, and humanitarian and infrastructure planning.2026-09-10T09:48:42ZDataset linked to that paper: https://zenodo.org/records/21594337Till KoebeTheophilus AidooAli El ChamiAli KansoAkansh MauryaPurushottam SharmaIngmar WeberRidhi Kashyaphttp://arxiv.org/abs/2609.11202v1Automated Identification of Competing Narratives in Political Discourse on Social Media2026-09-10T08:10:33ZSocial media platforms have become central to shaping political discourse, serving as arenas where narratives form and evolve, influencing public opinion. Identifying and analyzing these narratives, particularly when they compete across different political ideologies, is crucial for understanding the dynamics of modern political communication. This paper presents an unsupervised framework for identifying and characterizing competing narratives in political discourse on social media, focusing on German politicians' tweets. The framework employs a multi-stage pipeline that integrates natural language processing techniques such as topic modeling, event detection, and event linking. By forming data into coherent stories and uncovering the distinct perspectives of user communities, the system is able to detect the key competing narratives, highlighting the divergent framings and conflicts surrounding trending political topics. Two case studies on polarizing political issues demonstrate the efficacy of the methodology, showcasing its ability to uncover and analyze divergent viewpoints. The findings contribute to the broader understanding of how narratives propagate within the digital public sphere and offer insights for policymakers, social media platforms, and researchers interested in monitoring political discourse.2026-09-10T08:10:33Z11 pages, 5 figures. Published in the proceedings of Text2Story 2025, held with ECIR 2025In: Proceedings of Text2Story - Eighth Workshop on Narrative Extraction From Texts (Text2Story 2025), CEUR Workshop Proceedings, Vol. 3964, 2025, pp. 137-147Sergej WildemannErick Elejaldehttp://arxiv.org/abs/2609.11144v1Same Day, Same Story; One Day Ahead, a Different Signal: The Dual Validity of Financial Sentiment2026-09-10T06:37:33ZFinancial NLP has a standard workflow: validate a sentiment tool against human labels, then trust it to extract market signal. This assumes the two evaluations measure the same thing. We test that assumption in a setting where both can be measured at once: a corpus of securities class actions (2002-2025) linking 70,500 X messages to abnormal stock returns, with a single-annotator human labelled gold sample. Running five instruments (VADER, Loughran-McDonald, FinBERT, Twitter-RoBERTa, and an LLM annotator) through one identical pipeline, we find that the relationship between construct and predictive validity depends on the sampling convention and score representation. Under conventional method-specific sampling, human agreement aligns more closely with graded same-day associations than with one-day leads. On a fixed-n panel, however, agreement has similar graded rank correlations at both horizons, while the coarse ordering remains weak. Benchmark agreement therefore establishes semantic validity but does not by itself determine predictive rankings. In a conversation that is 17.6% spam, message volume predicts neither market damage nor settlement size.2026-09-10T06:37:33ZAS AravinthkakshanLaven SrivastavaHarsh Nandwanihttp://arxiv.org/abs/2609.10742v1SynCo: Synthetic Community-Aware Attributed Graph Generator for Graph Neural Network Benchmarking2026-09-09T18:38:48ZGraph Neural Networks (GNNs) are powerful models for handling attributed graphs in tasks such as classification, link prediction, and community detection, as they enable the aggregation of information from both structural and semantic sources. However, progress in community detection is hindered by the lack of high-quality datasets, since ground-truth community labels are often unavailable and most algorithms proposed in recent literature rely on the same benchmark datasets for model training and evaluation. To address this issue, attributed random graph generators are commonly employed to create synthetic graphs for assessing the strengths and limitations of GNN-based models. Nevertheless, most existing generators rely heavily on power-law degree distributions, despite recent evidence indicating that scale-free networks are rare, particularly in social network contexts. Moreover, state-of-the-art attributed graph generators provide limited flexibility, as they do not allow users to construct communities with varying densities, degree distributions, and sub-community structures. To overcome these limitations, we introduce the Synthetic Community-Aware Attributed Graph Generator (SynCo), a graph generation algorithm that allows users to control the node degree distribution and sub-community structure. We evaluate SynCo across three different tasks: graph mimicking, hyperparameter evaluation, and node clustering tuning. The results show that our model outperforms state-of-the-art approaches in synthetic graph generation and data augmentation, while preserving the original distributions of duplicated and augmented datasets, as confirmed by statistical tests well know in literature. We also demonstrate the ability of SynCo to generate nodes in large scale, up to 2.1 million nodes.2026-09-09T18:38:48Z15 pagesGuilherme Henrique MessiasMariana Caravanti de SouzaSylvia IasulaitisAlan Demétrius Baria Valejohttp://arxiv.org/abs/2609.10399v1Fundamental limits to identifying node and tie memory in temporal networks: marginal artefacts and spreading dynamics2026-09-09T16:20:09ZTemporal-network models attribute memory in contact data to either node self-excitation (branching ratio n_node) or tie reinforcement (kappa), carrying major consequences for epidemic spreading. We prove that when event initiators are observed, the two mechanisms are orthogonal: the Fisher information is block-diagonal and neither trades off against the other. In undirected proximity data, where initiators are unobserved, marginalising over them couples the mechanisms into a structural confound that survives posterior smoothing. On empirical proximity, messaging, and email records, however, a cruder failure dominates: fitted node memory is pinned to the inter-event marginal law and remains virtually invariant across latent label posterior samples (coefficient of variation below 1%). An inter-event-order shuffle test and burstiness-memory diagnostics reveal that exponential-Hawkes node memory is recovered from none, while tie reinforcement remains identifiable throughout. This near-unidentifiability is intrinsic, not an artefact of the exponential kernel: refitting flexible scale-free (sum-of-exponentials) kernels on synthetic power-law self-exciting processes fails to distinguish genuine node memory from memoryless renewal controls, with identical collapses recurring on algorithmic networks (edit bots, cloud microservices) and cortical spiking. Downstream epidemic consequences are quantitative: simulations fitted to empirical contact records under-predict outbreak sizes by up to a factor of 2.5 and shift the epidemic threshold. We conclude that observational temporal networks face a two-fold identifiability boundary: contact directionality is essential to decouple tie reinforcement, whereas heavy-tailed node self-excitation is intrinsically unidentifiable from contact timings alone.2026-09-09T16:20:09Z25 pages, 7 figures; 33 pages Supplementary MaterialMichele Tizzanihttp://arxiv.org/abs/2609.10277v1How neighbourhood ideology shapes misinformation belief in densely tied social networks2026-09-09T14:58:41ZWith the rapid spread of news on social media, understanding the propagation of misinformation is becoming increasingly important. One factor that affects individuals' vulnerability to false information is their ideological predisposition. Despite the large number of agent-based models that focus on social influence as a driver of the spread of false claims, they often fail to explicitly integrate personal ideological biases into belief formation. In this work, we explore how misinformation spreads through the interaction between individuals' ideological biases and social influence. Our model accounts for both the strength of individuals' ideological biases and the extent to which a false claim aligns with their ideology. Social influence modifies the effects of ideological intensity and false claim alignment through network interactions. Notably, the influence of neighbours' ideological intensity on belief is strongly affected by how well those neighbours are connected to one another. These results highlight the importance of considering both network structure and personal ideological biases when modelling misinformation propagation.2026-09-09T14:58:41Z18 pages, 9 figures, 1 tableSoroush KarimiMarcos OliveiraDiogo Pachecohttp://arxiv.org/abs/2609.05442v2Role differentiation as ignition of a collective information engine: Structuration in Agent Populations2026-09-09T13:14:27ZInformational active matter shows how measurement-informed decisions produce collective order, so far in systems that reach consensus. We design collective information engines structured by differentiation instead, and construct a minimal instance using anti-coordination games where differentiated role information has value. Within many coexisting games, agents infer their role from a noisy social signal grounded in a persistent identity, and role-following action feeds back into that signal, which shapes the incentive to follow roles. Resources accrued through coordinated role-play combine with identity variability to reinforce the schemas that generated them. The model thereby operationalizes Sewell's duality of schemas and resources in Structuration, a resolution to structure--agency debates across social science. The engine ignites when a social loop gain---the product of identity persistence, cognitive capacity, channel fidelity, and schema strength---exceeds one. For a repertoire of such schemas, roles emerge with increasing gain in a bifurcation cascade whose functional form is fixed by the repertoire's eigenvalue spectrum, ranging from monitorable logarithmic sequences to avalanches that arrive without warning. Resource accumulation supplies the fitness of a replicator dynamics on schema strengths, which selects the cascade type endogenously. Subcritical identity covariance reveals that type before onset, enabling early detection, while feedback channel parameters bias which type is selected. Platform design then becomes a control lever to throttle emergent coordination. This theory grounds distributional AGI takeoff in a mechanism and provides a monitor-based solution. Joining game theory, collective dynamics, and information engines, we open a route to an information thermodynamics of agent populations.2026-07-26T19:13:17ZMaximilian Puelma Touzelhttp://arxiv.org/abs/2603.05367v2Shock Propagation and Macroeconomic Fluctuations2026-09-09T12:53:28ZWe study how idiosyncratic firm-level shocks generate aggregate volatility and tail risk when they propagate through a production network under overlapping adjustment: new productivity draws arrive before the economy reaches the static equilibrium associated with earlier draws. Each innovation generates a `productivity wave' that mixes and dissipates over time as it travels through the production network. Macroeconomic fluctuations emerge from the interference between these waves of different vintages. The interference between these waves is governed by the dominant transient eigenvalue of the production network, and therefore so are the macroeconomic fluctuations they generate. In such a dynamic regime, the tail of the degree distribution is a markedly weaker determinant of macro fluctuations than in the fully adjusted static benchmark. And the macroeconomic significance of the degree-heterogeneity of production networks cannot be known without knowing the rate at which the economy converges to equilibrium or equivalently the spectral properties of the production network. More concretely, once we permit the time-averaging of shocks, granular shocks may account for only a small fraction of the empirically observed aggregate volatility.2026-03-05T16:44:54ZAntoine MandelVipin P. Veetilhttp://arxiv.org/abs/2512.02362v4Reconstructing Large Scale Production Networks2026-09-09T07:03:00ZFirm-to-firm production networks matter for aggregate propagation, but they are rarely observed. This paper reconstructs national-scale, weighted firm-to-firm networks from two public objects: a sectoral input--output table and the distribution of firm sizes by sector. The algorithm first draws a binary buyer-seller backbone from a sector-aware gravity model and then assigns weights by a minimum-energy program. A Markov closure makes the reconstructed network primitive, so it has a unique stationary distribution. The weighting program keeps one-step firm balances and sectoral flows close to the data; the stationary money vector is then checked ex post and remains close in aggregate. For the United States we reconstruct a network with about 6.5 million firms and 340 million links in roughly four hours on a single workstation. We also reconstruct the networks of Japan, the United Kingdom, Australia, Finland, and Denmark. The Japanese reconstruction, built without any link data, reproduces the heavy-tailed degree regime documented in the country's observed production network. The reconstructed networks exhibit customer tails heavier than supplier tails, though the algorithm treats the two sides symmetrically. We also run computational experiments on the reconstructed networks to assess the systemic risk posed by the failure of individual firms. These experiments show that neither firm size nor degree nor sectoral position is a good proxy for the aggregate losses generated by a firm's failure. For such questions, there is no good substitute for the complete weighted buyer-seller network that we reconstruct. We release the reconstruction code, the generated networks, a Python library, and a graphical2025-12-02T03:12:12ZAshwin BhattathiripadVipin P Veetilhttp://arxiv.org/abs/2405.00636v4Robustness of shallow graph embedding methods for community detection2026-09-09T05:37:55ZThis study investigates the robustness of shallow graph embedding methods for community detection in the face of network perturbations, specifically node deletions. Graph embedding techniques, which represent nodes as low-dimensional vectors, are widely used for various graph machine learning tasks due to their ability to capture structural properties of networks effectively. However, the impact of perturbations on the performance of these methods remains relatively understudied. The research considers state-of-the-art shallow graph embedding methods from two families: matrix factorization (e.g., LE, LLE, HOPE, M-NMF) and random walk-based (e.g., DeepWalk, LINE, node2vec). Through experiments conducted on both synthetic and real-world networks, the study reveals varying degrees of robustness within each family of shallow graph embedding methods. The robustness is found to be influenced by factors such as network size, initial community partition strength, and the type of perturbation. Notably, node2vec and LLE consistently demonstrate higher robustness for community detection across different scenarios, including networks with degree and community size heterogeneity. These findings highlight the importance of selecting an appropriate shallow graph embedding method based on the specific characteristics of the network and the task at hand, particularly in scenarios where robustness to perturbations is crucial.2024-05-01T17:04:20ZAccepted manuscript. Published in Applied Network ScienceAppl Netw Sci (2026)Zhi-Feng WeiPablo MorianoRamakrishnan Kannan10.1007/s41109-026-00825-zhttp://arxiv.org/abs/2602.15964v2Approximate Pareto Frontiers for Submodular Utility and Cost Tradeoffs2026-09-09T05:11:23ZIn many data-mining applications, including recommender systems, influence maximization, and team formation, the goal is to pick a subset of elements (e.g., items, nodes in a network, experts to perform a task) to maximize a monotone submodular utility function while simultaneously minimizing a cost function. Classical formulations model this tradeoff via cardinality or knapsack constraints, or by combining utility and cost into a single weighted objective. However, such approaches require committing to a specific tradeoff in advance and return only a single solution, offering limited insight into the space of viable utility-cost tradeoffs.
In this paper, we depart from the single-solution paradigm and examine the problem of computing representative sets of high-quality solutions that expose different tradeoffs between submodular utility and cost. For this, we introduce $(α_1,α_2)$-approximate Pareto frontiers that provably approximate the achievable tradeoffs between submodular utility and cost. Specifically, we formalize the Pareto-$\langle f,c \rangle$ problem and develop efficient algorithms for multiple instantiations arising from different combinations of submodular utility $f$ and cost functions $c$. We also provide an adaptive search algorithm that computes only a small subset of points that collectively summarize the entire Pareto frontier.
Our results offer a principled and practical framework for understanding and exploiting utility-cost tradeoffs in submodular optimization. Experiments on datasets from diverse application domains demonstrate that our algorithms efficiently compute approximate Pareto frontiers in practice.2026-02-17T19:28:55Z11 pages35th ACM International Conference on Information and Knowledge Management, CIKM 2026Karan VombatkereEvimaria Terzi10.1145/3799682.3841064http://arxiv.org/abs/2609.09687v1Chance, Persistent Advantage, and the Generative-AI Era in Open-Source Package Careers2026-09-09T04:03:18ZStudies of careers in science, film, music, and books report a common pattern. When a person's most successful work arrives is close to a random draw over the works they produce. How large their successes tend to be, in contrast, follows a stable, person-specific factor. We test whether this pattern holds for open-source software careers and whether it changed when generative AI coding tools arrived. From the complete public record of GitHub push events (2015-2025), we reconstruct 102.2M career works by 6.15M contributors, and for the 908k contributors whose repositories publish packages, we measure each work's impact by how many downstream packages come to depend on it. First, we find that the timing of a career's biggest hit is close to a lottery over their works, as in science and the arts, with a small, replicable lean toward early career that grows as careers get longer. Second, some coders reliably produce higher-impact work than others, but this lasting personal factor accounts for only part of why impact persists (about a fifth in our primary specification); the rest behaves like momentum, success feeding on itself for a period of time. Third, within the same contributors, this structure did not change after ChatGPT's release. The stable factor's weight grew by about as much as it grew for an earlier cohort that simply aged, and subtracting the effect of aging from the effect of generative AI puts the shift at +0.03 (95% CI [-0.22, +0.23]), indistinguishable from zero. The success pattern documented in science and the arts therefore describes open-source careers too, and it shows no detectable break across the arrival of generative AI. These results have implications for how track records on open platforms should be read and on what to expect from generative AI for the careers built on them.2026-09-09T04:03:18Z25 pages, 11 figuresHazem IbrahimYasir Zakihttp://arxiv.org/abs/2609.09658v1Link prediction in complex networks via fusing node centrality and local similarity indices2026-09-09T03:15:51ZLocal similarity indices are widely used in link prediction on complex networks owing to their low computational cost; however, in sparse networks they assign a zero score to every node pair lacking common neighbors, which severely limits their predictive power. A natural remedy is to fuse node centrality indices with local similarity indices: the former provide global importance for the node pair, while the latter capture fine-grained local topology, and the two can be combined into complementary scores within a unified framework. This paper uses PageRank and DomiRank as two representative centrality measures and constructs a centrality--local-similarity fusion framework. The PageRank-based fusion proposed by Charikhi is first generalized to seven classical local similarity indices, and the universality of its improvement is systematically verified on nine real-world network datasets. Furthermore, the DomiRank centrality is introduced to build the DR-MD series of fused indices under a unified weighting coefficient, which overcomes the drawback that the PageRank-based fusion requires index-by-index weight tuning. Results of five-fold cross-validation together with Wilcoxon signed-rank tests show that, under the unified experimental protocol, all DR-MD indices consistently outperform the corresponding local baselines and their PR-MD counterparts on all nine datasets ($p=0.002$), and that the improvements remain robust against perturbations of $σ$ and the weighting coefficients within the near-critical parameter plateau; in particular, DR-RA achieves an average AUC of 0.7084, surpassing global methods such as Katz and RWR as well as several advanced similarity indices. The framework is inherently extensible, and its fusion paradigm can be straightforwardly generalized to couple other node centrality indices with local similarity indices.2026-09-09T03:15:51Zlink prediction; complex networks; node centrality; local similarity; PageRank; DomiRankYingying ZhangChengye Zhao