https://arxiv.org/api/HdHS3UP/wDYevZu3/oxOCP6sjH0 2026-10-02T17:59:47Z 12468 150 15 http://arxiv.org/abs/2609.13802v1 Numerical Stability of Linear Algebra Operations over Relational Databases 2026-09-12T08:28:02Z A large body of work in the database literature develops efficient algorithms for linear algebra and machine learning over matrices defined by relational joins, yet the numerical stability of such computations has so far received no attention. This is a practical concern: a join matrix can be much larger than the input database, and the repeated copies of input values it contains compound the floating-point errors incurred by the numerical operations performed over it. This paper initiates a formal investigation of numerical stability for linear algebra over database joins. We first show that backward stability, the standard yardstick of numerical stability, loses its effectiveness in this setting: join matrices form a structured subspace of the ambient matrix space, so a perturbation explaining a computed result need not correspond to any perturbed input database. This failure already occurs for operations as simple as matrix-vector multiplication. To overcome this limitation, we introduce projected backward stability, a generalization of backward stability from the computation of one function to that of a composition of two functions, and establish its connection to classical backward stability. In our database setting, the two functions are the join query and the numerical operation. We further introduce the database condition number as the square root of the ratio of maximal to minimal number of copies of input data values into the join matrix, and show that it quantifies how a perturbation of the join matrix is amplified into a perturbation of the input database, independently of the computation used. The database condition number coincides with the classical condition number of the expansion matrix that replicates input values into the join matrix. 2026-09-12T08:28:02Z Andrei Draghici Yuchen He Dan Olteanu http://arxiv.org/abs/2609.13664v1 Cost Characterization of Vertically Partitioned Federated Knowledge Graphs 2026-09-12T02:47:14Z Knowledge graphs are increasingly distributed across autonomous organizations that share an entity space but own disjoint subsets of relations, forming a vertical partition. Answering a multi-hop query may require combining facts from several silos, making the partitioning strategy a key data management decision that affects communication, indexing, load balance, and query latency. However, the costs associated with different partitioning strategies remain insufficiently studied. We formalize vertical partitioning as a design space and compare four strategies: semantic domain grouping, frequency-balanced partitioning, co-occurrence graph-cut partitioning, and random partitioning. We evaluate them using five metrics: communication cost, candidate index size, cross-silo path length, load balance, and end-to-end query latency. Three of the five prove to be determined by the graph and the silo count rather than by the partition, which reduces the design problem to two conflicting axes, cross-silo path length and load balance. Experiments on MetaQA and PathQuestion use a fixed federated knowledge graph question-answering architecture based on TransE embeddings and a frozen BERT encoder across three silo configurations. By keeping the learning model unchanged, we isolate the effect of partitioning and show that the trade-off between locality and balance holds only where each silo can hold several relations, weakening as the number of silos increases. The study provides practical guidance for deployments constrained by cross-silo reasoning or by silo load. 2026-09-12T02:47:14Z Accepted at DMKG'26: 2nd International Workshop on Data Management for Knowledge Graphs, co-located with ISWC 2026; to appear in CEUR-WS proceedings Md Saikat Islam Khan Bappy Oshani Seneviratne http://arxiv.org/abs/2609.13452v1 COMPASS: Steering Distributed Vector Search with Scientific Knowledge Graphs 2026-09-11T19:13:09Z Vector databases use hashing to partition data across "shards," logical units for distributed execution. This placement, however, destroys semantic locality, forcing each query into scatter-gather limited by the slowest shard. Vector-space clustering can help, but scientific evidence is often connected by factual relations that do not align with embedding distance. We present COMPASS, a framework that uses a knowledge graph (KG) to determine data placement and query-time shard selection. COMPASS detects communities, splits oversized communities, inserts embeddings by subject entity, and routes queries to a small set of shards. Across four biomedical KGs, our method searches only 13-18% of the corpus while preserving broadcast recall and recovering up to 2.6x more multi-hop evidence than an embedding-based baseline. On 15 HPC nodes, COMPASS sustains 7.9x higher throughput with lower tail latency than hash-based broadcast. These results show that KG structure provides a compact complement to embedding geometry for scalable vector search. 2026-09-11T19:13:09Z Song Young Oh Amal Gueroudji Seth Ockerman Rob Latham Orcun Yildiz Ian Foster Kyle Chard Robert Ross http://arxiv.org/abs/2605.19246v4 Example-Driven Intent Synthesis for Constrained Data Bundle Retrieval: Focused Text Snippet Extraction and Beyond 2026-09-11T17:24:47Z Selecting a bundle of items that collectively satisfies constraints is a fundamental task across databases, recommender systems, and text summarization. Unlike traditional retrieval that returns individual or top-k items, bundle retrieval is inherently combinatorial and, in general, NP-hard. Although package queries can efficiently retrieve bundles given a well-formed query, two key user-centric challenges remain: (1) expressing and tuning multi-dimensional bundle intent through a user-friendly interface, and (2) ensuring feasibility when the query yields empty results. We introduce Ex2Bundle, an Example-driven Bundle retrieval framework that enables users to specify their intent through example bundles and automatically synthesizes package queries that capture the intent implicit in those example bundles via aggregate constraints. Ex2Bundle also addresses a challenge unique to bundle retrieval: when inferred aggregate constraints are infeasible over the target data, our data-aware constraint relaxation minimally adjusts the constraint bounds while preserving alignment with user intent. We instantiate a specific application of focused text snippet extraction by example to demonstrate the efficacy of the Ex2Bundle framework. Extensive experiments over real-world datasets and a user study demonstrate that Ex2Bundle improves usability and consistently returns intent-aligned bundles even under distributional shifts of the target database. 2026-05-19T01:40:08Z Whanhee Cho Kuangfei Long Mahmood Jasim Matteo Brucato Alexandra Meliou Peter J. Haas Anna Fariha http://arxiv.org/abs/2609.12745v1 How Do Data Collection Strategy and Data Quality Influence the Outcomes of Digital Technology Adoption? 2026-09-11T11:45:45Z In the era of Industry 4.0 (I4.0), data has become the essential foundation for digital transformation, yet many organizations still struggle to link data practices with digital performance outcomes. This study investigates how data collection strategy and data quality jointly influence the success of digital technology adoption (DTA) in manufacturing firms. Drawing on survey data from 86 firms, the research employs Partial Least Squares Structural Equation Modeling (PLS-SEM) to examine the relationships among data collection strategy, data quality, implementation performance, and operational performance. The results show that both data collection strategy and data quality significantly influence implementation and operational performance. However, the effect of data collection strategy on implementation performance is indirect, fully mediated by data quality. The study also finds that data collection strategy influences data quality. These findings demonstrate that data quality acts as a critical bridge between upstream data practices and downstream digital outcomes. The study contributes to the digital transformation and data management literature by empirically validating the central role of data quality and offering practical insights for managers to design and govern data processes strategically. It also sets a foundation for future research on data governance frameworks that integrate data quality assurance, standardization, and lifecycle management to sustain data-driven and digital transformation. 2026-09-11T11:45:45Z Xuejiao Li Cheng Yang http://arxiv.org/abs/2608.00501v5 Machine-Checked Dual-Write Recovery from a Commit Log 2026-09-11T09:16:18Z Applications often need to make related facts durable in two independent systems without a transaction spanning both. If the process crashes after the second system accepts an operation but before a source-side checkpoint is written, recovery cannot tell from source state alone whether to retry. Transactional outboxes and change data capture move this dual write out of an application process, but relay delivery and checkpointing remain separate durable operations. Systems address the problem with retries, checkpoints, idempotency keys, and fencing, and call the result exactly-once delivery. Whether that guarantee holds depends on which event it counts, what evidence recovery requires, and how long that evidence must survive. The closest formal studies model-check particular outbox and log-delivery designs, and their results hold only for the designs and instances they check. Answering the three questions in general requires statements about arbitrary recovery policies, which finite enumeration cannot reach. We prove them in Isabelle/HOL, and to our knowledge they have not been machine-checked before. The main result is an impossibility theorem for source-only recovery. We construct two reachable post-crash states with the same durable source-side state and different sink acceptance records. Any recovery policy based only on the source side must duplicate an effect in one state or leave it undelivered in the other. The same holds for a deterministic deliver-then-checkpoint protocol whose only nondeterminism is crash timing. An authoritative, complete, and current sink acceptance record lets recovery compute the missing operations when source coordinates distinguish them. We also prove arrival and claim fences for in-flight requests and concurrent recoverers. Finally, we show how bounded deduplication state and truncated source history limit the lifetime of the guarantee. 2026-08-01T07:52:16Z 23 pages, 5 figures. Machine-checked Isabelle/HOL formal development archived at https://doi.org/10.5281/zenodo.22700396 Andreas Andreakis http://arxiv.org/abs/2609.12597v1 Invisible Yet Dominant: Big Stalls of Kernel I/O Mechanisms in Cloud OLTP Databases 2026-09-11T08:49:32Z Most databases, including PostgreSQL, RocksDB, and recent AI KV-cache middleware, rely on buffered I/O, delegating write-back to the Linux kernel. On the distributed block storage standard in the cloud, this delegation inherits a hidden bottleneck: each device is drained by a single kernel flusher thread over a high-latency, shallow-queue path. When the drain falls behind, dirty throttling pauses write() system calls, and even reads that must evict dirty pages stall. These stalls are invisible to iostat and every standard counter. This poster observes the stall from inside the kernel, using the multi-volume data placement proposed in SteelDB as the experimental lever. eBPF probes on writeback and block tracepoints separate write-back by issuing context and count every throttle pause. Across three configurations with identical provisioned IOPS and bandwidth but 1, 2, and 4 devices, we show that adding drains, not bandwidth, cuts throttle pauses by 70%, reduces maximum transaction latency by 59%, and raises throughput by 23%. 2026-09-11T08:49:32Z Poster presented at SOSP 2026. Companion to arXiv:2603.29052 (SteelDB) Mitsumasa Kondo http://arxiv.org/abs/2609.12535v1 QEmbed: A Deep Learning Based Cardinality Estimator for Efficient Query Processing 2026-09-11T07:43:48Z Cardinality estimation is at the core of any commercial database system for efficient query processing. Over the decades, non-learning-based estimation techniques (e.g., histogram-based, sampling-based) have been widely used in both commercial and open-source database platforms. However, these techniques are only effective when the number of columns in a table is small, as they cannot properly capture dependencies between multiple attributes. Recently, learning-based approaches have been shown to perform significantly better than the heuristic methods that have been used for the past three decades. Despite this success, existing learned models often struggle to balance memory efficiency and accuracy when dealing with datasets that mix high and low cardinality attributes. In this paper, we propose a deep learning model formally called QEmbed. Our model is built upon the Masked Autoencoder for Distribution Estimation (MADE) auto-regressive framework to learn joint data distributions for selectivity estimation. To improve data representation and overcome the limitations of using a single encoding method, we design a hybrid encoding scheme that combines one-hot and embedding encodings. This hybrid design enables QEmbed to retain fine-grained attribute information for smaller domains while capturing compact semantic patterns for large, sparse domains. We capture attribute correlations by factoring the joint data distribution into a series of conditional probabilities. This approach naturally accommodates both point and range queries. Through extensive experiments, we show that while QEmbed faces a latency trade-off on extremely wide schemas, it provides highly reliable cardinality estimates overall. A key advantage of our model is that it reduces extreme tail errors (maximum Q-errors), avoiding catastrophic estimation failures on complex, highly correlated workloads. 2026-09-11T07:43:48Z Pooja Rajput Suman Banerjee http://arxiv.org/abs/2604.13050v2 Exploring Urban Land Use Patterns by Pattern Mining and Unsupervised Learning 2026-09-10T23:41:25Z Comparative planning needs reproducible methods for identifying recurring land-use configurations across cities. Using Urban Atlas 2018 data for 100 European urban areas, we construct 290,396 focal-neighborhood transactions and 1,543 frequent-itemset support features at 10\% minimum support. Ward clustering is applied in the original normalized feature space, with UMAP used only for visualization. A seven-cluster descriptive solution is retained through multi-criterion evaluation and 500 paired feature-subsampling repetitions. Sensitivity analyses show stable city-similarity geometry across 5--15\% thresholds but greater membership sensitivity to representation choices and non-artificial land-use content. The framework supports peer-city comparison while making scale and boundary limitations explicit. 2026-03-17T01:29:33Z Zdena Dobesova Tai Dinh Pavel Novak http://arxiv.org/abs/2608.03477v2 Getting to the Root: A Combined Complexity Perspective on Consistent Query Answering 2026-09-10T21:12:50Z We study the combined complexity of consistent query answering for Boolean self-join-free conjunctive queries with unary primary keys and acyclic attack graphs. Although every fixed query in this class admits a first-order rewriting [21], we show that allowing the query to vary makes the problem Pi_2^P-complete. To isolate the query structure governing this transition, we introduce the closure generator size: the minimum number of query variables whose closure under the functional dependencies induced by the primary keys contains every query variable. Our main theorem shows that every fixed bound on this parameter yields polynomial-time combined complexity. Parameterized by the closure generator size, the problem is in XP and co-W[t]-hard for every fixed t. The polynomial-time result is specific to unary primary keys: with binary primary keys, the problem becomes coNP-hard already for self-join-free queries with acyclic attack graphs and closure generator size zero. Regarding the upper bound, CQA is in coNP for arbitrary Boolean conjunctive queries of bounded closure treewidth or bounded hyperclosure treewidth; these width measures originate in [1]. Our polynomial-time algorithm realizes quantifier alternation dynamically by interleaving existential and universal substitutions with variable-level reductions. To implement this approach, we introduce concepts and techniques such as attack-propagation graphs, variable guarding and guarded pruning, database and query saturation, and query normalization. 2026-08-04T11:15:16Z Miika Hannula http://arxiv.org/abs/2605.01564v3 The TripleA Principle: Making Knowledge Actionable, Applicable, and Auditable in Post-FAIR Infrastructures 2026-09-10T17:19:10Z Ecological restoration, species distribution modelling, and invasive species management share a difficulty: knowledge that is findable and reusable carries no explicit account of the conditions under which it can be validly applied, or of the evidence grounding them. Applying it correctly is therefore demanding and expert-dependent, and misapplication usually goes unrecorded. The FAIR and CLEAR principles improved the findability, accessibility, interoperability, reusability, and human-interpretability of knowledge, but these address properties of representation, and reliable action requires more. Bridging the knowledge-action gap requires characterizing knowledge in terms of the operations it supports. Analysing what an operation needs, we derive three capabilities a knowledge representation must support. Actionability is the capacity to supply the knowledge and objective an operation executes. Applicability is the capacity to assess whether it can be reliably performed, through explicit conditions evaluated against context. Auditability is the capacity to assess the empirical grounding for that reliability, through documented success and failure. These form the three criteria of the TripleA Principle, an implementation-indipendent guide for next-generation knowledge infrastructures, jointly sufficient for the representational preconditions of reliably grounded action though not for its justification. Building on the Semantic Units Framework, we realize the principle as action units, typed components in which the knowledge an operation executes, the conditions under which it may validly be applied, and its documented successes and failures are addressable and evaluable. Action units form a nested hierarchy in which documented failure refines the conditions of valid use, letting knowledge graphs act as context-sensitive, evidentially accountable decision-support systems. 2026-05-02T18:25:27Z Lars Vogt Robert Fruehstueckl Tim Alamenciak http://arxiv.org/abs/2607.18029v2 Natural Language Access to Domain-Specific Metadata: A Reusable Framework for LLM Query Generation 2026-09-10T15:49:22Z Researchers need to answer ad-hoc questions about the contents of domain-specific archives but often lack the expertise to write structured queries on the metadata. We show that when domain vocabulary and semantics are captured in a well-designed Web Ontology Language (OWL) ontology, Large Language Models (LLMs) can generate accurate structured queries zero-shot, without task-specific fine-tuning, retrieval augmentation, or multi-agent orchestration. We present the Natural Language Knowledge Graph Query (NLKGQ) system, a framework and development process that enables natural language access to metadata in such archives. The framework includes a web interface that helps researchers pose natural language questions, which a domain-agnostic harness translates to SPARQL via an LLM and executes against a knowledge graph. The development process begins with capturing domain vocabulary and semantics in a formal OWL ontology. Domain-specific code then extracts metadata from archive sources and imports it into a knowledge graph defined by the ontology. Both are designed for reuse across domains. We demonstrate the system on metadata derived from a large-scale neuroimaging research archive, evaluating multiple LLMs and ontology representations. The best configurations achieve 100% accuracy on a 21-question competency and regression test set developed with domain experts. An ablation study across eight ontology representations reveals that readable entity names and semantic annotations are the dominant factors in accuracy, more significant than model choice or prompt engineering. We also compare SPARQL to an auto-generated SQL database as query backends, showing that OWL's structural features provide a substantial advantage over SQL DDL for LLM-driven query generation. Our demonstration domain requires local LLMs on modest institutional hardware to address privacy concerns for human subject data. 2026-07-20T14:59:33Z Blake G. Fitch Cato Elia Kurtz http://arxiv.org/abs/2603.23070v2 Knowledge management in House of Graphs 2026-09-10T11:58:04Z The House of Graphs is an online database of graphs which can be accessed at https://houseofgraphs.org/. It serves as a central repository for complete lists of graphs for various graph classes. However, its main feature is a searchable database of so-called "interesting" graphs. The development of the original House of Graphs started in 2010 and it was completely rebuilt in 2021-2022. Each graph in the database is accompanied by a significant amount of meta-data such as a name, drawings, precomputed graph invariants, and comments. Given this amount of information and the importance of reliability in the scientific world, robust data management is essential to ensure accuracy and consistency across the database. In this article, we therefore focus on knowledge management in the House of Graphs and describe the inner workings of the House of Graphs and how we ensure that its data is coherent, qualitative and stable. 2026-03-24T11:07:05Z 19 pages Gauvain Devillez Sven D'hondt Jan Goedgebeur http://arxiv.org/abs/2609.11390v1 VikingRAG: Accurate and Token-efficient Retrieval-augmented Generation over Structured Documents 2026-09-10T11:25:59Z State-of-the-art retrieval-augmented generation (RAG) methods exploit document structures to acquire sufficient evidence, but often incur substantial token costs. To reduce structural-context tokens without compromising high RAG accuracy, we present {\sf VikingRAG}, a directory-aware semantic data management system that tightly integrates semantic and structural access to support structural-context-efficient, evidence-gap-driven multi-round retrieval. To further reduce token overhead of multi-round interaction, we materialize agentic multi-round retrieval traces as experience edges, and reuse these edges for similar queries, avoiding repeated multi-round exploration. To additionally reduce token costs when agentic multi-round retrieval is unnecessary, we introduce an adaptive escalation strategy that answers from one-round experience-augmented retrieval when the evidence is sufficient, and invokes agentic multi-round retrieval only otherwise. Experiments on real datasets show that the base system {\sf VikingRAG} matches high accuracy of state-of-the-art methods while consuming only 11.6\%--51.9\% of their tokens. With retrieval-trace reuse and adaptive escalation, token costs drop to 5.1\%--32.5\% while maintaining competitive accuracy and practical document-storage performance, showing the utility of this work for emerging AI knowledge bases. 2026-09-10T11:25:59Z Peiyuan Gao Gaoyuan Zhang Haojie Qin Yahui Sun Qianyi Zhang Yunhao Zhang Zeyu Wang Wei Lu http://arxiv.org/abs/2605.19320v3 TextAlign: Preference Alignment for Text Rendering with Hierarchical Rewards 2026-09-10T09:59:25Z Faithful text rendering remains a persistent weakness of large text-to-image generative models, as it requires both semantic instruction following and fine-grained glyph-level structure. Prior methods often improve this ability through architecture-specific modules or encoder modifications, which complicate deployment across foundation models. We study text rendering as a post-training preference-alignment problem and propose TextAlign, a non-invasive framework that keeps the generator architecture unchanged. The key component is a hierarchical vision-language model (VLM)-based reward that decomposes rendering errors into global, word, and glyph levels, then converts binary defect judgments into a scalar preference signal. The resulting signal supports both Group Relative Policy Optimization (GRPO) and Direct Preference Optimization (DPO). Experiments on FLUX.1-dev and Z-Image-Turbo show consistent gains in OCR-based text accuracy without degrading general generation quality. Compared with strong foundation and text-rendering baselines, including SD3.5, Qwen-Image, AnyText, and TextDiffuser, these results indicate that reward design offers a scalable alternative to model redesign for improving text rendering. 2026-05-19T03:55:59Z Mingxuan Cui Jingpu Yang Fengxian Ji Qian Jiang Zhecheng Shi Jiaming Wang Zirui Song Zhuohan Xie Fajri Koto Xiuying Chen