https://arxiv.org/api/Oxbj0pZRrG6y3Zgu60K6z5vSzAc 2026-10-02T17:24:13Z 12468 135 15 http://arxiv.org/abs/2609.15058v1 Fast Label-Filtering Approximate Nearest Neighbor Search via Progressive Label Set Stratification 2026-09-14T05:23:52Z Approximate nearest neighbor search (ANNS) retrieves the most similar vectors to a query vector in high-dimensional space. Label-filtering ANNS (LFANNS) extends ANNS with a label filter that the labels of base vectors must satisfy a set relation (e.g., equality, containment, or overlap) with the query labels. Existing LFANNS indices suffer from inconsistent performance across different filter types and degraded scalability under varying label scale and distribution. In this paper, we define label-stratified similarity graph (LSSG), where edges connect neighboring vectors whose label sets fall within stratified similarity thresholds. To implement LSSG efficiently, we design an incremental insertion algorithm to prune redundant edges in both vector and label spaces, and leverage a MinHash structure to ensure scalability for large-scale labels. We analyze stepwise probabilities under explicit label models and explain why stricter label tiers reduce ineffective in-filtering expansions. Benchmark experiments show that LSSG achieves ideal optimality for equality queries, and 1.06x-92.9x and 1.08x-84.1x faster than the best competing index for containment and overlap, respectively, in query speed with identical accuracy and 0.35x index size. 2026-09-14T05:23:52Z Accepted in the ACM SIGMOD/PODS International Conference on Management of Data (SIGMOD 2027) Ziqi Wang Jingzhe Zhang Shuo Shen Wei Hu http://arxiv.org/abs/2509.00303v4 Access Paths for Efficient Ordering with Large Language Models 2026-09-14T05:13:18Z In this work, we present the \texttt{LLM ORDER BY} semantic operator as a logical abstraction and conduct a systematic study of its physical implementations. First, we propose several improvements to existing semantic sorting algorithms and introduce a semantic-aware external merge sort algorithm. Our extensive evaluation reveals that no single implementation offers universal optimality on all datasets. From our evaluations, we observe a general scaling relationship between sorting cost and the ordering quality for comparison-based algorithms. Building on these insights, we design a budget-aware optimizer that utilizes heuristic rules, LLM-as-Judge evaluation, and consensus aggregation to dynamically select the near-optimal access path for LLM ORDER BY. In our extensive evaluations, our optimizer consistently achieves ranking accuracy on par with or superior to the best static methods across all benchmarks. We believe that this work provides foundational insights into the principled optimization of semantic operators essential for building robust, large-scale LLM-powered analytic systems. 2025-08-30T01:44:36Z Fuheng Zhao Jiayue Chen Yiming Pan Tahseen Rabbani Sohaib Divyakant Agrawal Amr El Abbadi Paritosh Aggarwal Anupam Datta Dimitris Tsirogiannis http://arxiv.org/abs/2604.08849v4 SatIR: Scalable High-Recall Constraint-Satisfaction-Based Information Retrieval for Clinical Trials Matching 2026-09-14T05:01:37Z Many real-world retrieval and matching problems require more than topical relevance: a candidate must satisfy the specific constraints of one profile among many, not just be relevant to it. Clinical trials are a high-stakes instance of this challenge: they are central to evidence-based medicine, yet many struggle to meet enrollment targets, despite the availability of over half a million trials listed on ClinicalTrials.gov, which attracts approximately two million users monthly. Existing retrieval techniques, largely based on keyword and embedding-similarity matching, treat eligibility constraints as soft signals rather than binding requirements, resulting in low recall, low precision, and limited interpretability. We propose SatIR, a scalable, efficient, high-precision, high-recall, interpretable clinical trial retrieval method based on formal constraint satisfaction. Leveraging established medical ontologies, we use Large Language Models (LLMs) to convert informal reasoning -- regarding ambiguity, implicit clinical assumptions, and incomplete patient records -- into explicit, precise, controllable, and interpretable formal Satisfiability Modulo Theories (SMT) constraints. For scalable and efficient retrieval, we project the SMT matching problem onto relational algebra, enabling an efficient database implementation that retains high recall while sacrificing little precision. SatIR consistently improves eligibility-aware retrieval over similarity-based baselines on the SIGIR 2016 dataset and a benchmark derived from TREC 2022. Relative to TrialGPT-style retrieval, SatIR retrieves 32%-72% more relevant-and-eligible trials per patient on SIGIR 2016 and achieves 1.8-3.2x higher eligible-trial recall on the TREC benchmark. Retrieval is fast, requiring only 146 milliseconds per patient over 3,621 SIGIR trials. 2026-04-10T01:13:44Z Accepted at COLM 2026. 125 pages (11 pages main text, 110 pages appendix), 24 figures, 27 tables. Code: https://github.com/stanford-oval/clinical-trial-matching Project page: https://satir.genie.stanford.edu Zikai Zhou Yufei Jin Yilin Xu Yu-Chiang Wang Chieh-Ju Chao Monica S. Lam http://arxiv.org/abs/2609.15034v1 FastPair: GPU-Optimized String Decoding 2026-09-14T04:51:52Z Modern data systems compress data at rest and decompress it only when needed to preserve interconnect bandwidth. This design is often inefficient on GPU-based compute platforms because many conventional compression techniques exhibit serial data dependencies that limit GPU parallelism, leaving resources idle. Recent NVIDIA GPUs address this decoding deficiency through the Decompression Engine (DE), an on-die, fixed-function decompression accelerator for general-purpose compression formats such as Deflate, LZ4, and Snappy. Recent work has proposed string codecs that replace frequent substrings with fixed-width codes from a small, trained dictionary, making each code's lookup independent. While these lookups can run in parallel, the resulting scattered reads and short output writes still do not align well with GPU hardware, which handles contiguous memory accesses more efficiently. We present FastPair, a GPU decoder that optimizes the existing dictionary decoding process by reorganizing lookups and assembling decoded substrings for contiguous output writes. On a B300, FastPair decodes ten real-world columns 2.4 to 4.2x faster than the DE, reaching up to 1.6 TB/s. 2026-09-14T04:51:52Z Joseph Isaacs Francesco Gargiulo Peter Boncz Robert Kruszewski Nicholas Gates Rossano Venturini Will Manning Martin Prammer http://arxiv.org/abs/2607.23881v2 Answering Conjunctive Queries with Aggregations under Updates 2026-09-14T04:36:00Z Dynamic query processing keeps query answers up to date during insertions and deletions. For conjunctive queries (CQs) under set semantics, the classes maintainable in constant amortized time are known exactly: the $q$-hierarchical CQs under arbitrary updates, and the free-connex CQs under insertion-only updates. Many analytics tasks, including \textsf{SUM}/\textsf{COUNT} aggregations, provenance, and access control, are captured by evaluating a CQ over a positive commutative semiring. We thus ask whether aggregation changes what can be maintained efficiently, and if so, when. Under \emph{insertion-only} updates, it does: the boundary retreats from free-connex to a new class we call \emph{strong-connex}, with $q\text{-hierarchical} \subsetneq \text{strong-connex} \subsetneq \text{free-connex} \subsetneq \text{acyclic}$. For every \emph{strictly monotone} semiring, including the sum-product and tropical semirings, no free-connex but non-strong-connex CQ is maintainable in $O(|D|^{1/2-ε})$ time under the OuMv and OMv conjectures, whereas every strong-connex CQ is maintainable in $O(1)$ amortized time over every semiring. Under \emph{arbitrary} updates, the boundary stays at the $q$-hierarchical CQs for every semiring with $O(1)$-deletable aggregates, and maintenance over any semiring is at least as hard as over the Boolean semiring. We further strengthen the lower bounds to semirings that fall outside the class and to query with different \emph{height} and \emph{dimension}, under the combinatorial $k$-clique and generalized OuMv conjectures. All upper bounds come from a single framework, obtained by adapting CROWN to annotated relations; together with the lower bounds, they yield dichotomies parameterized by both the query and the semiring, recovering the Boolean results as a special case. 2026-07-26T22:58:57Z Qichen Wang Xiao Hu http://arxiv.org/abs/2609.14780v1 The Stochastic Deputy: Structural Tenant Isolation for Tool-Using LLM Agents 2026-09-13T20:40:11Z Multi-tenant tools commonly accept a tenant identifier and validate it against the caller's entitlement. For a large language model (LLM) agent, that pattern delegates resource selection to a process whose context may contain attacker controlled instructions. We formalize this stochastic deputy problem and present a structural defense: remove tenant identity from the Model Context Protocol (MCP) tool schema, bind scope to a verified credential, and enforce it below the agent. In a 373-trial ablation across eight model configurations and two transports, a correctly validated tenant parameter served every out-of-scope attempt: 26 of 26, or 26 of 41 plausible-pretext trials overall. With the parameter removed, no tool signature could express the read. Twelve of 56 trials instead escaped the interface by forging writable scope, showing that interface invariance requires cryptographically protected context. On a production dataset containing multiple GBs of data, set-valued scope caused a measured $57\times$ latency ratio under function-wrapped membership predicates; a JSON_TABLE lateral join recovered index access where the tenant key was indexed. The evaluation also exposes deployment limits, including an entitlement-size query-planner cliff and incomplete index coverage. The result is a tenant-isolation argument that depends on enforceable interfaces and credentials rather than model compliance. 2026-09-13T20:40:11Z 29 pages, 7 figures, 15 tables Mirza Samad Ahmed Baig Syeda Anshrah Gillani Asher Ali Muhammad Hamzah Siddiqui http://arxiv.org/abs/2609.14652v1 Natural Language Knowledge Graph Query Execution: Leveraging Controlled Semantics in the LLM Context Window 2026-09-13T16:36:40Z Large Language Model (LLM) applications often transfer domain concepts into the model's context informally, through prompt prose, schema dumps, and examples. We show that for database queries, data model concepts pass to LLMs more effectively through representations whose vocabulary terms carry declared, machine-readable semantics (controlled semantics). NLKGQ is a working system and reusable framework that does this for data modeled in a knowledge graph. A formal OWL ontology serves as the transfer mechanism, concentrating the meaning of the data into semantically precise tokens the model can use directly. In a single LLM call, NLKGQ places in the context a system prompt instructing on SPARQL, the complete domain OWL ontology, and a domain-specific prompt addition, together with the user's natural language query. The model then generates the SPARQL query directly, zero-shot. Where the native vocabulary of an existing database or federation of databases is opaque, a wrapper ontology substitutes clean terms and a runtime rewriter restores the native forms. Evaluating on DBLP-QuAD 2.0 showed that its scores depend on the graph snapshot, the endpoint used, and the wording of its machine-generated questions, so we propose DBLP-QuAD 3.1, which maintains the intent of 2.0 while making reference results deterministic, revising reference SPARQL where needed, and rewriting the natural language questions, with a frontier model, to state each reference query's intent clearly and completely. We evaluate on the DBLP-QuAD 2.0 benchmark (57.6% Match under deterministic re-scoring), DBLP-QuAD 3.1 (89.9% Match on 1,000 questions), SemOpenAlex (98% Match against a published baseline's 86% on the identical test set), and neuroimaging metadata (100%). 2026-09-13T16:36:40Z 9 pages, 4 tables Blake G. Fitch http://arxiv.org/abs/2603.21710v2 FGIM: a Fast Graph-based Indexes Merging Framework for Approximate Nearest Neighbor Search 2026-09-13T14:49:13Z As the state-of-the-art methods for high-dimensional data retrieval, Approximate Nearest Neighbor Search (ANNS) approaches with graph-based indexes have attracted increasing attention and play a crucial role in many real-world applications, e.g., retrieval-augmented generation (RAG) and recommendation systems. Unlike the extensive works focused on designing efficient graph-based ANNS methods, this paper delves into merging multiple existing graph-based indexes into a single one, which is also crucial in many real-world scenarios (e.g., cluster consolidation in distributed systems and read-write contention in real-time vector databases). We propose a Fast Graph-based Indexes Merging (FGIM) framework with three core techniques: (1) Proximity Graphs (PGs) to k Nearest Neighbor Graph (k-NNG) transformation used to extract potential candidate neighbors from input graph-based indexes through cross-querying, (2) k-NNG refinement designed to identify overlooked high-quality neighbors and maintain graph connectivity, and (3) k-NNG to PG transformation aimed at improving graph navigability and enhancing search performance. Then, we integrate our FGIM framework with the state-of-the-art ANNS method, HNSW, and other existing mainstream graph-based methods to demonstrate its generality and merging efficiency. Extensive experiments on six real-world datasets show that our FGIM framework is applicable to various mainstream graph-based ANNS methods, achieves up to 3.5X speedup over HNSW's incremental construction and an average of 7.9X speedup for methods without incremental support, while maintaining comparable or superior search performance. 2026-03-23T08:53:17Z 27 pages, accepted by SIGMOD 2026 Zekai Wu Jiabao Jin Peng Cheng Xiaoyao Zhong Lei Chen Yongxin Tong Zhitao Shen Jingkuan Song Heng Tao Shen Xuemin Lin 10.1145/3786651 http://arxiv.org/abs/2607.09149v2 Taxonomy Maintenance In The Wild Over Evolving Scholarly Data: Reliability, Efficiency, and Cost-Effectiveness 2026-09-13T12:14:55Z The rapid growth of scientific publications makes scholarly taxonomies quickly obsolete. We study taxonomy maintenance in the wild, a new problem that moves beyond static construction by continuously adapting taxonomies to evolving scholarly repositories, such as arXiv, for a given research topic. We propose GIST, a robust framework for maintaining evolving taxonomies. Unlike purely LLM-centric approaches, GIST grounds structure induction in expert-curated evidence by extracting partial hierarchies from the "Related Work" sections of papers. It integrates these partial taxonomies into a unified global taxonomy in a geometric box-embedding space, where box containment encodes the inductive bias of is-a relations. To connect semantics with geometric structure, GIST learns a bidirectional mapping between word embeddings and box embeddings. For efficient incremental updates, GIST uses novelty-aware coreset selection to update the model with representative historical signals and new evidence, avoiding costly full retraining. To handle high-velocity paper streams under user-specific token budgets, GIST further combines a hypothesized concept generator with a cost-effective evidence retrieval module. Experiments on real-world arXiv datasets show that GIST outperforms state-of-the-art baselines, improving Node F1 and Edge F1 by 11.0% and 13.1% over the strongest baseline while requiring only 9.6% of its runtime and 12.7% of its monetary cost. 2026-07-10T07:03:08Z The paper has been accepted by SIGMOD 2027 Daomin Ji Hui Luo Zhifeng Bao Junhao Gan Zi Huang http://arxiv.org/abs/2609.05014v2 Reducing the Cross-Model Tax: Query Optimization over Multi-Model Data 2026-09-13T09:27:52Z Querying across heterogeneous data models incurs overhead from query decomposition, result retrieval and conversion, and processing outside the underlying database systems. This paper investigates the extent to which, in a decomposition-based architecture, this cross-model tax results from decisions made by the unifying query processor rather than from heterogeneity alone. We present a mapping- and capability-aware optimization approach that moves applicable processing into native query parts. It combines model-aware predicate pushdown, cross-model dependent joins, and non-redundant query-part construction within a unified pipeline spanning relational, document, and graph databases. The approach is implemented in MM-quecat and evaluated using 20 read-only queries across PostgreSQL, MongoDB, and Neo4j, as well as a heterogeneous combination of the three systems in a single-machine, containerized deployment. For the query--environment combinations most affected by large intermediate results, predicate pushdown yields maximum observed latency reductions of up to two orders of magnitude and prevents the out-of-memory failures observed in the original single-DBMS experiments. Dependent execution further improves eligible external joins, while non-redundant construction reduces planning time for the largest evaluated graph plans, from hundreds of milliseconds to several milliseconds. The results show how established optimization principles can be applied across conceptual, mapping, data-model, and DBMS boundaries in decomposition-based multi-model query processing. 2026-09-04T11:23:57Z Jáchym Bártík Filip Štrobl Irena Holubová http://arxiv.org/abs/2609.14370v1 DiaLSM: Towards Write-Stall-Free Performance via Shard-based LSM-tree 2026-09-13T08:04:23Z Log-Structured Merge-tree (LSM) aims to achieve high write throughput, but is known to experience the write stall problems when subjected to sustained write pressure. We quantify the occurrence probability and average duration of write stalls in LSM using a queuing model in the write--flush--compaction pipeline, moving beyond existing empirical analysis. The proposed model demonstrates that a monolithic LSM with a single pipeline cannot eliminate write stalls, revealing that internal sharding within the LSM offers an opportunity for fundamental write stall mitigation. To break this structural bottleneck, we propose DiaLSM, an internally shard-based LSM architecture. Instead of forcing all writes through one pipeline, DiaLSM splits the write--flush--compaction path into multiple independent shards and employs dynamic fallback redirection, allowing writes to proceed even when some shards stall. Implemented on RocksDB, DiaLSM achieves up to 2.4x higher throughput, 94% lower stalls, and significantly lower latency than state-of-the-art methods ADOC and Sub-Compaction, as demonstrated by db_bench, YCSB, and Sysbench OLTP evaluations. 2026-09-13T08:04:23Z Accepted to the 43rd IEEE International Conference on Data Engineering (ICDE 2027) Hongsu Byun Safdar Jamil Honghyeon Yoo Sungyong Park Myungcheol Lee Xubin He Zhichao Cao Youngjae Kim http://arxiv.org/abs/2609.14255v1 Towards Anticipatory Databases Through Shared Data and Workload Semantics 2026-09-13T03:12:43Z Database management systems increasingly serve dynamic and exploratory workloads, yet many of their decisions still rely on low-level signals such as recency, frequency, and address locality. These signals capture how data was accessed, but not what is being examined or how an analytical focus evolves. We argue for treating workload semantics as a first-class control signal for anticipatory decision making. Central to this view, we introduce semantic locality and semantic trajectories, which capture relationships among nearby queries and how those relationships evolve across a session. We propose a framework that represents semantic context at the data, query, and session levels, models its evolution over time, and translates it into task-specific utility estimates. We instantiate this framework in semantic prefetching and semantic cache eviction, which share a semantic layer to make two separate decisions. Prefetching uses semantic trajectories to anticipate future accesses beyond what address-based locality can capture, while eviction uses semantic relevance to inform block replacement. These systems provide initial evidence that shared semantic context can support multiple DBMS components. We further outline how this principle can extend to other decisions and data systems, and discuss key challenges in representation, cost, adaptation, and evaluation. 2026-09-13T03:12:43Z 8 pages, 6 Figures Farzaneh Zirak Kasper Overgaard Mortensen Farhana Choudhury Renata Borovica-Gajic http://arxiv.org/abs/2609.14121v1 Semantic Knowledge Technologies: what the Semantic Web lost sight of, and what it never had 2026-09-12T19:55:46Z The Semantic Web set out to give information a machine-interpretable form so that software could integrate and reason over it. Its standards became scientific knowledge infrastructure, but the machine competence it promised did not follow, and the systems now answering questions over scientific knowledge are language models holding no inspectable account of what they know. This paper argues the original goal was right and the technical programme incomplete, states what is missing, and names the extended programme Semantic Knowledge Technologies: the same technical core carried out of its web-publishing origin and applied to knowledge wherever held. The diagnosis is that the standards formalised truth while omitting three things: the conditions under which a claim holds, the operations its terms permit, and any account of what a base covers. Without conditions, contradiction and applicability cannot be judged; without operational grounding, holding a statement confers no ability; without declared coverage, a system cannot recognise the boundary of its own content, which under the open-world assumption cannot be inferred. The paper fixes the word understanding to five measurable tests (check, connect, derive, act, delimit) and sets out a seven-layer architecture in which the first three layers are enabling and the rest the cognitive capabilities they make possible. It then defines three terms the programme implies: Large Knowledge Model, a model whose unit of output is a reference to an addressable claim, not a token; SLKM, the knowledge base an agent builds for itself from declared sources; and Semantic Artificial General Intelligence, stated as a falsifiable position about necessary conditions, not a system. A graded ladder replaces the untestable word general. It is offered as a research agenda, with its weakest points and refutation condition named. 2026-09-12T19:55:46Z Achille Zappa http://arxiv.org/abs/2609.13929v1 Specification-Driven Data Architecture Reconstruction: From Physical Code to Logical and Conceptual Specifications 2026-09-12T13:24:41Z Legacy database migrations often begin with incomplete or outdated documentation, leaving physical data definition language (DDL) as the principal evidence of data architecture. However, DDL does not fully encode conceptual intent, and model-generated completions can be plausible without being correct. This study proposes and evaluates a provenance-aware, deterministic-first pipeline for reconstructing logical and conceptual data specifications from Oracle-oriented DDL while explicitly separating observed facts, deterministic derivations, and large language model (LLM) suggestions. The pipeline performs DDL investigation, parsing, consolidation, primary-key backfilling, type normalization, and declared relationship-graph construction before optional LLM-assisted enrichment. It preserves source provenance in the deterministic catalog and declared relationship graph, and records inferred primary-key and foreign-key candidates in a separate reviewable overlay. We evaluated the implementation on 249 artifactized schema samples comprising 1,225 SQL files. The pipeline completed 244 samples (97.99%); 52 completed samples contained no extractable DDL. Across completed samples, the deterministic path reconstructed 208 tables and recovered 278 declared foreign-key records; 168 of the reconstructed tables lacked an explicitly parsed primary key before backfilling. LLM enrichment generated 100 foreign-key candidates in 36 samples, but the parent-table admissibility rate was only 17.9% for logical-specification candidates and 17.5% for conceptual-specification candidates. These findings show that the proposed deterministic-first architecture can preserve an auditable structural baseline, quantify observed primary-key and relationship gaps, and prevent model-generated hypotheses from being silently promoted to source-grounded architectural facts. 2026-09-12T13:24:41Z 15 pages, 3 figures, 7 tables, 35 references Oleg Grynets Olena Pochernina Vasyl Lyashkevych http://arxiv.org/abs/2606.07795v2 The Role of Semirings in Incremental View Maintenance 2026-09-12T08:39:46Z We study the problem of incremental view maintenance (IVM) under inserts to semiring-annotated databases. The key observation put forward in this paper is that the complexity of the IVM problem depends fundamentally on the underlying semiring. We introduce a class of conjunctive queries called p-hierarchical. For a zero-sum free and zero-divisor free commutative semiring $K$, we show that for any p-hierarchical query with fractional hypertree width fhtw and any insert-only update sequence of length N to an initially empty K-database, we can construct a data structure that can be updated in O(N^{fhtw-1}) amortized time and supports the enumeration of the query result with constant delay. In particular, the amortized update time for any p-hierarchical alpha-acyclic query is constant. For a class of semirings used to model a wide range of computational problems, we give conditional lower bounds showing that any conjunctive query without self-joins that is not p-hierarchical cannot be maintained with constant amortized update time and constant enumeration delay under inserts. This class includes the natural semiring and its generalizations to the provenance and covariance semirings, as well as idempotent and strictly ordered semirings such as the tropical semiring. When put together, our upper and lower bounds imply a dichotomy for the insert-only maintenance of conjunctive queries without self-joins over every semiring in our class: such a query can be maintained with constant amortized update time and constant enumeration delay if and only if it is p-hierarchical alpha-acyclic. 2026-06-05T19:19:27Z Eden Chmielewski Andrei Draghici Dan Olteanu Haozhe Zhang