https://arxiv.org/api/kEE1zgs4dTKNWmlbE0MExRPN1q0 2026-10-03T23:22:05Z 12468 240 15 http://arxiv.org/abs/2608.29993v1 Demystifying and Improving Lazy Promotion in Cache Eviction 2026-08-30T19:38:02Z Cache eviction algorithms play a critical role in the performance of modern data systems, yet their scalability is often limited by the high computational overhead associated with object promotions. Lazy Promotion techniques have emerged as relaxations of traditional Least-Recently-Used (LRU) methods, designed to alleviate lock contention and increase throughput. This work uses production traces from real-world systems to benchmark five Lazy Promotion strategies: Probabilistic-LRU, Batch-LRU, Delay-LRU, FIFO-reinsertion, and Random-LRU. We evaluate these techniques across miss ratio, scalability, promotion count, and a novel metric called promotion efficiency, which measures the number of hits per promotion. Our results reveal that Delay-LRU and FIFO-reinsertion significantly improve promotion efficiency, whereas Batch-LRU and Probabilistic-LRU struggle to reduce promotions without significantly increasing miss ratio. We further explore the impact of lazy promotion in advanced algorithms such as ARC and 2Q and make a similar observation. Moreover, we uncover substantial optimization potential, showing that most cache promotions are unnecessary when equipped with oracle knowledge. To further reduce promotions in LRU, we propose two novel enhancements-Delayed FIFO-reinsertion (D-FR) and Age-Guided Eviction (AGE)-that reduce promotions by 20-60% while achieving a similar or lower miss ratio. 2026-08-30T19:38:02Z 14 pages, 13 figures, and 3 tables. Published in PVLDB 19(4); VLDB 2026 conference paper Proceedings of the VLDB Endowment, 19(4): 549-562, 2025 Qinghan Chen Muhammad Haekal Muhyidin Al-Araby Ziyue Qiu Zhuofan Chen Rashmi Vinayak Juncheng Yang 10.14778/3785297.3785299 http://arxiv.org/abs/2607.25891v2 Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation 2026-08-30T18:22:18Z Comprehensively evaluating AI agents across interactive environments is difficult due to fragmented tasks, scaffolds, verifiers, and scoring rules. Unfortunately, existing efforts to unify these evaluations are limited in scale and domain, making costly reruns necessary and leaving available data incomparable. We introduce MESSIER, a unified corpus of 957,611 records spanning 30 benchmarks, 745 agents, 11,891 tasks, and 74,263 verifiers. MESSIER combines public evaluation results with new runs on six underrepresented professional and scientific benchmarks, standardizing their heterogeneous components into a common schema. Using this corpus, we show that frontier progress is uneven across benchmark groups, with function-calling evaluations largely saturated, programming improving fastest, and enterprise workflows remaining most challenging. Counterfactual rescoring further shows that strict all-pass scoring in multi-verifier tasks can alter agent rankings. Finally, we derive capability scores from our corpus that correlate with Epoch's Evaluation Capability Index rankings at Spearman \r{ho} = 0.84. The scores can also be estimated for subsets defined by domain, occupation, action space, or verifier type. In essence, MESSIER is a reusable resource for studying agent performance at scale, and a basis for designing better evaluations. 2026-07-28T15:50:19Z Stefan Krsteski Charlotte Meyer Guillaume Allegre Tony O'Halloran Alexandre Sallinen http://arxiv.org/abs/2608.29543v1 Evaluating LLMs on Conversational Text-to-SQL under Chain Ambiguity and Intent Drift 2026-08-30T04:24:08Z Recent advances in large language models (LLMs) have established conversational text-to-SQL as a practical interface between users and databases, often involving multiple turns of clarification and revision. However, existing benchmarks primarily evaluate execution accuracy, leaving the unfolding and shifting of user intent across turns largely uncovered. To address this, we introduce TIDE-Bench, a benchmark for conversational text-to-SQL under chain ambiguity and intent drift evaluation, targeting two recurring patterns: chain ambiguity, where an underspecified question triggers layered clarification with conditional dependencies, and intent drift, where the user retracts and replaces a previously committed request element. Built on 514 anchor SQLs from BIRD, TIDE-Bench comprises 1,542 samples and introduces dedicated metrics for chain identification and drift recognition-resolution beyond execution accuracy. Evaluating 12 advanced LLMs reveals a persistent chain identification bottleneck unaffected by clarification frequency, a wide drift recognition-resolution gap, and overlap between failure modes when jointly activated. The corresponding code of TIDE-Bench is released for further research. 2026-08-30T04:24:08Z Accepted to EMNLP2026 Main Yujia Liu Jiayan Lin Zijin Hong Zheng Yuan Shengyuan Chen Hao Chen Qinggang Zhang Xiao Huang Feiran Huang http://arxiv.org/abs/2609.29661v1 Ingest-Time Fact Compilation for Cost-Efficient and Reliable Question Answering over Revised Corpora 2026-08-29T23:04:37Z Most agentic question answering (QA) systems do an important part of their semantic work at the worst possible time: every time someone asks a question. When a corpus contains revisions, drafts, revocations, deletions, and sources with different levels of authority, the model must reconstruct the governed current state on every read - then throw that work away and repeat it on the next query. This is a bit like a database that rebuilds a materialized view every time someone reads from it. We present ingest-time fact compilation, an architecture that performs this work when corpus data is ingested or changed. Raw passages are rephrased into self-contained facts; rules governing revisions, deletions, effective dates, and source trust are resolved once; and the resulting state is stored as typed records carrying source and revision provenance. At query time, an inexpensive model reads the compiled record instead of reconstructing it from noisy candidates. In a controlled synthetic experiment across five seeds, the same low-cost model produced the correct value, source, and revision in only one of 30 trials under query-time reconstruction, but in all 30 trials from the compiled substrate, at 12.89 times lower mean read cost per question. On simpler revision questions both architectures were exact, but the compiled path used 21.6 times fewer tokens. A separate test found that fact rephrasing roughly halved verbose Federal Reserve dialogue while preserving high source entailment, but left concise Wikipedia prose essentially unchanged. These results support a narrow but practical claim: resolving a corpus state once can make subsequent QA cheaper and more reliable for inexpensive models. We release the open source, MIT-licensed implementation and experimental artifacts. 2026-08-29T23:04:37Z 6 pages, 5 tables, 1 figure. Accepted for presentation at the 2026 International Conference on Applied Science and Technology - Engineering Science (iCAST-ES 2026), Surabaya, Indonesia, October 2026. Code and frozen experimental artifacts: https://github.com/aix-sc/isc (tag data-freeze-2026-07-16) Kyle Wild Yusuke Takahashi Asako Uraki http://arxiv.org/abs/2607.20489v2 EvoSQL: Memory-Augmented Critic-Generator Co-Evolution for Text-to-SQL 2026-08-29T16:48:45Z Text-to-SQL has advanced rapidly with large language models, but complex database queries still require reasoning beyond one-shot generation, including multi-step decomposition, execution-based diagnosis, and targeted correction. We present EvoSQL, a co-evolution framework that formulates SQL synthesis as an iterative interaction between a generator and a critic. EvoSQL maintains a contextualized candidate memory, verifies SQL candidates with both execution signals and LLM-based critique, and updates its memory through utility-guided aggregation. To strengthen the underlying generator-critic pair, we further introduce a Self-Distillation Policy Optimization (SDPO) fine-tuning stage that injects execution-aware supervision into modern coding LLM backbones. Experiments on Spider and BIRD show that EvoSQL consistently improves open-source models over Maj@16 baselines, with particularly large gains on BIRD-Dev, ranging from +1.37% for Qwen3-4B to +9.19% for Qwen2.5-Coder-3B. SDPO initialization further improves selected backbones on Spider-Test and BIRD-Dev. These results suggest that memory-grounded co-evolution is an effective path toward more reliable and generalizable Text-to-SQL systems. Code is available at https://github.com/valleysprings/EvoSQL. 2026-06-04T14:21:19Z Accepted as EMNLP Findings (2026) Jiawei Zhou Jianwei Wang Chenyu Zhou Chaojian Shi Ming Dong Kai Wang http://arxiv.org/abs/2509.08395v5 SINDI: An Efficient Index for Sparse Vector Approximate Maximum Inner Product Search 2026-08-29T03:18:40Z Sparse vector Maximum Inner Product Search (MIPS) is crucial in multi-path retrieval for Retrieval-Augmented Generation (RAG). Recent inverted index-based and graph-based algorithms have achieved high search accuracy with practical efficiency. However, their performance in production environments is often limited by redundant distance computations and frequent random memory accesses. Furthermore, the compressed storage format of sparse vectors hinders the use of SIMD acceleration. In this paper, we propose the sparse inverted non-redundant distance index (SINDI), which incorporates three key optimizations: (i) Efficient Inner Product Computation: SINDI leverages SIMD acceleration and eliminates redundant identifier lookups, enabling batched inner product computation; (ii) Memory-Friendly Design: SINDI replaces random memory accesses to original vectors with sequential accesses to inverted lists, substantially reducing memory-bound latency. (iii) Vector Pruning: SINDI retains only the high-magnitude non-zero entries of vectors, improving query throughput while maintaining accuracy. We evaluate SINDI on multiple real-world datasets. Experimental results show that SINDI achieves state-of-the-art performance across datasets of varying scales, languages, and models. On the MsMarco dataset, when Recall@50 exceeds 99%, SINDI delivers single-thread query-per-second (QPS) improvements ranging from 4.2 to 26.4 times compared with SEISMIC and PyANNs. Notably, SINDI has been integrated into Ant Group's open-source vector search library, VSAG. 2025-09-10T08:38:32Z 18 pages, accepted by ICDE 2026. Due to submission limitation for ICDE 2026 (i.e., maximum 6 submissions per author), Lei Chen and Xuemin Lin are not included as authors Ruoxuan Li Xiaoyao Zhong Jiabao Jin Peng Cheng Wangze Ni Zhitao Shen Wei Jia Xiangyu Wang Heng Tao Shen Jingkuan Song http://arxiv.org/abs/2606.23667v2 The Table Says Otherwise: Testing LLMs with Counterfactual Relational Data 2026-08-29T01:27:58Z Large language models (LLMs) are increasingly used to answer natural-language questions over structured data. However, when a table contains familiar real-world facts, it is unclear whether the model answers by reading the provided data or by recalling knowledge learned during pretraining. This distinction is important for database applications, where the provided tables should be the source of truth. In this paper, we introduce ContraTable, a paired original-counterfactual benchmark for evaluating whether LLMs ground their answers in relational tables. We build the benchmark with two aligned versions: an original database with real-world facts and a counterfactual database that preserves the same schemas, identifiers, and relationships while changing selected country, club, and player attributes. We design 214 matched questions across three levels: single-table lookup, multi-table lookup, and multi-table temporal reasoning. Experiments on commercial closed-source and open-source models show that strong instruction-tuned models can often handle direct lookup, but their reliability drops as questions require joins, comparison, and temporal reasoning. The gap between original and counterfactual accuracy reveals that models may fall back on prior knowledge when table evidence conflicts with familiar facts. These results suggest that table-QA evaluation should measure not only accuracy, but also faithfulness to the provided database. 2026-06-22T17:52:14Z Xinzhi Wang Chunwei Liu http://arxiv.org/abs/2608.28963v1 RENSA: Rich Environment Metadata to Navigate Shared and Distributed Endpoints for Automated Federated SPARQL Query Generation 2026-08-29T00:32:13Z The number of knowledge graph databases has increased significantly with the proliferation of knowledge graph technologies. Knowledge graphs enable the dynamic integration of distributed data through federated SPARQL queries. However, constructing efficient queries in a federated environment is challenging due to the lack of detailed structural knowledge across decentralized datasets. While standards like VoID provide basic metadata, they often fail to capture the complex interlinks and authority distributions necessary for optimization. Consequently, current engines frequently rely on runtime ASK queries for source selection, increasing communication overhead. We propose RENSA, a federated SPARQL query generation framework that leverages an extension of SPARQL Builder Metadata (SBM). By integrating class and authority information, mapping subject and object usage to specific predicates, RENSA enables precise source selection and semantic constraint inference for query variables without runtime communication. The generated profiles represent less than 1\% of the original dataset triples in most cases, ensuring storage efficiency. Evaluation on the LargeRDFBench benchmark (13 datasets with >1B triples, 32 queries) shows that RENSA achieves source selection results comparable to state-of-the-art methods while eliminating ASK query overhead. Furthermore, we demonstrate that RENSA infers class and authority constraints for query variables, enabling the identification of data sources even across heterogeneous endpoints. These profiles additionally offer human-readable structural insights for semi-automated query generation. 2026-08-29T00:32:13Z Victor Eiti Yamamoto Takeda Hideaki Yamamoto Yasunori http://arxiv.org/abs/2608.28835v1 Engaging the scientific community in high-quality biocuration: a report on the International Society for Biocuration workshop, 'Maximizing community curation for the benefit of all' 2026-08-28T20:14:27Z Biological knowledgebases traditionally rely on expert, professional curation of the research literature to maintain up-to-date collections of data organized in machine-readable form. However, despite the increasing amount of curatable biomedical knowledge, support for knowledgebases is declining, leaving these resources no alternative but to explore additional ways of updating and maintaining content. One way in which knowledgebases have addressed this problem is by engaging researchers to help curate their published papers, a process generally known as 'community curation'. As helpful as community curation can be, though, it is not universally adopted and, for groups that do have it, there is a wide range of approaches. To learn about existing community curation pipelines and explore possibilities for working towards a common approach, we organized a workshop, Maximizing Community Curation for the Benefit of All, at the 18th International Biocuration Conference, hosted by the Stowers Institute for Medical Research. Our aim was to examine the different strategies that groups use, share successes, failures, and ongoing challenges, and produce suggested deliverables for broader adoption of common best practices and tools for effective community curation. Representatives from 18 different resources, ranging from model organism and specialty knowledgebases to journals and literature resources, presented their work. The result was a comprehensive assessment of the state-of-the-art for community curation and an in-depth discussion on how community curation can become standard practice for maintaining timely, highquality biological resources that will continue to provide scientists with the essential information they need for their research. 2026-08-28T20:14:27Z 26 pages, 7 tables. Preprint intended for publication in a journal Daniela Raciti Susan L. M. Coort Christian Grove Jade Hotchkiss Matt Jeffryes Nancy T. Li Zhiyong Lu Bastien Molcrette Sushma Naithani Maria Victoria Nugnes Jolene Ramsey Rene Ranzinger Leonore Reiser Karen E. Ross Garrett Stevens Courtney Thaxton Sabrina Toro Valerie Wood Karen Yook Kimberly Van Auken http://arxiv.org/abs/2608.28432v1 Are These Modules Worth Their Cost? A Paradigm-Level Accuracy-Cost Analysis of In-context Learning Text-to-SQL 2026-08-28T15:13:55Z Recent advances in in-context learning (ICL) text-to-SQL have substantially improved execution accuracy on public benchmarks by assembling increasingly elaborate pipelines around the base generator, yet existing studies typically report aggregate end-to-end accuracy, without quantifying the marginal accuracy-cost contribution of individual design choices. Consequently, providing a unified, paradigm-level cost-accuracy quantification remains a critical challenge for understanding and configuring modern text-to-SQL. To address this, we instantiate 17 paradigm-level configurations across five recurring modules of the ICL text-to-SQL pipeline under a single controlled implementation, and attribute each paradigm's marginal contribution and incurred cost across all four backbones spanning diverse capability levels and reasoning styles. Our analysis reveals that execution-feedback refinement is the only paradigm whose benefit holds universally at consistently low cost, while most other modules help only under backbone-dependent conditions. Token accounting shows that input demand is more closely tied to pipeline structure, whereas output demand is more sensitive to backbone generation behavior. Cross-module analysis further shows that stacking improves accuracy on most backbones, although how the gains compose varies with backbone capability. We also find that a fixed budget is often better spent engineering a more elaborate pipeline over a mid-tier backbone than upgrading to a frontier model with a lean pipeline. These findings distill into an actionable, cost-aware tiered guideline that transfers to five additional backbones without per-paradigm search. 2026-08-28T15:13:55Z Jiayan Lin Yujia Liu Zijin Hong Zheng Yuan Yilin Xiao Hao Chen Qinggang Zhang Xiao Huang Feiran Huang http://arxiv.org/abs/2602.23999v2 GPU-Native Approximate Nearest Neighbor Search with IVF-RaBitQ: Fast Index Build and Search 2026-08-28T14:04:33Z Approximate nearest neighbor search (ANNS) on GPUs is gaining increasing popularity for modern retrieval and recommendation workloads that operate over massive high-dimensional vectors. Graph-based indexes deliver high recall and throughput but incur heavy build-time and storage costs. In contrast, cluster-based methods build and scale efficiently yet often need many probes for high recall, straining memory bandwidth and compute. Aiming to simultaneously achieve fast index build, high-throughput search, high recall, and low storage requirement for GPUs, we present IVF-RaBitQ (GPU), a GPU-native ANNS solution that integrates the cluster-based method IVF with RaBitQ quantization into an efficient GPU index build/search pipeline. Specifically, for index build, we develop a scalable GPU-native RaBitQ quantization method that enables fast and accurate low-bit encoding at scale. For search, we develop GPU-native distance computation schemes for RaBitQ codes and a fused search kernel to achieve high throughput with high recall. With IVF-RaBitQ implemented and integrated into the NVIDIA cuVS Library, experiments on cuVS Bench across multiple datasets show that IVF-RaBitQ offers a strong performance frontier in recall, throughput, index build time, and storage footprint. For Recall approximately equal 0.95, IVF-RaBitQ achieves 3.0x higher QPS than the state-of-the-art graph-based method CAGRA, while also constructing indices 14.7x faster on average. Compared to the cluster-based method IVF-PQ, IVF-RaBitQ delivers on average over 4.5x higher throughput while avoiding accessing the raw vectors for reranking. 2026-02-27T13:23:30Z Jifan Shi Jianyang Gao James Xia Tamás Béla Fehér Cheng Long http://arxiv.org/abs/2608.28352v1 No Silver Bullet: Boosting GaussDB Performance on the 30TB TPC-H Workload 2026-08-28T14:04:04Z GaussDB is Huawei's premier database system, designed for large-scale deployments and the most demanding workloads. It is a distributed shared-nothing system, capable of handling all types of workloads. This paper outlines a series of modifications to GaussDB aimed at improving its performance on large-scale and complex analytical workloads. After these changes, its performance on the TPC-H workload exceeded the best published result by 40% at 30 TB. The key enhancements to achieve this elite performance include adopting a pipeline execution model, a faster and more scalable inter-node data shuffle, exploiting a unified bus and unified remote memory access. We also expanded the support of cost-based Bloom filter placement and implemented several Bloom filter streaming strategies, enabling their use across nodes. 2026-08-28T14:04:04Z Tim Zeyl Jason Lam Shu Lin Reza Pournaghi Qi Cheng Calvin Wong Kaixiang Du Yuliang He Yang Sun Weicheng Wang Paul Lee Chen Ruo Yang Xinyi Li Qunan Wang Junjie Hu Dongxing Chong Chen Per-Ake Larson 10.14778/3827998.3828014 http://arxiv.org/abs/2608.28206v1 NumBench: Diagnosing Counting Failures in Text-to-Image Models 2026-08-28T11:26:19Z Text-to-image (T2I) models often generate the wrong number of objects, yet existing benchmarks are too small or weakly controlled to explain why. We introduce \textbf{NumBench}, a benchmark of 640{,}000 prompts spanning 1{,}600 categories and counts from 1 to 100. Its factorial design varies object composition, spatial guidance, and appearance conditions while balancing counts and category exposure. We also develop a process model in which requested instances compete for a finite set of resolvable image regions. The model predicts a near-quadratic collision deficit at low occupancy and shows how coordinated placement reduces it. For scalable evaluation, we propose the Confidence-Weighted Numeric Precision Score (\cwnps), which aggregates three calibrated detectors and discounts uncertain proposals. Across five commercial systems, two open models, and two specialized counting methods, performance declines sharply with requested count; all evaluated methods are weak above 50 objects. Count range has the largest measured effect, followed by layout and composition. Grid guidance is strongest among guided layouts, consistent with the coordination prediction, although the analysis does not establish collision as the sole cause. A 14{,}400-image human study supports automated evaluation through count 50, while results on 243 natural-language prompts show transfer beyond NumBench templates. 2026-08-28T11:26:19Z The compiled main paper has seven technical-content pages; references start on page 8. The compiled supplement has three pages Sandeep Wadhwa Mayank Vatsa Richa Singh Parrva Chirag Shah Prakhar Galriya http://arxiv.org/abs/2609.29612v1 Towards Quantum Range Query for Spatial-Temporal-Semantic Trajectory Data 2026-08-28T08:33:07Z Range query is a fundamental task in geospatial data search and many other downstream applications. Classic range queries often rely on tree-based spatial indexes, of which the query speed depends on the number of indexed points $k$ within the queried range. For instance, a classical B+ tree answers a range query in O(log N+k). For a long time, this speed has long been considered asymptotically optimal in classic database systems, until the recent emergence of quantum computing, where a quantum B+ tree may requires only O(log_B N). This paper presents Quantum Range Query (QRQ) via a hybrid quantum-classic algorithm to return the range query results in quantum superpositions. In this context, QRQ is designed to accelerate classic range query on spatial-temporal-semantic trajectory geodata using quantum algorithms. Specifically, QRQ develops quantum variants of R-tree, TB-tree and KD-tree, where the physical slots of a node, including unused padding slots, are treated as an array that a quantum random-access memory (QRAM) can read in superpositions. Evaluations on three common trajectory datasets, namely GeoLife, T-Drive, and GDP Drifter, and 10,000 queries per setting, the QRQ speedup at 1% target selectivity ranges from 2.03 times) to 64.70 times. More importantly, QRQ shows a great potential in optimizing existing trajectory range query, so it becomes the shared primitive on which human mobility analysis, and later searches specified by large language models (LLMs) and geo-foundation models (GeoFMs), can rest. 2026-08-28T08:33:07Z Hao Li Zhihang Liu Liwei Zou Jinlin Wu Wufan Zhao http://arxiv.org/abs/2602.18154v3 FENCE: A Financial and Multimodal Jailbreak Detection Dataset 2026-08-28T05:23:02Z Jailbreaking poses a significant risk to the deployment of Large Language Models (LLMs) and Vision Language Models (VLMs). VLMs are particularly vulnerable because they process both text and images, creating broader attack surfaces. However, available resources for jailbreak detection are scarce, particularly in finance. To address this gap, we present FENCE, a bilingual (Korean-English) multimodal dataset for training and evaluating jailbreak detectors in financial applications. FENCE emphasizes domain realism through finance-relevant queries paired with image-grounded threats. Experiments with commercial and open-source VLMs reveal consistent vulnerabilities, with GPT-4o showing measurable attack success rates and open-source models displaying greater exposure. A baseline detector trained on FENCE achieves 99 percent in-distribution accuracy and maintains strong performance on external benchmarks, underscoring the dataset's robustness for training reliable detection models. FENCE provides a focused resource for advancing multimodal jailbreak detection in finance and for supporting safer, more reliable AI systems in sensitive domains. Warning: This paper includes example data that may be offensive. 2026-02-20T11:40:41Z lrec 2026 accepted paper Mirae Kim Seonghun Jeong Youngjun Kwak