https://arxiv.org/api/kEE1zgs4dTKNWmlbE0MExRPN1q02026-10-03T23:22:05Z1246824015http://arxiv.org/abs/2608.29993v1Demystifying and Improving Lazy Promotion in Cache Eviction2026-08-30T19:38:02ZCache eviction algorithms play a critical role in the performance of modern data systems, yet their scalability is often limited by the high computational overhead associated with object promotions. Lazy Promotion techniques have emerged as relaxations of traditional Least-Recently-Used (LRU) methods, designed to alleviate lock contention and increase throughput. This work uses production traces from real-world systems to benchmark five Lazy Promotion strategies: Probabilistic-LRU, Batch-LRU, Delay-LRU, FIFO-reinsertion, and Random-LRU. We evaluate these techniques across miss ratio, scalability, promotion count, and a novel metric called promotion efficiency, which measures the number of hits per promotion.
Our results reveal that Delay-LRU and FIFO-reinsertion significantly improve promotion efficiency, whereas Batch-LRU and Probabilistic-LRU struggle to reduce promotions without significantly increasing miss ratio. We further explore the impact of lazy promotion in advanced algorithms such as ARC and 2Q and make a similar observation. Moreover, we uncover substantial optimization potential, showing that most cache promotions are unnecessary when equipped with oracle knowledge. To further reduce promotions in LRU, we propose two novel enhancements-Delayed FIFO-reinsertion (D-FR) and Age-Guided Eviction (AGE)-that reduce promotions by 20-60% while achieving a similar or lower miss ratio.2026-08-30T19:38:02Z14 pages, 13 figures, and 3 tables. Published in PVLDB 19(4); VLDB 2026 conference paperProceedings of the VLDB Endowment, 19(4): 549-562, 2025Qinghan ChenMuhammad Haekal Muhyidin Al-ArabyZiyue QiuZhuofan ChenRashmi VinayakJuncheng Yang10.14778/3785297.3785299http://arxiv.org/abs/2607.25891v2Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation2026-08-30T18:22:18ZComprehensively evaluating AI agents across interactive environments is difficult due to fragmented tasks, scaffolds, verifiers, and scoring rules. Unfortunately, existing efforts to unify these evaluations are limited in scale and domain, making costly reruns necessary and leaving available data incomparable. We introduce MESSIER, a unified corpus of 957,611 records spanning 30 benchmarks, 745 agents, 11,891 tasks, and 74,263 verifiers. MESSIER combines public evaluation results with new runs on six underrepresented professional and scientific benchmarks, standardizing their heterogeneous components into a common schema. Using this corpus, we show that frontier progress is uneven across benchmark groups, with function-calling evaluations largely saturated, programming improving fastest, and enterprise workflows remaining most challenging. Counterfactual rescoring further shows that strict all-pass scoring in multi-verifier tasks can alter agent rankings. Finally, we derive capability scores from our corpus that correlate with Epoch's Evaluation Capability Index rankings at Spearman \r{ho} = 0.84. The scores can also be estimated for subsets defined by domain, occupation, action space, or verifier type. In essence, MESSIER is a reusable resource for studying agent performance at scale, and a basis for designing better evaluations.2026-07-28T15:50:19ZStefan KrsteskiCharlotte MeyerGuillaume AllegreTony O'HalloranAlexandre Sallinenhttp://arxiv.org/abs/2608.29543v1Evaluating LLMs on Conversational Text-to-SQL under Chain Ambiguity and Intent Drift2026-08-30T04:24:08ZRecent advances in large language models (LLMs) have established conversational text-to-SQL as a practical interface between users and databases, often involving multiple turns of clarification and revision. However, existing benchmarks primarily evaluate execution accuracy, leaving the unfolding and shifting of user intent across turns largely uncovered. To address this, we introduce TIDE-Bench, a benchmark for conversational text-to-SQL under chain ambiguity and intent drift evaluation, targeting two recurring patterns: chain ambiguity, where an underspecified question triggers layered clarification with conditional dependencies, and intent drift, where the user retracts and replaces a previously committed request element. Built on 514 anchor SQLs from BIRD, TIDE-Bench comprises 1,542 samples and introduces dedicated metrics for chain identification and drift recognition-resolution beyond execution accuracy. Evaluating 12 advanced LLMs reveals a persistent chain identification bottleneck unaffected by clarification frequency, a wide drift recognition-resolution gap, and overlap between failure modes when jointly activated. The corresponding code of TIDE-Bench is released for further research.2026-08-30T04:24:08ZAccepted to EMNLP2026 MainYujia LiuJiayan LinZijin HongZheng YuanShengyuan ChenHao ChenQinggang ZhangXiao HuangFeiran Huanghttp://arxiv.org/abs/2609.29661v1Ingest-Time Fact Compilation for Cost-Efficient and Reliable Question Answering over Revised Corpora2026-08-29T23:04:37ZMost agentic question answering (QA) systems do an important part of their semantic work at the worst possible time: every time someone asks a question. When a corpus contains revisions, drafts, revocations, deletions, and sources with different levels of authority, the model must reconstruct the governed current state on every read - then throw that work away and repeat it on the next query. This is a bit like a database that rebuilds a materialized view every time someone reads from it. We present ingest-time fact compilation, an architecture that performs this work when corpus data is ingested or changed. Raw passages are rephrased into self-contained facts; rules governing revisions, deletions, effective dates, and source trust are resolved once; and the resulting state is stored as typed records carrying source and revision provenance. At query time, an inexpensive model reads the compiled record instead of reconstructing it from noisy candidates. In a controlled synthetic experiment across five seeds, the same low-cost model produced the correct value, source, and revision in only one of 30 trials under query-time reconstruction, but in all 30 trials from the compiled substrate, at 12.89 times lower mean read cost per question. On simpler revision questions both architectures were exact, but the compiled path used 21.6 times fewer tokens. A separate test found that fact rephrasing roughly halved verbose Federal Reserve dialogue while preserving high source entailment, but left concise Wikipedia prose essentially unchanged. These results support a narrow but practical claim: resolving a corpus state once can make subsequent QA cheaper and more reliable for inexpensive models. We release the open source, MIT-licensed implementation and experimental artifacts.2026-08-29T23:04:37Z6 pages, 5 tables, 1 figure. Accepted for presentation at the 2026 International Conference on Applied Science and Technology - Engineering Science (iCAST-ES 2026), Surabaya, Indonesia, October 2026. Code and frozen experimental artifacts: https://github.com/aix-sc/isc (tag data-freeze-2026-07-16)Kyle WildYusuke TakahashiAsako Urakihttp://arxiv.org/abs/2607.20489v2EvoSQL: Memory-Augmented Critic-Generator Co-Evolution for Text-to-SQL2026-08-29T16:48:45ZText-to-SQL has advanced rapidly with large language models, but complex database queries still require reasoning beyond one-shot generation, including multi-step decomposition, execution-based diagnosis, and targeted correction. We present EvoSQL, a co-evolution framework that formulates SQL synthesis as an iterative interaction between a generator and a critic. EvoSQL maintains a contextualized candidate memory, verifies SQL candidates with both execution signals and LLM-based critique, and updates its memory through utility-guided aggregation. To strengthen the underlying generator-critic pair, we further introduce a Self-Distillation Policy Optimization (SDPO) fine-tuning stage that injects execution-aware supervision into modern coding LLM backbones. Experiments on Spider and BIRD show that EvoSQL consistently improves open-source models over Maj@16 baselines, with particularly large gains on BIRD-Dev, ranging from +1.37% for Qwen3-4B to +9.19% for Qwen2.5-Coder-3B. SDPO initialization further improves selected backbones on Spider-Test and BIRD-Dev. These results suggest that memory-grounded co-evolution is an effective path toward more reliable and generalizable Text-to-SQL systems. Code is available at https://github.com/valleysprings/EvoSQL.2026-06-04T14:21:19ZAccepted as EMNLP Findings (2026)Jiawei ZhouJianwei WangChenyu ZhouChaojian ShiMing DongKai Wanghttp://arxiv.org/abs/2509.08395v5SINDI: An Efficient Index for Sparse Vector Approximate Maximum Inner Product Search2026-08-29T03:18:40ZSparse vector Maximum Inner Product Search (MIPS) is crucial in multi-path retrieval for Retrieval-Augmented Generation (RAG). Recent inverted index-based and graph-based algorithms have achieved high search accuracy with practical efficiency. However, their performance in production environments is often limited by redundant distance computations and frequent random memory accesses. Furthermore, the compressed storage format of sparse vectors hinders the use of SIMD acceleration. In this paper, we propose the sparse inverted non-redundant distance index (SINDI), which incorporates three key optimizations: (i) Efficient Inner Product Computation: SINDI leverages SIMD acceleration and eliminates redundant identifier lookups, enabling batched inner product computation; (ii) Memory-Friendly Design: SINDI replaces random memory accesses to original vectors with sequential accesses to inverted lists, substantially reducing memory-bound latency. (iii) Vector Pruning: SINDI retains only the high-magnitude non-zero entries of vectors, improving query throughput while maintaining accuracy. We evaluate SINDI on multiple real-world datasets. Experimental results show that SINDI achieves state-of-the-art performance across datasets of varying scales, languages, and models. On the MsMarco dataset, when Recall@50 exceeds 99%, SINDI delivers single-thread query-per-second (QPS) improvements ranging from 4.2 to 26.4 times compared with SEISMIC and PyANNs. Notably, SINDI has been integrated into Ant Group's open-source vector search library, VSAG.2025-09-10T08:38:32Z18 pages, accepted by ICDE 2026. Due to submission limitation for ICDE 2026 (i.e., maximum 6 submissions per author), Lei Chen and Xuemin Lin are not included as authorsRuoxuan LiXiaoyao ZhongJiabao JinPeng ChengWangze NiZhitao ShenWei JiaXiangyu WangHeng Tao ShenJingkuan Songhttp://arxiv.org/abs/2606.23667v2The Table Says Otherwise: Testing LLMs with Counterfactual Relational Data2026-08-29T01:27:58ZLarge language models (LLMs) are increasingly used to answer natural-language questions over structured data. However, when a table contains familiar real-world facts, it is unclear whether the model answers by reading the provided data or by recalling knowledge learned during pretraining. This distinction is important for database applications, where the provided tables should be the source of truth. In this paper, we introduce ContraTable, a paired original-counterfactual benchmark for evaluating whether LLMs ground their answers in relational tables. We build the benchmark with two aligned versions: an original database with real-world facts and a counterfactual database that preserves the same schemas, identifiers, and relationships while changing selected country, club, and player attributes. We design 214 matched questions across three levels: single-table lookup, multi-table lookup, and multi-table temporal reasoning. Experiments on commercial closed-source and open-source models show that strong instruction-tuned models can often handle direct lookup, but their reliability drops as questions require joins, comparison, and temporal reasoning. The gap between original and counterfactual accuracy reveals that models may fall back on prior knowledge when table evidence conflicts with familiar facts. These results suggest that table-QA evaluation should measure not only accuracy, but also faithfulness to the provided database.2026-06-22T17:52:14ZXinzhi WangChunwei Liuhttp://arxiv.org/abs/2608.28963v1RENSA: Rich Environment Metadata to Navigate Shared and Distributed Endpoints for Automated Federated SPARQL Query Generation2026-08-29T00:32:13ZThe number of knowledge graph databases has increased significantly with the proliferation of knowledge graph technologies. Knowledge graphs enable the dynamic integration of distributed data through federated SPARQL queries. However, constructing efficient queries in a federated environment is challenging due to the lack of detailed structural knowledge across decentralized datasets. While standards like VoID provide basic metadata, they often fail to capture the complex interlinks and authority distributions necessary for optimization. Consequently, current engines frequently rely on runtime ASK queries for source selection, increasing communication overhead. We propose RENSA, a federated SPARQL query generation framework that leverages an extension of SPARQL Builder Metadata (SBM). By integrating class and authority information, mapping subject and object usage to specific predicates, RENSA enables precise source selection and semantic constraint inference for query variables without runtime communication. The generated profiles represent less than 1\% of the original dataset triples in most cases, ensuring storage efficiency. Evaluation on the LargeRDFBench benchmark (13 datasets with >1B triples, 32 queries) shows that RENSA achieves source selection results comparable to state-of-the-art methods while eliminating ASK query overhead. Furthermore, we demonstrate that RENSA infers class and authority constraints for query variables, enabling the identification of data sources even across heterogeneous endpoints. These profiles additionally offer human-readable structural insights for semi-automated query generation.2026-08-29T00:32:13ZVictor Eiti YamamotoTakeda HideakiYamamoto Yasunorihttp://arxiv.org/abs/2608.28835v1Engaging the scientific community in high-quality biocuration: a report on the International Society for Biocuration workshop, 'Maximizing community curation for the benefit of all'2026-08-28T20:14:27ZBiological knowledgebases traditionally rely on expert, professional curation of the research literature to maintain up-to-date collections of data organized in machine-readable form. However, despite the increasing amount of curatable biomedical knowledge, support for knowledgebases is declining, leaving these resources no alternative but to explore additional ways of updating and maintaining content. One way in which knowledgebases have addressed this problem is by engaging researchers to help curate their published papers, a process generally known as 'community curation'. As helpful as community curation can be, though, it is not universally adopted and, for groups that do have it, there is a wide range of approaches. To learn about existing community curation pipelines and explore possibilities for working towards a common approach, we organized a workshop, Maximizing Community Curation for the Benefit of All, at the 18th International Biocuration Conference, hosted by the Stowers Institute for Medical Research. Our aim was to examine the different strategies that groups use, share successes, failures, and ongoing challenges, and produce suggested deliverables for broader adoption of common best practices and tools for effective community curation. Representatives from 18 different resources, ranging from model organism and specialty knowledgebases to journals and literature resources, presented their work. The result was a comprehensive assessment of the state-of-the-art for community curation and an in-depth discussion on how community curation can become standard practice for maintaining timely, highquality biological resources that will continue to provide scientists with the essential information they need for their research.2026-08-28T20:14:27Z26 pages, 7 tables. Preprint intended for publication in a journalDaniela RacitiSusan L. M. CoortChristian GroveJade HotchkissMatt JeffryesNancy T. LiZhiyong LuBastien MolcretteSushma NaithaniMaria Victoria NugnesJolene RamseyRene RanzingerLeonore ReiserKaren E. RossGarrett StevensCourtney ThaxtonSabrina ToroValerie WoodKaren YookKimberly Van Aukenhttp://arxiv.org/abs/2608.28432v1Are These Modules Worth Their Cost? A Paradigm-Level Accuracy-Cost Analysis of In-context Learning Text-to-SQL2026-08-28T15:13:55ZRecent advances in in-context learning (ICL) text-to-SQL have substantially improved execution accuracy on public benchmarks by assembling increasingly elaborate pipelines around the base generator, yet existing studies typically report aggregate end-to-end accuracy, without quantifying the marginal accuracy-cost contribution of individual design choices. Consequently, providing a unified, paradigm-level cost-accuracy quantification remains a critical challenge for understanding and configuring modern text-to-SQL. To address this, we instantiate 17 paradigm-level configurations across five recurring modules of the ICL text-to-SQL pipeline under a single controlled implementation, and attribute each paradigm's marginal contribution and incurred cost across all four backbones spanning diverse capability levels and reasoning styles. Our analysis reveals that execution-feedback refinement is the only paradigm whose benefit holds universally at consistently low cost, while most other modules help only under backbone-dependent conditions. Token accounting shows that input demand is more closely tied to pipeline structure, whereas output demand is more sensitive to backbone generation behavior. Cross-module analysis further shows that stacking improves accuracy on most backbones, although how the gains compose varies with backbone capability. We also find that a fixed budget is often better spent engineering a more elaborate pipeline over a mid-tier backbone than upgrading to a frontier model with a lean pipeline. These findings distill into an actionable, cost-aware tiered guideline that transfers to five additional backbones without per-paradigm search.2026-08-28T15:13:55ZJiayan LinYujia LiuZijin HongZheng YuanYilin XiaoHao ChenQinggang ZhangXiao HuangFeiran Huanghttp://arxiv.org/abs/2602.23999v2GPU-Native Approximate Nearest Neighbor Search with IVF-RaBitQ: Fast Index Build and Search2026-08-28T14:04:33ZApproximate nearest neighbor search (ANNS) on GPUs is gaining increasing popularity for modern retrieval and recommendation workloads that operate over massive high-dimensional vectors. Graph-based indexes deliver high recall and throughput but incur heavy build-time and storage costs. In contrast, cluster-based methods build and scale efficiently yet often need many probes for high recall, straining memory bandwidth and compute. Aiming to simultaneously achieve fast index build, high-throughput search, high recall, and low storage requirement for GPUs, we present IVF-RaBitQ (GPU), a GPU-native ANNS solution that integrates the cluster-based method IVF with RaBitQ quantization into an efficient GPU index build/search pipeline. Specifically, for index build, we develop a scalable GPU-native RaBitQ quantization method that enables fast and accurate low-bit encoding at scale. For search, we develop GPU-native distance computation schemes for RaBitQ codes and a fused search kernel to achieve high throughput with high recall. With IVF-RaBitQ implemented and integrated into the NVIDIA cuVS Library, experiments on cuVS Bench across multiple datasets show that IVF-RaBitQ offers a strong performance frontier in recall, throughput, index build time, and storage footprint. For Recall approximately equal 0.95, IVF-RaBitQ achieves 3.0x higher QPS than the state-of-the-art graph-based method CAGRA, while also constructing indices 14.7x faster on average. Compared to the cluster-based method IVF-PQ, IVF-RaBitQ delivers on average over 4.5x higher throughput while avoiding accessing the raw vectors for reranking.2026-02-27T13:23:30ZJifan ShiJianyang GaoJames XiaTamás Béla FehérCheng Longhttp://arxiv.org/abs/2608.28352v1No Silver Bullet: Boosting GaussDB Performance on the 30TB TPC-H Workload2026-08-28T14:04:04ZGaussDB is Huawei's premier database system, designed for large-scale deployments and the most demanding workloads. It is a distributed shared-nothing system, capable of handling all types of workloads. This paper outlines a series of modifications to GaussDB aimed at improving its performance on large-scale and complex analytical workloads. After these changes, its performance on the TPC-H workload exceeded the best published result by 40% at 30 TB.
The key enhancements to achieve this elite performance include adopting a pipeline execution model, a faster and more scalable inter-node data shuffle, exploiting a unified bus and unified remote memory access. We also expanded the support of cost-based Bloom filter placement and implemented several Bloom filter streaming strategies, enabling their use across nodes.2026-08-28T14:04:04ZTim ZeylJason LamShu LinReza PournaghiQi ChengCalvin WongKaixiang DuYuliang HeYang SunWeicheng WangPaul LeeChen RuoYang XinyiLi QunanWang JunjieHu DongxingChong ChenPer-Ake Larson10.14778/3827998.3828014http://arxiv.org/abs/2608.28206v1NumBench: Diagnosing Counting Failures in Text-to-Image Models2026-08-28T11:26:19ZText-to-image (T2I) models often generate the wrong number of objects, yet existing benchmarks are too small or weakly controlled to explain why. We introduce \textbf{NumBench}, a benchmark of 640{,}000 prompts spanning 1{,}600 categories and counts from 1 to 100. Its factorial design varies object composition, spatial guidance, and appearance conditions while balancing counts and category exposure. We also develop a process model in which requested instances compete for a finite set of resolvable image regions. The model predicts a near-quadratic collision deficit at low occupancy and shows how coordinated placement reduces it. For scalable evaluation, we propose the Confidence-Weighted Numeric Precision Score (\cwnps), which aggregates three calibrated detectors and discounts uncertain proposals. Across five commercial systems, two open models, and two specialized counting methods, performance declines sharply with requested count; all evaluated methods are weak above 50 objects. Count range has the largest measured effect, followed by layout and composition. Grid guidance is strongest among guided layouts, consistent with the coordination prediction, although the analysis does not establish collision as the sole cause. A 14{,}400-image human study supports automated evaluation through count 50, while results on 243 natural-language prompts show transfer beyond NumBench templates.2026-08-28T11:26:19ZThe compiled main paper has seven technical-content pages; references start on page 8. The compiled supplement has three pagesSandeep WadhwaMayank VatsaRicha SinghParrva Chirag ShahPrakhar Galriyahttp://arxiv.org/abs/2609.29612v1Towards Quantum Range Query for Spatial-Temporal-Semantic Trajectory Data2026-08-28T08:33:07ZRange query is a fundamental task in geospatial data search and many other downstream applications. Classic range queries often rely on tree-based spatial indexes, of which the query speed depends on the number of indexed points $k$ within the queried range. For instance, a classical B+ tree answers a range query in O(log N+k). For a long time, this speed has long been considered asymptotically optimal in classic database systems, until the recent emergence of quantum computing, where a quantum B+ tree may requires only O(log_B N). This paper presents Quantum Range Query (QRQ) via a hybrid quantum-classic algorithm to return the range query results in quantum superpositions. In this context, QRQ is designed to accelerate classic range query on spatial-temporal-semantic trajectory geodata using quantum algorithms. Specifically, QRQ develops quantum variants of R-tree, TB-tree and KD-tree, where the physical slots of a node, including unused padding slots, are treated as an array that a quantum random-access memory (QRAM) can read in superpositions. Evaluations on three common trajectory datasets, namely GeoLife, T-Drive, and GDP Drifter, and 10,000 queries per setting, the QRQ speedup at 1% target selectivity ranges from 2.03 times) to 64.70 times. More importantly, QRQ shows a great potential in optimizing existing trajectory range query, so it becomes the shared primitive on which human mobility analysis, and later searches specified by large language models (LLMs) and geo-foundation models (GeoFMs), can rest.2026-08-28T08:33:07ZHao LiZhihang LiuLiwei ZouJinlin WuWufan Zhaohttp://arxiv.org/abs/2602.18154v3FENCE: A Financial and Multimodal Jailbreak Detection Dataset2026-08-28T05:23:02ZJailbreaking poses a significant risk to the deployment of Large Language Models (LLMs) and Vision Language Models (VLMs). VLMs are particularly vulnerable because they process both text and images, creating broader attack surfaces. However, available resources for jailbreak detection are scarce, particularly in finance. To address this gap, we present FENCE, a bilingual (Korean-English) multimodal dataset for training and evaluating jailbreak detectors in financial applications. FENCE emphasizes domain realism through finance-relevant queries paired with image-grounded threats. Experiments with commercial and open-source VLMs reveal consistent vulnerabilities, with GPT-4o showing measurable attack success rates and open-source models displaying greater exposure. A baseline detector trained on FENCE achieves 99 percent in-distribution accuracy and maintains strong performance on external benchmarks, underscoring the dataset's robustness for training reliable detection models. FENCE provides a focused resource for advancing multimodal jailbreak detection in finance and for supporting safer, more reliable AI systems in sensitive domains. Warning: This paper includes example data that may be offensive.2026-02-20T11:40:41Zlrec 2026 accepted paperMirae KimSeonghun JeongYoungjun Kwak