https://arxiv.org/api/w3AD3Y0l38LWBL69itjFPg4PDgo 2026-10-05T02:43:23Z 12483 270 15 http://arxiv.org/abs/2608.25210v2 Bolt-on, Verifiable Provenance for LLM-Powered Data Processing 2026-08-28T03:03:25Z Large Language Models (LLMs) are powerful tools for processing data. However, LLMs are also complex black-boxes, returning answers to queries on data, without any indication for where the answer came from or whether it is trustworthy. We introduce the notion of provenance for data processing with LLMs. While existing heuristics (such as embedding similarity or directly asking an LLM) could provide some hints for where the answer was derived, they provide no guarantees that the answer can be derived using the identified provenance, and indeed, are often incorrect. Instead, we propose the notion of verifiable provenance wherein we identify a subset of the input text that reproduces the same (or equivalent) answer as that on the complete text, and introduce the notion of minimality, where the verifiable provenance is as small as possible. To identify such a provenance, a naive solution would require checking all possible subsets of the source data with the LLM, which is prohibitively expensive. We present BLIP, a bolt-on framework for efficiently inferring a small-sized verifiable provenance for any LLM-powered data processing task, with any LLM. As part of BLIP, we introduce eight strategies, each guaranteed to find a minimal verifiable provenance, as well as an adaptive strategy that combines their strengths to reduce cost further. We further extend BLIP to produce multiple minimal verifiable provenances. Experiments on seven datasets show that the provenance generated by BLIP is always guaranteed to reproduce the answer, achieving over 30% higher accuracy than the best-performing baseline with a comparable provenance size. Moreover, BLIP incurs a low cost, comparable to the original query on the original data. 2026-08-25T22:59:08Z Yiming Lin Sepanta Zeighami Aditya G. Parameswaran http://arxiv.org/abs/2608.27822v1 DBRepro: Automated Database Synthesis via a Hybrid Constraint-Solving Approach for Reproducing Slow Queries 2026-08-28T01:41:18Z Slow queries frequently cause severe performance bottlenecks in database management systems. Diagnosing their root causes online risks exacerbating resource contention, while data privacy regulations often prohibit copying production data to test environments. Synthesizing a proxy database from non-intrusive metadata that induces the query optimizer to generate the same physical execution plans is therefore critical for offline diagnosis. High-fidelity reproduction requires preserving global statistical distributions while enforcing exact local cardinalities. Existing data-driven and workload-aware approaches cannot satisfy both requirements simultaneously. We present DBRepro, an automated end-to-end framework that formulates database generation as a constrained distribution synthesis problem. DBRepro initializes a global distribution from lightweight column statistics, extracts execution constraints from target queries, and progressively adjusts the distribution to satisfy these constraints while preserving the global distribution. Experiments on TPC-H and SSB show that DBRepro reduces cardinality error by up to 20.3% over a data-driven baseline while maintaining identical plan consistency. Compared with a workload-aware baseline, it reproduces 15% more consistent execution plans and reduces latency proportion error by 21.5%. We further validate DBRepro on a nearly 1 TB real-world dataset managed by KingbaseES, where it reproduces the execution performance of complex slow queries with high fidelity. 2026-08-28T01:41:18Z 13 pages, 10 figures, and 3 tables. Accepted for publication in the Industry Showcase Track of the 41st IEEE/ACM International Conference on Automated Software Engineering (ASE 2026) Zhaoyang Zhang Shuang Liu Dengfeng Xu Wei Lu Jianquan Leng Sheng Du Xiaoyong Du http://arxiv.org/abs/2608.27819v1 ANCHOR: A Vision for Secure Persistent Key-Value Stores in Disaggregated Data Centers 2026-08-28T01:37:51Z Persistent key-value stores (PKVS) are increasingly deployed in disaggregated settings that split compute, memory, and storage across separate server pools. This shift redraws the trust boundary: data that would remain within a single machine is now transported, cached, and rewritten across multiple hosts, expanding exposure to both network attackers and intra-infrastructure adversaries. This paper presents ANCHOR, a vision for end-to-end integrity and freshness in disaggregated PKVS. ANCHOR proposes a two-part semantics-aware architecture: 1) Persistence path: ANCHOR outlines encrypting and authenticating PKVS persistent files and preventing rollback with manifest versioning. 2) Volatile path: ANCHOR treats caches, indexes, and filters as untrusted hints unless accompanied by verifiable provenance, enforced by a TEE-resident policy. Finally, we outline key invariants and discuss enclave-friendly batching and asynchronous I/O to amortize verification without undermining disaggregation's performance and elasticity benefits. 2026-08-28T01:37:51Z Viraj Thakkar Dongha Kim Hokeun Kim Zhichao Cao 10.1145/3807894.3810275 http://arxiv.org/abs/2608.27790v1 Credo: Reusable Declarative Primitives for Agentic Workflows 2026-08-28T00:06:36Z An LLM application depends on both a model and a harness: the program that determines what each call sees, how many calls to make, and which answers to trust. Coding agents can now discover strong harnesses by searching over candidate programs, but the resulting artifact is an opaque block of imperative code whose logical steps, runtime signals, physical execution decisions, and prompt strategies remain implicit and task-specific, forcing subsequent tasks to start the harness search process from scratch. The potential for reuse, however, is substantial. A searched harness encodes significant knowledge, such as the logical steps that work, the signals that matter, the physical operator decisions that adapt execution, and the prompt strategies that are effective, yet this knowledge is buried in imperative code with no inspectable or reusable structure, nor does it carry any provenance or metadata. Credo addresses this problem by recovering a structured declarative description of a searched harness, tagging each extracted primitive with relevant metadata, and cataloguing all of it with provenance. A compiler can then bind stored primitives to generate harnesses for new tasks without having to start the search over from scratch. This paper provides preliminary results demonstrating the potential of our approach and lays out a related research agenda that the database community is well-positioned to tackle, including cost-based compilation over declarative catalogs and catalog maintenance under model and workload drift. 2026-08-28T00:06:36Z Duo Lu Andrew Crotty Uğur Çetintemel http://arxiv.org/abs/2608.27758v1 Real-time SQL Plan Management in Oracle 2026-08-27T22:42:58Z Consistent query performance is essential for mission critical database applications, yet SQL execution plans can change due to factors such as database upgrades, DML changes, new indexes, etc. While plan stability mechanisms such as stored outlines prevent regressions by freezing execution plans, they also inhibit performance improvements by disallowing plan evolution. We introduced SQL Plan Management (SPM) in Oracle 11g to address this trade-off by maintaining a set of accepted execution plans and allowing plan evolution only when new plans demonstrably outperform existing baselines. However, prior implementations of SPM primarily rely on background performance verification processes, delaying regression detection and recovery. This issue is amplified in autonomous cloud database systems, where several automatic actions that could cause plan change driven regressions are performed with limited customer control. Timely detection and remediation is paramount, but the constrained background resources on cloud may not keep pace. To overcome these limitations, we introduce Real-Time SPM in Oracle 26ai, a novel extension of SPM that performs foreground verification of new execution plans during user query execution. Real-Time SPM leverages runtime session context to immediately validate plan changes, enabling rapid adoption of superior plans while promptly detecting and preventing regressions. This paper presents the architecture and design of Real-Time SPM - including technical challenges like reliably comparing performance of previous plans - and contrasts it with traditional background plan evolution. 2026-08-27T22:42:58Z Sunil Chakkappen Mohamed Ziauddin Hong Su Shreya Kunjibettu Nigel Bayliss http://arxiv.org/abs/2608.25061v2 DataKernelBench: Can LLMs Optimize Database Queries on GPUs? 2026-08-27T17:07:30Z GPUs increasingly accelerate database systems, but query-specific peak performance still often relies on hand-written kernels. Existing LLM kernel benchmarks focus on machine learning operators, leaving irregular, heterogeneous, data-movement-heavy database-style operators untested. We introduce DataKernelBench, which translates SQL into validated PyTorch TorchPlan programs and evaluates LLMs that optimize either the core tensor-bounded snippet or the full query in CUDA or Triton through execution-guided repair. Across ten proprietary and open-weight models on TPC-H SF10 with an H100 GPU, the strongest full-query CUDA configuration achieves $2.11\times$ speedup over the TorchPlan baseline at full pass rate. We find that higher-performing implementations commonly use kernel fusion and execution-strategy changes, stronger models benefit most from full-query specialization, and workload context matters more than hardware context. To handle data larger than GPU memory, we extend TorchPlan with Dask-cuDF for on-demand partition loading on TPC-H SF100 with four H100 GPUs, achieving $2.54\times$ speedup. Project page: https://kerneldf.github.io/datakernelbench 2026-08-25T18:57:39Z Accepted at EMNLP 2026 Gokul Karthik Kumar Yotam Perlitz Corey Lammie Andrea Giovannini Katja Hose http://arxiv.org/abs/2608.27244v1 Compositional Online Learning for Semantic Data Processing Systems 2026-08-27T15:25:05Z An LLM call in a semantic data processing system is expensive enough to dominate query cost, yet slow enough to hide a CPU-side learner's update behind its round-trip. In production, LLM compute accounts for $80-90\%$ of query cost, and each call costs $10^5-10^7\times$ a relational predicate. The latency window inverts a design constraint of classical adaptive query processing, where online learners had to stay lightweight to avoid dominating the predicates they optimize. At LLM latency, per-call gradient steps and per-batch threshold solves fit inside the round-trip. We develop compositional online learning at the LLM call boundary: a framework for combining online-learning components in semantic data processing systems. Each component makes execution-time decisions and refines its learned artifacts online. The design space spans two axes, decision granularity and learner update cadence, and the components share a single learning pattern that hides each trainer step inside the next LLM round-trip. A production case study in Cortex AISQL composes three components: a memoization layer, an online per-call filter-ordering learner, and an online per-batch cascade-routing learner. A conditional cost decomposition assigns each learning component to a distinct factor of per-row LLM cost. Under independence, the two learning components compose multiplicatively to an $11.4\times$ upper bound on a representative conjunction-filter workload. Self-selection at the cascade boundary, sample-budget shrinkage, and selectivity-estimation drift reduce it to a realistic figure near $8\times$. 2026-08-27T15:25:05Z Paweł Liskowski Fuheng Zhao Benjamin Han Anupam Datta Dimitris Tsirogiannis http://arxiv.org/abs/2608.07945v2 ScaleSense: Cost-Intelligent Scaling Framework via Learned Resource Estimation in Alibaba AnalyticDB 2026-08-27T10:56:54Z Cloud-native serverless data warehouses achieve fine-grained elasticity by decoupling storage from compute, yet determining the optimal resource allocation for highly heterogeneous ad-hoc queries remains a formidable industrial challenge. Our analysis of production workloads in Alibaba AnalyticDB exposes a costly ``provisioning trap'': the fear of catastrophic resource depletion drives users to blindly over-provision resources, wasting immense monetary budgets without alleviating non-CPU bottlenecks (e.g., I/O saturation). To break this impasse, we propose ScaleSense, a proactive, query-level resource scaling framework. Specifically, it features a multi-faceted query encoder that jointly models plan topologies and hardware specifications. Crucially, a quantile-based resource predictor estimates multi-dimensional physical footprints, acting as a reliable safety net for optimal resource scaling. An auto-scaling controller then navigates the performance-cost Pareto frontier, dynamically tailoring allocations to specific business priorities without requiring model retraining. Evaluations on over 1.36 million production queries show that ScaleSense achieves state-of-the-art prediction accuracy with good prediction interval coverage. By achieving a 76.7% relative improvement in optimal resource configuration selection over the best baseline, this approach addresses the critical performance-cost trade-off while maintaining low-overhead inference latency, confirming its practical performance in production deployments. Under the performance-optimization policy, ScaleSense satisfies user-defined performance requirements while reducing monetary cost by up to 5.22x. 2026-08-08T06:10:35Z This paper has been accepted for presentation at VLDB 2026 Yifan Wu Yuhan Li Zhenhua Wang Ke Chen Lidan Shou Zonghao Chen Liang Lin Huan Li Gang Chen http://arxiv.org/abs/2608.26930v1 Incremental Delta-Shapley: A Standalone Runtime for Predicate Attribution on Sliding Windows 2026-08-27T10:30:21Z Continuous aggregate queries over sliding windows are common in real-time analytics, but most systems report \emph{what} an aggregate is doing without attributing \emph{which} predicates account for the result. A companion paper~\cite{khani2026closedformpredicatelevelshapleyattribution} shows that exact predicate-level Shapley attribution for SUM, COUNT, AVG, and variance needs only three additive predicate summaries with closed-form coefficients. Those results settle the mathematics, not how a runtime maintains summaries across slides, exposes attribution, answers unregistered predicates, or amortizes repeated ad hoc ones. We present \textbf{IDS} (Incremental Delta-Shapley), a standalone single-node runtime that turns those closed forms into a deployable explanation system. IDS consumes window-maintenance deltas, updates global, marginal, and atom summaries, and evaluates any closed form in constant time. Overlapping predicates use atomic refinement, and a restricted SQL-like API exposes attribution and its per-slide change as first-class operators. Unregistered predicates are answered by a retained-state scan, an inverted index, or an amortized sliding-window sample with concentration guarantees; frequent ones are promoted by rebuilding the refinement. On synthetic, adversarial, NEXMark-style, and NYC taxi workloads, attribution matches exhaustive Shapley enumeration to floating-point precision; incremental maintenance is flat in $N$ and up to $4.3\times10^{5}\times$ faster than per-window scans of the same form; and adaptive promotion cuts ad hoc cost by up to $9.2\times$ on Zipfian traces. 2026-08-27T10:30:21Z Pouya Khani Ira Assent http://arxiv.org/abs/2608.26775v1 Size Bounds for CQs Under Acyclic Constraints 2026-08-27T08:06:14Z We study size bounds for conjunctive query (CQ) results which in recent years have played a crucial role in database theory. In particular, we compare the so-called entropic bound which is known to be asymptotically tight but not known to be computable, and its computable relaxation called the polymatroid bound, which is generally not tight. We focus here on conjunctive queries under acyclic functional dependencies. These queries are known to be well-behaved in the sense that, in the case without projections, both bounds coincide. We show that this picture changes when projections are allowed: in this case, even for acyclic functional dependencies, there is in general a polynomial gap between the two bounds. We complement this negative result by showing a special case for which the polymatroid bound is tight under acyclic functional dependencies and projections that is characterized by the position of the output variables in a topological order of the query variables. 2026-08-27T08:06:14Z Stefan Mengel Andrei Romashchenko http://arxiv.org/abs/2606.20318v4 AgenticDB: Self-Evolving Reconfiguration Framework for Database Workloads 2026-08-27T06:40:46Z Configuration tuning is critical to database performance but remains difficult in real deployments. Despite notable advances, prior methods still leave substantial performance potential unexplored, suffer from low tuning efficiency, and provide limited support for configuration validation and failure recovery. To address these limitations, we propose AgenticDB, a self-evolving agentic framework for database workload reconfiguration. AgenticDB uses a large language model (LLM)-based DBA Planner to jointly reconfigure database knobs and operating system (OS) parameters through two key mechanisms. First, context-grounded bottleneck diagnosis uses workload characteristics, configuration state, and observed runtime behavior to identify the current performance bottleneck and recommend targeted database management system (DBMS)/OS reconfiguration actions. Second, closed-loop context evolution uses observed performance and runtime-state changes as feedback to update the bottleneck diagnosis, guide subsequent decisions, and terminate the reconfiguration loop when performance plateaus. It also consolidates accumulated reconfiguration experience for reuse on workloads with similar characteristics. Beyond these two mechanisms, AgenticDB improves reliability by validating each proposed configuration before applying it and automatically recovering from failures. We evaluate AgenticDB on MySQL and PostgreSQL using YCSB, Sysbench, and TPC-H. Compared with SOTA methods, AgenticDB outperforms the best-performing baseline by 118.1% on average and reduces the total time-to-best across workloads by 22.6%. Further analyses show that validation and recovery improve reconfiguration reliability. The consolidated experience also helps AgenticDB reach high-performing configurations earlier on workloads with similar characteristics. 2026-06-18T14:57:10Z Xinyue Yang Chaozheng Wang Chen Zheng Heng Zhang Yanjun Wu http://arxiv.org/abs/2608.26464v1 VoS: Variate Ordering Strategies for Skyline Query Optimization 2026-08-26T23:37:09Z Efficiency of skyline algorithms is highly influenced by the underlying data characteristics. Traditionally, optimization efforts have focused on minimizing the total number of tuple-pair dominance checks to improve query performance. However, in practice, a dominance check between two tuples does not necessarily require evaluating dominance relationships for each and every preference attribute of the data and this creates a disconnect between dominance checks optimization and query execution performance. In this paper, we argue that skyline algorithms need to optimize total per-attribute dominance checks, along with per-tuple dominance checks and that, for both of these goals, the ordering of the attributes (or variates) can have a substantial impact on the efficiency of skyline computation. Based on this premise, we present several strategies for identifying an effective variate order to minimize redundant attribute comparisons. Extensive experiments on both synthetic and real-world datasets, and on both scalar and SIMD architectures, confirm the effectiveness of the proposed approach in reducing computational overhead and improving skyline query performance. 2026-08-26T23:37:09Z Abhinav Gorantla Pratanu Mandal K. Selçuk Candan Maria Luisa Sapino http://arxiv.org/abs/2606.19751v2 DeQL: A Decision Query Language for Prescriptive Analytics over Relational Data 2026-08-26T20:21:27Z DeQL (Decision Query Language) extends SQL to express decision queries: given options drawn from relational data, constraints from policy, and a measurable objective, a DeQL query computes the best course of action. Two constructs carry the core of the extension: CREATE CANDIDATES, which defines the space of options from relational sources, and DECIDE, which declares decision variables, named constraints, and an objective over them. The design follows SQL's principles: the user states what to optimize while the engine chooses how to solve it, every query consumes and produces relations, and the structure of a problem stays visible to the engine. This document specifies the language: its design principles, syntax, formal grammar, and execution model. The examples span subset selection, allocation, assignment, scheduling, and decisions at multiple levels of aggregation. The specification also covers extensions for optimization under uncertainty, inline model scoring, and time- and quality-bounded solving. It is an early version of the specification; the language is under active development, and this version fixes the core constructs on which later revisions will build. 2026-06-18T03:31:28Z 33 pages, 5 tables, no figures. Version 0.2. Corrects several rules stated incorrectly in version 0.1 (SELECTION desugaring scope, auto-multiply over MIN, MAX and AVG, minimax objective directions, the PaQL REPEAT translation, and the grain of scheduling constraints) and closes several gaps in the formal grammar Matteo Brucato Fjodor Kholodkov Soren Little Jakob Mayer Duc Nguyen http://arxiv.org/abs/2603.08612v2 Query-Guided Analysis and Mitigation of Data Verification Errors (Extended Version) 2026-08-26T20:14:26Z Data verification, the process of labeling data items as correct or incorrect, is a preprocessing step that may critically affect the quality of results in data-driven pipelines. Despite recent advances, verification can still produce erroneous labels that propagate to downstream query results in complex ways. We present a framework that complements existing verification tools by assessing the impact of potential labeling errors on query outputs and guiding additional verification steps to improve result reliability. To this end, we introduce Maximal Error Score (MES), a worst-case uncertainty metric that quantifies the reliability of query output tuples independently of the underlying data distribution. As an auxiliary indicator, we identify risky tuples - input tuples for which reducing label uncertainty may counterintuitively increase the output uncertainty. We then develop efficient algorithms for computing MES and detecting risky tuples, as well as a generic algorithm, named MESReduce, that builds on both indicators and interacts with external verifiers to select effective additional verification steps. We implement our techniques in a prototype system and evaluate them on real and synthetic datasets, demonstrating that MESReduce can substantially and effectively reduce the MES and improve the accuracy of verification results. 2026-03-09T16:57:45Z This paper is an extended version of our paper published in the IEEE International Conference on Data Engineering (ICDE) 2026 Extended version of "Query-Guided Analysis and Mitigation of Data Verification Errors", published in the 2026 IEEE 42nd International Conference on Data Engineering (ICDE), pp. 1994-2007 Ran Schreiber Yael Amsterdamer 10.1109/ICDE65706.2026.00151 http://arxiv.org/abs/2608.26335v1 Realistic Counterfactual Explanations via Denial Constraints 2026-08-26T19:16:10Z In the realm of Explainable AI, classification results are often explained via counterfactuals (CFs for short), which are (ideally small) perturbations to an instance that lead to a change of classification label. Such CFs may serve as explanations for the prediction, pinpointing the features that were important. Existing explainability solutions typically aim at minimizing the distance of CFs from the original instance so that they are specific to it, and/or maximizing the diversity of CFs to cover multiple facets of the reasons underlying the prediction. In this paper, we note that in pursuing these aims, state-of-the-art explainability solutions may (and often do) yield counterfactual explanations that do not correspond to realistic instances. This limits their applicability and usefulness in practice. To remedy this, we combine ideas from Explainable AI with ideas from data management. Specifically, we capture realism of CFs via logical constraints that hold with respect to a dataset of examples (e.g., training set); the class of such constraints that we focus on is that of denial constraints, extensively studied in the context of relational databases. Algorithmically, we then combine explainable AI solutions to yield CFs, with ideas from data cleaning that we adapt to this unique setting, to transform CFs into realistic ones. Extensive experiments across four datasets validate that our solutions achieve realism with relatively minor compromise in terms of distance and diversity. They further validate that the dedicated optimizations that we have developed to speed up the search for CFs are indeed highly effective. 2026-08-26T19:16:10Z Proc. 32nd ACM SIGKDD Conf. on Knowledge Discovery and Data Mining (KDD '26), Jeju Island, South Korea, Aug 2026, pp. 33-44 Avia Asael Nave Frost Amir Gilad Daniel Deutch 10.1145/3770855.3817712