https://arxiv.org/api/iY0toizNFC1O0+o6zv7ZirVmC7U2026-10-03T22:40:43Z1246822515http://arxiv.org/abs/2609.01818v1Zeta-Lite: A Concurrent, Branchable In-Browser SQL Database for Agentic Memory2026-09-01T19:47:01ZThe browser has become a first-class database host: applications increasingly want to store, query, and reason over structured data entirely on the client - for privacy, offline operation, local-first collaboration, and, most recently, as durable memory for in-browser AI agents. One way to get SQL in the browser, compiling PostgreSQL to WebAssembly (PGlite), inherits PostgreSQL's process model: a single backend connection that executes one statement at a time and blocks. That model cannot express concurrent transactions, and it leaves richer capabilities - graph queries, database branching - to whatever the compiled server happens to include. We present zeta-lite, the browser form factor of the Zeta database engine: a WebAssembly build that compiles the same Zeta server down to a 2.87 MB gzipped artifact. Zeta-lite keeps the engine's log-centric asynchronous MVCC core, which yields two capabilities no other in-browser SQL engine provides. First, overlapping snapshot-isolated transactions on a single thread: multiple transactions hold distinct read/commit timestamps and interleave, with snapshot-isolation conflict detection between them. Second, copy-on-write database branching - whole-database fork, merge, and rebase - is unique in a browser SQL database and rare even in servers. On top of these, zeta-lite exposes a feature-complete PostgreSQL surface (joins, CTEs, window functions, JSONB with GIN indexes, full-text search, HNSW vector search, SQL/PGQ graph queries, multi-database) and snapshot-to-OPFS durability. Across Chrome, Firefox, and a native reference runtime, zeta-lite sustains 268k-315k point reads/s and holds a mixed read/write workload flat over millions of operations. This small, fully-featured, concurrent SQL database is an especially good fit for agentic memory - where cheap branchable state lets an agent explore, inspect, and commit or discard speculative work.2026-09-01T19:47:01ZGene Zhanghttp://arxiv.org/abs/2609.01525v1Relational-Core Graph Analytics Querying graphs at SQL scale, and why the node/edge model is a performance tax, not a truer picture of connected data2026-09-01T16:55:11ZA durable assumption holds that graph analytics requires a purpose-built graph engine, and that relational systems are ill-suited to connected data. We argue the opposite for the workloads enterprises actually run. A columnar relational engine fronted by a graph query language matches or exceeds native graph engines on analytical graph queries, and - decisively - scales past the point where in-memory graph engines fail. We further argue that the node/edge property graph is not a more faithful model of connected data but a re-encoding of relationships that already exist explicitly in relational tables; reconstructing them at query time is pure overhead. We present ClickGraph and its Databricks-dialect sibling DeltaGraph, systems that translate Cypher directly onto the native relational schema - the tables, columns, and foreign keys as they already exist - and execute in place on ClickHouse, Databricks, or in-process on lakehouse files, with no import and no separate cluster. Because the output is ordinary SQL, an underperforming query is an open optimization surface: it can be rewritten, and the engine itself extended. We support the argument with a peer system's own published benchmark, in which a columnar engine outruns Neo4j by two-to-four orders of magnitude, and with reproducible measurements across the LDBC Social Network Benchmark suite.2026-09-01T16:55:11ZGene Zhanghttp://arxiv.org/abs/2609.01292v1Relational Task Generation Language: A Declarative Specification Framework for Relational Deep Learning2026-09-01T14:27:06ZRelational Deep Learning (RDL) has become a powerful paradigm for learning from multi-tabular data. However, manually defining RDL prediction tasks is a laborious process that frequently results in data leakage. To address this issue, we introduce Relational Task Generation Language (RTGL) - an open-source declarative language that streamlines RDL task formulation by abstracting away low-level SQL details. We showcase RTGL by reconstructing existing RDL benchmark tasks and uncovering their inconsistencies stemming from manually crafted SQL definitions of RDL prediction targets, thereby underscoring the value of a dedicated declarative language. In addition, we demonstrate the practical utility of RTGL by designing various new tasks with diverse forms and target types. Our experiments confirm the robustness and usability of RTGL, as well as its seamless integration with the existing RDL frameworks, making it widely accessible to the community.2026-09-01T14:27:06ZAccepted to MLG 2026Oleksii KolesnichenkoJakub PeleškaGustav Šírhttp://arxiv.org/abs/2609.00783v1Efficient discovery of unique column combinations on disk-resident data with limited memory2026-09-01T06:27:37ZThe discovery of unique column combinations (UCCs) is a core task in data profiling, describing the key constraints of a table. The existing algorithms cannot deal with large-scale disk-resident data well due to high memory consumption and computational cost. In this paper, a novel DUD algorithm is developed to efficiently discover UCCs on disk-resident data with limited memory, which is inspired by the relationship between UCC discovery and transversal hypergraph. Rather than complete difference set generation of quadratic complexity, DUD only generates partial difference sets for hypergraph construction, followed by minimal hitting set enumeration to generate candidates and a validation process. DUD devises a strategy to generate full useful difference sets by pairwise comparisons of tuples having the same values with respect to some selected attributes. A novel theorem is developed and proved in this paper to report the candidates including the selected attributes as true UCCs directly without validation, which reduces the number of candidates to be validated significantly. A hash-based batch validation strategy is devised to validate a set of candidates on the relation instance, which only needs to maintain a small number of tuples in memory at a time. The extensive experimental results, conducted on synthetic and real-life data sets, show that DUD can discover UCCs on disk-resident data with high efficiency and low memory consumption.2026-09-01T06:27:37ZXiaolong WanXixian Hanhttp://arxiv.org/abs/2609.00749v1ContextPipe: Database-Inspired Context Assembly for Long-Horizon Agents2026-09-01T05:27:38ZLong-horizon large language model (LLM) agents require context assembly: the runtime must decide what to include in each prompt, in what order, and when to compact history under a hard context-window budget and a byte-sensitive prompt cache. In production agentic systems, this logic is scattered across prompt builders, ad hoc compaction routines, cache-break workarounds, and per-provider shims. We argue that context assembly is structurally isomorphic to query execution in a relational database: both execute under a hard budget, exploit a tiered cache, and leverage statistics. We adopt this discipline in ContextPipe: a five-phase pipeline (Plan Bind Optimize Execute Feedback) backed by a structured data-source catalog, a deterministic cache-aware optimizer, and an EXPLAIN ANALYZE trace. We show that context in ContextPipe is auditable, replayable, and failure-isolated. A preliminary evaluation using the SWE-bench Pro Qutebrowser subset shows that, compared with the append-only context construction policy, ContextPipe reduces total token volume by 31%, LLM calls by 23%, and response time by 9%, at the cost of a lower KV cache-hit ratio.2026-09-01T05:27:38ZPeng XuZuyu ZhangYuze SunFeng TianLong WangChen Zhanghttp://arxiv.org/abs/2608.30607v2ByteX: A Unified AI Search Engine at ByteDance2026-09-01T04:48:30ZSince 2016, ByteX has been the foundation of ByteDance's search infrastructure, scaling to more than 7,000 clusters and 300 PB of indexed data. Driven by the demands of AI workloads, ByteX has evolved from a text search engine into a unified AI search system supporting vector retrieval, lexical matching, and predicate filtering. Its largest deployment indexes nearly one trillion high-dimensional vectors. This scale exposes two central bottlenecks in AI-era retrieval: memory-intensive graph-index construction under sustained ingestion, and the prohibitive cost of keeping vector indexes entirely in memory. ByteX addresses these bottlenecks with two techniques. First, it introduces a quantization-aware vector kernel based on SymRaBitQ, a new symmetric quantization scheme with tight theoretical guarantees that allows index construction to run directly in the quantized space accurately and efficiently without retaining a copy of full-precision vectors. Second, it provides a hybrid storage engine that supports memory-resident, hybrid, and SSD-resident deployments, with fine-grained record-level caching to trade memory for latency under operational control. On large-scale benchmarks, ByteX improves throughput by up to 3x, reduces indexing memory by 80%, and lowers operating cost by 86% compared with prior systems, while supporting trillion-vector scale, write-heavy or latency-sensitive workloads in production.2026-08-31T11:19:16ZYao TianYuncheng LuLiyao XiongYuming XuHao ZhangWeichen ZhaoXi ZhaoBo KuangDongyu WangJiehui LiYakun LiLei Zhanghttp://arxiv.org/abs/2609.00548v1Time-Decayed Vector Search in the Rhythm of TANGO: Jointly Modeling Semantic Similarity and Temporal Freshness2026-09-01T01:33:45ZVector search typically measures relevance through semantic similarity under a fixed scoring function. However, in a growing range of applications, relevance may evolve over time, making temporal freshness an additional signal beyond semantic similarity. In this paper, we formalize time-decayed vector search (TDVS), which incorporates continuous temporal decay into the search objective so that relevance is jointly determined by semantic similarity and temporal freshness. We design Score-Preserving Temporal Reduction (STR) that enables existing Maximum Inner Product Search indexes to directly support TDVS. We further present Chronos, a TDVS-native framework that derives an exact metric formulation and introduces Query-Orthogonal TimeLift to control data--data geometry while preserving all query--data scores and rankings. Building on Chronos, we propose TANGO, a hierarchical graph index that adopts layer-specific TimeLift geometries to preserve temporal locality at the base layer while strengthening long-range semantic connectivity in upper layers. TANGO traverses the hierarchy using the exact TDVS score, caches temporal factors to reduce computation, and supports efficient online insertion. Extensive experiments show that TANGO achieves up to 3.5$\times$ higher query throughput and 4.05$\times$ faster index construction than state-of-the-art graph-based competitors. TANGO also maintains its advantage over all competitors across diverse temporal settings and enables efficient online insertion, demonstrating its robustness and practicality.2026-09-01T01:33:45ZJiuqi WeiQiyao LuoQuanqing XuChuanhui YangThemis Palpanashttp://arxiv.org/abs/2609.00181v1Intelligent Edge Computing2026-08-31T18:04:32ZThe number of edge devices in large-scale edge systems is rapidly increasing. Edge devices have limited processing power, memory, and network bandwidth, making resource utilization and data management during edge query processing challenging. Joins are among the costliest database operations in terms of time and resources. The State-of-the-Art edge query processing, Column Imprint-Hash Join CI-HJ, addresses this challenge using equi-height binning to accelerate hash joins. However, it lacks efficiency in real-time processing and scans unnecessary cachelines. This paper presents Workload Aware Column Imprint-Hash Join WACI-HJ, which uses a workload-aware approach to accelerate hash joins. Predicting the upcoming query workload in advance further improves its suitability for real-time edge query processing. WACI-HJ comprises two phases: WACI-HJ Generation Phase, including Pre-processing, Prediction, and Blocking and Hashing modules to compute bins based on the predicted workload before query arrival, and Query Processing and Resource Utilization, which handles query processing and CPU, RAM, and I/O utilization. Evaluations on a benchmark dataset and a real-world Smart Transportation dataset show a 54% reduction in cachelines read and 10% improved query execution time. The proposed technique is effective for both scaled and skewed data. Although PCR is an indirect measure of energy consumption, the work also directly measures energy consumption through energy-efficiency experiments. WACI-HJ shows 1%, 38%, and 49% gain in CPU, RAM, and I/O, respectively. Optimizing cache usage and query execution speeds up real-time traffic analysis, congestion management, and routing in Smart Transportation. Additionally, this technology can be applied to other domains to accelerate edge query processing.2026-08-31T18:04:32ZKalgi GandhiMinal Bhisehttp://arxiv.org/abs/2608.31150v1Local Private Information Retrieval for Graph-Based Replicated Systems2026-08-31T17:50:47ZWe rethink the definition of privacy in multi-server, graph-replicated private information retrieval (PIR) systems, by introducing a novel setting where the user's privacy is governed by the servers' storage structure. In classical graph-replicated PIR, the user retrieves a single message stored at the servers, while hiding the message index from each server. In our proposed privacy setting, the user is concerned with hiding the message index from a particular server, only if that server stores the message being retrieved, and privacy is not imposed otherwise. We coin this relaxed privacy requirement as local user privacy and the resulting PIR problem as local PIR on the graph. Our focus is on two-replicated PIR systems, where every message is replicated twice and stored on two distinct servers. Specifically, we study local PIR systems where the storage is represented by simple graphs, i.e., every pair of vertices is associated with at most one edge, and by their multigraph extension, i.e., $r$ parallel edges replace every edge. For these settings, we establish bounds on the local PIR capacity, defined as the maximum number of message symbols retrieved, per downloaded symbol. The local privacy requirement yields significant capacity gain over the classical PIR capacity under the same storage structure. For instance, in settings where the graph is a disjoint union of multiple identical sub-graphs, the gain in the local PIR capacity over classical PIR capacity is multiplicative in the number of sub-graphs. Further, for connected graphs, we derive capacity lower bounds for edge-transitive and bipartite graphs, which are greater than the best-known PIR capacity bounds. From these and by establishing matching upper bounds, we exactly characterize the capacity for star graphs, cyclic graphs, and path graphs with odd number of vertices. We introduce two local PIR schemes for general graphs.2026-08-31T17:50:47ZShreya MeelMohamed NomeirSennur Ulukushttp://arxiv.org/abs/2608.31027v1GRBench: A Comprehensive Benchmark Evaluation for Graph-relational Data Management2026-08-31T16:09:04ZModern data-intensive applications increasingly require database systems to manage structured records and graph data. This demand gives rise to graph-relational data management, spanning storage, query processing, and optimization across relational and graph data. In response, relational database extensions, multi-model databases, and dedicated graph-relational systems have emerged with diverse architectures. However, evaluation methodologies have not kept pace. Existing relational and graph benchmarks assess the two models largely in isolation, while multi-model benchmarks provide limited coverage of graph-relational workloads. Available graph-relational workloads mainly support functional validation and end-to-end latency measurement, revealing little about how storage, operator, and optimization designs affect performance. To evaluate system capabilities in graph-relational data management, we present GRBench. First, GRBench constructs a linked graph-relational schema from the real-world SciSciNet-v2 dataset and derives scalable instances through consistency-preserving subset extraction. Second, it organizes purpose-built query series for controlled evaluation of query processing and system components. Third, GRBench provides semantically equivalent native query formulations and evaluates representative system architectures through a unified, multidimensional methodology. Based on this evaluation, we analyze design trade-offs and identify open challenges to guide future system design and optimization.2026-08-31T16:09:04ZZepeng LiuXinxin HuangXuanming LiuSheng WangZhiyong Penghttp://arxiv.org/abs/2604.12431v3VeriX-Anon: A Multi-Layered Framework for Mathematically Verifiable Outsourced Target-Driven Data Anonymization2026-08-31T15:59:57ZOrganisations increasingly outsource privacy-sensitive data transformations to cloud providers, yet no practical mechanism lets the data owner verify that the contracted algorithm was faithfully executed. VeriX-Anon is a multi-layered verification framework for outsourced Target-Driven k-anonymization combining three orthogonal mechanisms: deterministic verification via Merkle-style hashing of an Authenticated Decision Tree, probabilistic verification via Boundary Sentinels and exact-duplicate Twins with cryptographic identifiers, and utility-based verification via Explainable AI fingerprinting that compares SHAP value distributions before and after anonymization using the Wasserstein distance. Across seven cross-domain datasets and four cloud profiles (28 scenarios), against Lazy (drops records), Dumb (fake hash), and Approximate (valid hash) adversaries, VeriX-Anon detects 25 of 28 deviations under a fixed threshold and 27 of 28 once the threshold is calibrated per dataset, with no false alarms. No single layer achieved this alone. The XAI layer was the only mechanism that caught the Approximate adversary, succeeding on six of seven datasets and missing only a high-dimensional case where honest generalization shifts SHAP as much as the attack. Target-Driven anonymization preserved significantly more utility than blind splitting, with mean F1 gaps of +0.058 to +0.362 and Wilcoxon p <= 0.001 on six of seven datasets. Client-side verification completes under one second at one million rows. The threat model covers three empirically evaluated profiles and one theoretical Informed Attacker unable to defeat the cryptographic salt. Sentinel evasion probability ranges from near-zero to 0.82 for the most imbalanced data, which the twin layer offsets in every scenario.2026-04-14T08:22:18Zv2: revised after peer review. Evaluation expanded from 3 to 7 datasets, per-dataset Wasserstein-threshold calibration added, effect-size CIs and cross-dataset statistics reported, and analytical zk-SNARK/MPC/TEE baselines added. Minor errors correctedArray 31 (2026) 101169Miit DagaSwarna Priya Ramu10.1016/j.array.2026.101169http://arxiv.org/abs/2606.19319v2Data Intelligence Agents: Interpreting, Modeling, and Querying Enterprise Data via Autonomous Coding Agents2026-08-31T09:22:23ZProduction data integration is bottlenecked by repeated, lossy handoffs between data owners, engineers, and analysts who must collaboratively discover, structure, and query enterprise data. We present Data Intelligence Agents (DIA), a system of three agents (Data Interpreter, Schema Creator, and Query Generator) that compresses this workflow by treating autonomous coding agents (ACAs) as a first-class abstraction: rather than emitting text, the agents generate, execute, validate, and repair concrete artifacts, draw on a shared memory for experience reuse, and surface each for review by domain experts. DIA is deployed in production for enterprise customers. We study the Query Generator in depth and evaluate it in fully autonomous mode across seven SQL benchmarks spanning four task categories and four dialects. It matches or surpasses the best published results on all seven, demonstrating that an architecture grounded in execution, built on ACAs and a shared memory, generalizes across the data intelligence workload with adaptation confined to natural-language instructions.2026-06-17T17:45:32ZAnoushka VyasAarushi DhanukaSina Khoshfetrat PakazadHenrik Ohlssonhttp://arxiv.org/abs/2608.30465v1Strengthening LargeRDFBench for Interoperable Federated SPARQL Evaluation2026-08-31T08:51:16ZLargeRDFBench is one of the most comprehensive benchmarks for evaluating federated SPARQL query engines, combining real, interlinked datasets with a rich query suite that has made it a reference point for the community. Evaluations of federated engines are published by comparing engine results against the benchmark's expected results, so those expected results must themselves be reproducible. Moreover, several of its data dumps violate the RDF specifications, so only engines that parse RDF leniently can host them, and its expected results are distributed in an ad hoc format. We identify, categorize and repair these data-quality issues with a reproducible cleaning pipeline, producing standards-conformant serializations of every affected dataset. Furthermore, we re-encode the benchmark's expected results in the W3C SPARQL 1.1 Query Results JSON Format and correct their discrepancies. Every dataset now parses under strict RDF parsers, and the expected results are machine-verifiable through a standard format, extending the benchmark's reach to the full range of conformant engines while staying faithful to the original data. Reproducing the expected results end-to-end with an independent implementation uncovers corruption in the published reference, and discrepancies between our results and the original ones, some not trivial to resolve, others open questions. We further perform a preliminary comparison, not previously explored, of ASK- and COUNT-based source selection in the FedX algorithm. This work strengthens an already valuable community resource by aligning its artifacts with the RDF standards, broadening the set of engines that can be fairly and reproducibly compared. We also raise the question of how the results of federated queries under automatic source selection can be made reproducible.2026-08-31T08:51:16ZBryan-Elliott TamMuhammad SaleemRuben Taelmanhttp://arxiv.org/abs/2608.30385v1Detecting DBMS Bugs by Constructing Equivalent Representations of Intermediate Query Results2026-08-31T07:41:09ZDatabase Management Systems (DBMSs) support multiple SQL mechanisms for representing intermediate query results, including VIEWs, Common Table Expressions (CTEs), and Temporary Tables (TEMPTs). When these mechanisms are used to represent the same intermediate query result, the corresponding queries are expected to produce consistent results. However, we observe that such queries can return inconsistent results, indicating potential DBMS logic bugs. Existing approaches for detecting DBMS logic bugs have never explored result consistency across such equivalent representations. In this paper, we propose ERIQ, a novel testing approach for detecting DBMS logic bugs from the perspective of checking result consistency across Equivalent Representations of Intermediate Query Results. ERIQ constructs SQL variants using a VIEW, a CTE, or a TEMPT to represent the same intermediate query result, executes these variants, and compares their returned results. We evaluated ERIQ on four widely used open-source DBMSs: MySQL, MariaDB, Percona, and OceanBase. In total, ERIQ detected 64 bugs, 63 of which were confirmed by developers, and two have been fixed. Among the confirmed bugs, 54 were unique and previously unknown logic bugs, and one was a documentation issue.2026-08-31T07:41:09Z12 pages, 3 figuresXiaoxu NiuGong ChenJinfu ChenXiaoyuan Xiehttp://arxiv.org/abs/2608.30227v1ELASTIC: Trajectory-Based Synchronization of Event and Tracking Data in Soccer2026-08-31T04:23:30ZCombining event and tracking data is fundamental to modern soccer analytics, yet the two sources are rarely well aligned: event timestamps recorded by human annotators often miss the true moment of the action, distorting the spatiotemporal context that downstream models rely on. Existing synchronization methods depend on noisy human-annotated event locations and fail to detect ball receptions, obscuring when each player gains ball possession. To address these limitations, we propose ELASTIC (Event-Location-AgnoSTIC synchronizer), a framework that infers the start and end timestamps of events solely from player and ball trajectories, without relying on annotated event locations. To recover ball receptions, ELASTIC enriches the event sequence by inserting virtual termination events between consecutive events, so that the end of each event is detected jointly with its start. It then extracts a sparse set of candidate frames where ball touches are physically plausible, and aligns the termination-inserted event sequence with the candidate-frame sequence using an extended Needleman-Wunsch algorithm. For reproducible evaluation, we construct a publicly available benchmark by annotating ground-truth timestamps on the Sportec Open DFL Dataset, on which ELASTIC substantially outperforms existing methods. Through downstream task evaluation, we further show that improved synchronization translates into measurable gains in soccer analytics. The source code and benchmark are available at https://github.com/hyunsungkim-ds/elastic.git.2026-08-31T04:23:30ZHyunsung KimHoyoung ChoiKunhee LeeSangwoo SeoTom BoomstraJinsung YoonChanyoung Park