collaborators
Showing cs.DBShow all

6 papers · 1 filter

cs.DB2026

Evergreen: Efficient Claim Verification for Semantic Aggregates

Alexander W. Lee, Benjamin Han, Shayak Sen +3

With recent semantic query processing engines, semantic aggregation has become a primitive operator, enabling the reduction of a relation into a natural language aggregate using an…

cs.DB2026

AvalancheBench: Evaluating Enterprise Data Agents Through Latent World Recovery

Darek Kleczek, Fuheng Zhao, Alexander W. Lee +4

We introduce AvalancheBench, a benchmark for evaluating enterprise data agents through \emph{latent world recovery}. AvalancheBench improves on existing benchmarks in three ways. F…

cs.DB2026

VectraFlow: Long-Horizon Semantic Processing over Data and Event Streams with LLMs

Shu Chen, Junhan Liu, Deepti Raghavan +1

Monitoring continuous data for meaningful signals increasingly demands long-horizon, stateful reasoning over unstructured streams. However, today's LLM frameworks remain stateless…

cs.DB2026

Making Prompts First-Class Citizens for Adaptive LLM Pipelines

Ugur Cetintemel, Shu Chen, Alexander W. Lee +3

Modern LLM pipelines increasingly resemble complex data-centric applications: they retrieve data, correct errors, call external tools, and coordinate interactions between agents. Y…

cs.DB2025

Continuous Prompts: LLM-Augmented Pipeline Processing over Unstructured Streams

Shu Chen, Deepti Raghavan, Uğur Çetintemel

Monitoring unstructured streams increasingly requires persistent, semantics-aware computation, yet today's LLM frameworks remain stateless and one-shot, limiting their usefulness f…

cs.DB2025

Semantic Integrity Constraints: Declarative Guardrails for AI-Augmented Data Processing Systems

Alexander W. Lee, Justin Chan, Michael Fu +4

AI-augmented data processing systems (DPSs) integrate large language models (LLMs) into query pipelines, allowing powerful semantic operations on structured and unstructured data.…