4 papers
What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents
Ashwin Gerard Colaco, Nada Lahjouji
Large language models, and the agents built on them, spend an ever-growing share of their compute and memory on remembering: caching attention keys and values, carrying long prompt…
Agents That Know Too Much: A Data-Centric Survey of Privacy in LLM Agents
Nada Lahjouji, Ashwin Gerard Colaco
Large language model agents increasingly query databases, search document collections, call external APIs, remember past interactions, and act on a user's behalf. As they move from…
GenIE - Simulator-Driven Iterative Data Exploration for Scientific Discovery
Ashwin Gerard Colaco, Martin Boissier, Sriram Rao +3
Physics-based simulators play a critical role in scientific discovery and risk assessment, enabling what-if analyses for events like wildfires and hurricanes. Today, databases trea…
NOMAD -- Navigating Optimal Model Application to Datastreams
Ashwin Gerard Colaco, Sharad Mehrotra, Michael J De Lucia +5
NOMAD (Navigating Optimal Model Application for Datastreams) is an intelligent framework for data enrichment during ingestion that optimizes realtime multiclass classification by d…