5 papers
Reuse, Don't Recompute: Efficient Large Reasoning Model Inference via Memory Orchestration
Daivik Patel, Shrenik Patel
Large reasoning models (LRMs) achieve strong accuracy through test-time scaling, generating longer chains of thought or sampling multiple solutions, but at steep costs in tokens an…
ENGRAM: Effective, Lightweight Memory Orchestration for Conversational Agents
Daivik Patel, Shrenik Patel
Large language models (LLMs) deployed in user-facing applications require long-horizon consistency: the ability to remember prior interactions, respect user preferences, and ground…
DoubleTake: Contrastive Reasoning for Faithful Decision-Making in Medical Imaging
Daivik Patel, Shrenik Patel
Accurate decision making in medical imaging requires reasoning over subtle visual differences between confusable conditions, yet most existing approaches rely on nearest neighbor r…
CacheFlow: Compressive Streaming Memory for Efficient Long-Form Video Understanding
Shrenik Patel, Daivik Patel
Long-form video question answering (VQA) overwhelms current vision-language models (VLMs) because attention and key-value (KV) caches grow with runtime, forcing either expensive in…
Creating a Cooperative AI Policymaking Platform through Open Source Collaboration
Aiden Lewington, Alekhya Vittalam, Anshumaan Singh +48
Advances in artificial intelligence (AI) present significant risks and opportunities, requiring improved governance to mitigate societal harms and promote equitable benefits. Curre…