3 citations · 5 across the 9 of their papers we have counts for
10 papers
Fast Recovery for LLM Serving via Decoupled Device Memory Lifetime in Dynamo
Schwinn Saereesitthipitak, Mohammed Abdulwahhab, Hannah Zhang +6
Large language model (LLM) inference replicas run across tightly coupled GPUs and serve traffic continuously for weeks. Hardware and software failures are therefore inevitable, and…
IRB: Automated Generation of Robust Factuality Benchmarks
Lam Thanh Do, Bhagyashree Taleka, Hozaifa Ammar Bhutta +3
Static benchmarks for RAG systems often suffer from rapid saturation and require significant manual effort to maintain robustness. To address this, we present IRB, a framework for…
VibeTensor: System Software for Deep Learning, Fully Generated by AI Agents
Bing Xu, Terry Chen, Fengzhe Zhou +12
VIBETENSOR is an open-source research system software stack for deep learning, generated by LLM-powered coding agents under high-level human guidance. In this paper, "fully generat…
Five-Minute Rule 40 Years Later: A First-Principles Revisit for Modern Memory Hierarchy
Tong Zhang, Vikram Sharma Mailthody, Fei Sun +7
In 1987, Jim Gray and Gianfranco Putzolu introduced the five-minute rule, a simple, storage-memory-economics-based heuristic for deciding when data should live in DRAM rather than…
AUGUSTUS: An LLM-Driven Multimodal Agent System with Contextualized User Memory
Jitesh Jain, Shubham Maheshwari, Ning Yu +2
Riding on the success of LLMs with retrieval-augmented generation (RAG), there has been a growing interest in augmenting agent systems with external memory databases. However, the…
Generalized Neighborhood Attention: Multi-dimensional Sparse Attention at the Speed of Light
Ali Hassani, Fengzhe Zhou, Aditya Kane +13
Many sparse attention mechanisms such as Neighborhood Attention have typically failed to consistently deliver speedup over the self attention baseline. This is largely due to the l…