4 citations · 5 across the 2 of their papers we have counts for
2 papers
cs.CL2024★ 1 cited
Cost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttention
Bin Gao, Zhuomin He, Puru Sharma +6
Interacting with humans through multi-turn conversations is a fundamental feature of large language models (LLMs). However, existing LLM serving engines executing multi-turn conver…
cs.ET2022★ 4 cited
Efficiently Enabling Block Semantics and Data Updates in DNA Storage
Puru Sharma, Cheng-Kai Lim, Dehui Lin +2
We propose a novel and flexible DNA-storage architecture, which divides the storage space into fixed-size units (blocks) that can be independently and efficiently accessed at rando…