5 papers
Mosaic: Runtime-Efficient Multi-Agent Embodied Planning
Kunjal Panchal, Saayan Mitra, Sunav Choudhary +3
LLM-based multi-agent embodied planning remains impractical due to prohibitively high execution latency. We identify failed actions as the dominant bottleneck, stemming from two co…
Memory Savings at What Cost? A Study of Alternatives to Backpropagation
Kunjal Panchal, Sunav Choudhary, Yuriy Brun +1
Forward-mode automatic differentiation (FmAD) and zero-order (ZO) optimization are increasingly proposed as memory-efficient, backpropagation-free alternatives for large language m…
CATTO: Balancing Preferences and Confidence in Language Models
Nisarg Parikh, Ananya Sai, Pannaga Shivaswamy +2
Large language models (LLMs) often make accurate next token predictions but their confidence in these predictions can be poorly calibrated: high-confidence predictions are frequent…
Atom: Efficient On-Device Video-Language Pipelines Through Modular Reuse
Kunjal Panchal, Saayan Mitra, Somdeb Sarkhel +4
Recent advances in video-language models have enabled powerful applications like video retrieval, captioning, and assembly. However, executing such multi-stage pipelines efficientl…
Thinking Forward: Memory-Efficient Federated Finetuning of Language Models
Kunjal Panchal, Nisarg Parikh, Sunav Choudhary +3
Finetuning large language models (LLMs) in federated learning (FL) settings has become increasingly important as it allows resource-constrained devices to finetune a model using pr…