2 papers
cs.CV2026
Attend Before Attention: Efficient and Scalable Video Understanding via Autoregressive Gazing
Baifeng Shi, Stephanie Fu, Long Lian +10
Multi-modal large language models (MLLMs) have advanced general-purpose video understanding but struggle with long, high-resolution videos -- they process every pixel equally in th…
cs.CL2024
Efficient In-Domain Question Answering for Resource-Constrained Environments
Isaac Chung, Phat Vo, Arman C. Kizilkale +1
Retrieval Augmented Generation (RAG) is a common method for integrating external knowledge into pretrained Large Language Models (LLMs) to enhance accuracy and relevancy in questio…