6 papers
EvoComp: Learning Visual Token Compression for Multimodal Large Language Models via Semantic-Guided Evolutionary Labeling
Jiafei Song, Fengwei Zhou, Jin Qu +7
Recent Multimodal Large Language Models (MLLMs) have demonstrated strong performance on vision-language understanding tasks, yet their inference efficiency is often hampered by the…
CoGR-MoE: Concept-Guided Expert Routing with Consistent Selection and Flexible Reasoning for Visual Question Answering
Xiyin Zeng, Yi Lu, Hao Wang
Visual Question Answering (VQA) requires models to identify the correct answer options based on both visual and textual evidence. Recent Mixture-of-Experts (MoE) methods improve op…
GLIMPSE : Real-Time Text Recognition and Contextual Understanding for VQA in Wearables
Akhil Ramachandran, Ankit Arun, Ashish Shenoy +8
Video Large Language Models (Video LLMs) have shown remarkable progress in understanding and reasoning about visual content, particularly in tasks involving text recognition and te…
MOOSComp: Improving Lightweight Long-Context Compressor via Mitigating Over-Smoothing and Incorporating Outlier Scores
Fengwei Zhou, Jiafei Song, Wenjin Jason Li +4
Recent advances in large language models have significantly improved their ability to process long-context input, but practical applications are challenged by increased inference t…
EgoQR: Efficient QR Code Reading in Egocentric Settings
Mohsen Moslehpour, Yichao Lu, Pierce Chuang +7
QR codes have become ubiquitous in daily life, enabling rapid information exchange. With the increasing adoption of smart wearable devices, there is a need for efficient, and frict…
Lumos : Empowering Multimodal LLMs with Scene Text Recognition
Ashish Shenoy, Yichao Lu, Srihari Jayakumar +11
We introduce Lumos, the first end-to-end multimodal question-answering system with text understanding capabilities. At the core of Lumos is a Scene Text Recognition (STR) component…