activity
20242026
collaborators

6 papers

cs.CV2026

EvoComp: Learning Visual Token Compression for Multimodal Large Language Models via Semantic-Guided Evolutionary Labeling

Jiafei Song, Fengwei Zhou, Jin Qu +7

Recent Multimodal Large Language Models (MLLMs) have demonstrated strong performance on vision-language understanding tasks, yet their inference efficiency is often hampered by the…

cs.CV2026

CoGR-MoE: Concept-Guided Expert Routing with Consistent Selection and Flexible Reasoning for Visual Question Answering

Xiyin Zeng, Yi Lu, Hao Wang

Visual Question Answering (VQA) requires models to identify the correct answer options based on both visual and textual evidence. Recent Mixture-of-Experts (MoE) methods improve op…

cs.CV2026

GLIMPSE : Real-Time Text Recognition and Contextual Understanding for VQA in Wearables

Akhil Ramachandran, Ankit Arun, Ashish Shenoy +8

Video Large Language Models (Video LLMs) have shown remarkable progress in understanding and reasoning about visual content, particularly in tasks involving text recognition and te…

cs.CL2025

MOOSComp: Improving Lightweight Long-Context Compressor via Mitigating Over-Smoothing and Incorporating Outlier Scores

Fengwei Zhou, Jiafei Song, Wenjin Jason Li +4

Recent advances in large language models have significantly improved their ability to process long-context input, but practical applications are challenged by increased inference t…

cs.CV2024

EgoQR: Efficient QR Code Reading in Egocentric Settings

Mohsen Moslehpour, Yichao Lu, Pierce Chuang +7

QR codes have become ubiquitous in daily life, enabling rapid information exchange. With the increasing adoption of smart wearable devices, there is a need for efficient, and frict…

cs.CV2024

Lumos : Empowering Multimodal LLMs with Scene Text Recognition

Ashish Shenoy, Yichao Lu, Srihari Jayakumar +11

We introduce Lumos, the first end-to-end multimodal question-answering system with text understanding capabilities. At the core of Lumos is a Scene Text Recognition (STR) component…