4 citations · 9 across the 4 of their papers we have counts for
4 papers · 1 filter
VideoZoomer: Reinforcement-Learned Temporal Focusing for Long Video Reasoning
Yang Ding, Yizhen Zhang, Xin Lai +2
Multimodal Large Language Models (MLLMs) have achieved remarkable progress in vision-language tasks yet remain limited in long video understanding due to the limited context window…
VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning
Senqiao Yang, Junyi Li, Xin Lai +3
Recent advancements in vision-language models (VLMs) have improved performance by increasing the number of visual tokens, which are often significantly longer than text tokens. How…
LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model
Senqiao Yang, Tianyuan Qu, Xin Lai +4
While LISA effectively bridges the gap between segmentation and large language models to enable reasoning segmentation, it poses certain limitations: unable to distinguish differen…
Mask-Attention-Free Transformer for 3D Instance Segmentation
Xin Lai, Yuhui Yuan, Ruihang Chu +3
Recently, transformer-based methods have dominated 3D instance segmentation, where mask attention is commonly involved. Specifically, object queries are guided by the initial insta…