2 papers
cs.CV2026
Learning Trajectory-Aware Multimodal Large Language Models for Video Reasoning Segmentation
Jingnan Luo, Mingqi Gao, Jun Liu +2
The prosperity of Multimodal Large Language Models (MLLMs) has stimulated the demand for video reasoning segmentation, which aims to segment video objects based on human instructio…
cs.CV2026
Show Me When and Where: Towards Referring Video Object Segmentation in the Wild
Mingqi Gao, Jinyu Yang, Jingnan Luo +4
Referring video object segmentation (RVOS) has recently generated great popularity in computer vision due to its widespread applications. Existing RVOS setting contains elaborately…