collaborators

5 papers

cs.CV2026

Not All Modalities Are Equal: Instruction-Aware Gating for Multimodal Videos

Bonan Ding, Umair Nawaz, Ufaq Khan +5

Pre-trained video large language models excel at visual reasoning. However, they struggle when videos arrive with auxiliary streams, such as audio, depth map, or dense temporal evi…

cs.CL2026

STDec: Spatio-Temporal Stability Guided Decoding for dLLMs

Yuzhe Chen, Jiale Cao, Xuyang Liu +3

Diffusion Large Language Models (dLLMs) have achieved rapid progress, viewed as a promising alternative to the autoregressive paradigm. However, most dLLM decoders still adopt a gl…

cs.CV2024

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation

Lin Sun, Jiale Cao, Jin Xie +2

Contrastive Language-Image Pre-training (CLIP) exhibits strong zero-shot classification ability on various image-level tasks, leading to the research to adapt CLIP for pixel-level…

cs.CV2024

iSeg: An Iterative Refinement-based Framework for Training-free Segmentation

Lin Sun, Jiale Cao, Jin Xie +2

Stable diffusion has demonstrated strong image synthesis ability to given text descriptions, suggesting it to contain strong semantic clue for grouping objects. The researchers hav…

cs.CV2024

CLIP-VIS: Adapting CLIP for Open-Vocabulary Video Instance Segmentation

Wenqi Zhu, Jiale Cao, Jin Xie +2

Open-vocabulary video instance segmentation strives to segment and track instances belonging to an open set of categories in a videos. The vision-language model Contrastive Languag…