4 papers
Unified Energy for Invariant and Independent Decoding in Diffusion Language Models
Yuchen Yan, Minkai Xu, Zaiquan Yang +1
Diffusion Language Models (DLMs) enable parallel text generation by iteratively denoising a full sequence, offering attractive flexibility compared to auto-regressive (AR) decoding…
Visual Enhanced Depth Scaling for Multimodal Latent Reasoning
Yudong Han, Yong Wang, Zaiquan Yang +3
Multimodal latent reasoning has emerged as a promising paradigm that replaces explicit Chain-of-Thought (CoT) decoding with implicit feature propagation, simultaneously enhancing r…
Unleashing the Potential of Multimodal LLMs for Zero-Shot Spatio-Temporal Video Grounding
Zaiquan Yang, Yuhao Liu, Gerhard Hancke +1
Spatio-temporal video grounding (STVG) aims at localizing the spatio-temporal tube of a video, as specified by the input text query. In this paper, we utilize multimodal large lang…
Boosting Weakly-Supervised Referring Image Segmentation via Progressive Comprehension
Zaiquan Yang, Yuhao Liu, Jiaying Lin +2
This paper explores the weakly-supervised referring image segmentation (WRIS) problem, and focuses on a challenging setup where target localization is learned directly from image-t…