4 papers
GuideMe: Multi-Domain Task Guidance and Intervention in Streaming Video
Fang Liu, Jinpeng Chen, Ke Xu +7
While multimodal Large Language Models (MLLMs) excel at offline video understanding, an interesting question of how far they are from serving as a real-time procedural coach remain…
EgoCS-400K: An Egocentric Gameplay Dataset for World Models
Rongjin Guo, Dong Liang, Yuhao Liu +4
The shift from video generation to interactive world modeling places new demands on data: beyond captioned videos, world models require temporally aligned video-action-language tra…
Revisiting the Integration of Convolution and Attention for Vision Backbone
Lei Zhu, Xinjiang Wang, Wayne Zhang +1
Convolutions (Convs) and multi-head self-attentions (MHSAs) are typically considered alternatives to each other for building vision backbones. Although some works try to integrate…
Boosting Weakly-Supervised Referring Image Segmentation via Progressive Comprehension
Zaiquan Yang, Yuhao Liu, Jiaying Lin +2
This paper explores the weakly-supervised referring image segmentation (WRIS) problem, and focuses on a challenging setup where target localization is learned directly from image-t…