4 papers
Towards an Effective Action-Region Tracking Framework for Fine-grained Video Action Recognition
Baoli Sun, Yihan Wang, Xinzhu Ma +3
Fine-grained action recognition (FGAR) aims to identify subtle and distinctive differences among fine-grained action categories. However, current recognition methods often capture…
Referring Video Object Segmentation with Cross-Modality Proxy Queries
Baoli Sun, Xinzhu Ma, Ning Wang +2
Referring video object segmentation (RVOS) is an emerging cross-modality task that aims to generate pixel-level maps of the target objects referred by given textual expressions. Th…
Propagating Sparse Depth via Depth Foundation Model for Out-of-Distribution Depth Completion
Shenglun Chen, Xinzhu Ma, Hong Zhang +2
Depth completion is a pivotal challenge in computer vision, aiming at reconstructing the dense depth map from a sparse one, typically with a paired RGB image. Existing learning bas…
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens
Zhuqiang Lu, Zhenfei Yin, Mengwei He +4
Recently, Vision Large Language Models (VLLMs) integrated with vision encoders have shown promising performance in vision understanding. The key of VLLMs is to encode visual conten…