collaborators

5 papers

cs.CV2025

3D Weakly Supervised Semantic Segmentation via Class-Aware and Geometry-Guided Pseudo-Label Refinement

Xiaoxu Xu, Xuexun Liu, Jinlong Li +5

3D weakly supervised semantic segmentation (3D WSSS) aims to achieve semantic segmentation by leveraging sparse or low-cost annotated data, significantly reducing reliance on dense…

cs.CV2025

Weakly-Supervised 3D Visual Grounding based on Visual Language Alignment

Xiaoxu Xu, Yitian Yuan, Qiudan Zhang +4

Learning to ground natural language queries to target objects or regions in 3D point clouds is quite essential for 3D scene understanding. Nevertheless, existing 3D visual groundin…

cs.CV2025

InstructionBench: An Instructional Video Understanding Benchmark

Haiwan Wei, Yitian Yuan, Xiaohan Lan +2

Despite progress in video large language models (Video-LLMs), research on instructional video understanding, crucial for enhancing access to instructional content, remains insuffic…

cs.CV2024

TimeMarker: A Versatile Video-LLM for Long and Short Video Understanding with Superior Temporal Localization Ability

Shimin Chen, Xiaohan Lan, Yitian Yuan +2

Rapid development of large language models (LLMs) has significantly advanced multimodal large language models (LMMs), particularly in vision-language tasks. However, existing video…

cs.CV2024

VidCompress: Memory-Enhanced Temporal Compression for Video Understanding in Large Language Models

Xiaohan Lan, Yitian Yuan, Zequn Jie +1

Video-based multimodal large language models (Video-LLMs) possess significant potential for video understanding tasks. However, most Video-LLMs treat videos as a sequential set of…