4 papers
3D Weakly Supervised Semantic Segmentation via Class-Aware and Geometry-Guided Pseudo-Label Refinement
Xiaoxu Xu, Xuexun Liu, Jinlong Li +5
3D weakly supervised semantic segmentation (3D WSSS) aims to achieve semantic segmentation by leveraging sparse or low-cost annotated data, significantly reducing reliance on dense…
Weakly-Supervised 3D Visual Grounding based on Visual Language Alignment
Xiaoxu Xu, Yitian Yuan, Qiudan Zhang +4
Learning to ground natural language queries to target objects or regions in 3D point clouds is quite essential for 3D scene understanding. Nevertheless, existing 3D visual groundin…
InstructionBench: An Instructional Video Understanding Benchmark
Haiwan Wei, Yitian Yuan, Xiaohan Lan +2
Despite progress in video large language models (Video-LLMs), research on instructional video understanding, crucial for enhancing access to instructional content, remains insuffic…
TimeMarker: A Versatile Video-LLM for Long and Short Video Understanding with Superior Temporal Localization Ability
Shimin Chen, Xiaohan Lan, Yitian Yuan +2
Rapid development of large language models (LLMs) has significantly advanced multimodal large language models (LMMs), particularly in vision-language tasks. However, existing video…