6 citations · 7 across the 4 of their papers we have counts for
4 papers · 1 filter
View-on-Graph: Zero-shot 3D Visual Grounding via Vision-Language Reasoning on Scene Graphs
Yuanyuan Liu, Haiyang Mei, Dongyang Zhan +4
3D visual grounding (3DVG) identifies objects in 3D scenes from language descriptions. Existing zero-shot approaches leverage 2D vision-language models (VLMs) by converting 3D spat…
Q-CLIP: Unleashing the Power of Vision-Language Models for Video Quality Assessment through Unified Cross-Modal Adaptation
Yachun Mi, Yu Li, Yanting Li +6
Accurate and efficient Video Quality Assessment (VQA) has long been a key research challenge. Current mainstream VQA methods typically improve performance by pretraining on large-s…
Explore Contextual Information for 3D Scene Graph Generation
Yuanyuan Liu, Chengjiang Long, Zhaoxuan Zhang +4
3D scene graph generation (SGG) has been of high interest in computer vision. Although the accuracy of 3D SGG on coarse classification and single relation label has been gradually…
Pyramid Network with Online Hard Example Mining for Accurate Left Atrium Segmentation
Cheng Bian, Xin Yang, Jianqiang Ma +5
Accurately segmenting left atrium in MR volume can benefit the ablation procedure of atrial fibrillation. Traditional automated solutions often fail in relieving experts from the l…