3 papers
cs.CV2026
PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet
Xiaopei Wu, Chenshu Hou, Liang Peng +9
3D dense captioning, an emerging vision-language task, aims to generate descriptive sentences for each object in the 3D scene. Despite the impressive results achieved by previous m…
cs.CV2026
Heterogeneous and Adept Snapshot Distillation for 3D Semantic Segmentation
Xiaopei Wu, Yuenan Hou, Junkai Xu +7
Multi-modal fusion and multi-model ensembling are prevalent in enhancing the performance of 3D semantic segmentation. Despite the impressive performance, these methods either rely…
cs.CV2025
Do MLLMs Exhibit Human-like Perceptual Behaviors? HVSBench: A Benchmark for MLLM Alignment with Human Perceptual Behavior
Jiaying Lin, Shuquan Ye, Dan Xu +2
While Multimodal Large Language Models (MLLMs) excel at many vision tasks, it is unknown if they exhibit human-like perceptual behaviors. To evaluate this, we introduce HVSBench, t…