3 papers
cs.CV2026
PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet
Xiaopei Wu, Chenshu Hou, Liang Peng +9
3D dense captioning, an emerging vision-language task, aims to generate descriptive sentences for each object in the 3D scene. Despite the impressive results achieved by previous m…
cs.CV2025
Multi-Task Label Discovery via Hierarchical Task Tokens for Partially Annotated Dense Predictions
Jingdong Zhang, Hanrong Ye, Xin Li +2
In recent years, simultaneous learning of multiple dense prediction tasks with partially annotated label data has emerged as an important research area. Previous works primarily fo…
cs.CV2025
MM-Ego: Towards Building Egocentric Multimodal LLMs for Video QA
Hanrong Ye, Haotian Zhang, Erik Daxberger +9
This research aims to comprehensively explore building a multimodal foundation model for egocentric video understanding. To achieve this goal, we work on three fronts. First, as th…