collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2026

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet

Xiaopei Wu, Chenshu Hou, Liang Peng +9

3D dense captioning, an emerging vision-language task, aims to generate descriptive sentences for each object in the 3D scene. Despite the impressive results achieved by previous m…

cs.CV2024

Multi-Task Label Discovery via Hierarchical Task Tokens for Partially Annotated Dense Predictions

Jingdong Zhang, Hanrong Ye, Xin Li +2

In recent years, simultaneous learning of multiple dense prediction tasks with partially annotated label data has emerged as an important research area. Previous works primarily fo…

cs.CV2024

MM-Ego: Towards Building Egocentric Multimodal LLMs for Video QA

Hanrong Ye, Haotian Zhang, Erik Daxberger +9

This research aims to comprehensively explore building a multimodal foundation model for egocentric video understanding. To achieve this goal, we work on three fronts. First, as th…

cs.CV2024

X-VILA: Cross-Modality Alignment for Large Language Model

Hanrong Ye, De-An Huang, Yao Lu +8

We introduce X-VILA, an omni-modality model designed to extend the capabilities of large language models (LLMs) by incorporating image, video, and audio modalities. By aligning mod…

cs.CV2024

DiffusionMTL: Learning Multi-Task Denoising Diffusion Model from Partially Annotated Data

Hanrong Ye, Dan Xu

Recently, there has been an increased interest in the practical problem of learning multiple dense scene understanding tasks from partially annotated data, where each training samp…

cs.CV2023

SegGen: Supercharging Segmentation Models with Text2Mask and Mask2Img Synthesis

Hanrong Ye, Jason Kuen, Qing Liu +3

We propose SegGen, a highly-effective training data generation method for image segmentation, which pushes the performance limits of state-of-the-art segmentation models to a signi…