activity
20172024
most cited3D Reconstruction of Simple Objects from A Single View Silhouette Image

8 citations · 10 across the 6 of their papers we have counts for

collaborators

6 papers

cs.CV2024

OCC-MLLM-Alpha:Empowering Multi-modal Large Language Model for the Understanding of Occluded Objects with Self-Supervised Test-Time Learning

Shuxin Yang, Xinhan Di

There is a gap in the understanding of occluded objects in existing large-scale visual language multi-modal models. Current state-of-the-art multi-modal models fail to provide sati…

cs.CV20241 cited

OCC-MLLM:Empowering Multimodal Large Language Model For the Understanding of Occluded Objects

Wenmo Qiu, Xinhan Di

There is a gap in the understanding of occluded objects in existing large-scale visual language multi-modal models. Current state-of-the-art multimodal models fail to provide satis…

cs.CV2024

Self-Supervised Learning of Deviation in Latent Representation for Co-speech Gesture Video Generation

Huan Yang, Jiahui Chen, Chaofan Ding +5

Gestures are pivotal in enhancing co-speech communication. While recent works have mostly focused on point-level motion transformation or fully supervised motion representations th…

cs.CL20241 cited

Bailing-TTS: Chinese Dialectal Speech Synthesis Towards Human-like Spontaneous Representation

Xinhan Di, Zihao Chen, Yunming Liang +3

Large-scale text-to-speech (TTS) models have made significant progress recently.However, they still fall short in the generation of Chinese dialectal speech. Toaddress this, we pro…

cs.CV2022

LWA-HAND: Lightweight Attention Hand for Interacting Hand Reconstruction

Xinhan Di, Pengqian Yu

Recent years have witnessed great success for hand reconstruction in real-time applications such as visual reality and augmented reality while interacting with two-hand reconstruct…

cs.CV20178 cited

3D Reconstruction of Simple Objects from A Single View Silhouette Image

Xinhan Di, Pengqian Yu

While recent deep neural networks have achieved promising results for 3D reconstruction from a single-view image, these rely on the availability of RGB textures in images and extra…