8 citations · 10 across the 6 of their papers we have counts for
6 papers
OCC-MLLM-Alpha:Empowering Multi-modal Large Language Model for the Understanding of Occluded Objects with Self-Supervised Test-Time Learning
Shuxin Yang, Xinhan Di
There is a gap in the understanding of occluded objects in existing large-scale visual language multi-modal models. Current state-of-the-art multi-modal models fail to provide sati…
OCC-MLLM:Empowering Multimodal Large Language Model For the Understanding of Occluded Objects
Wenmo Qiu, Xinhan Di
There is a gap in the understanding of occluded objects in existing large-scale visual language multi-modal models. Current state-of-the-art multimodal models fail to provide satis…
Self-Supervised Learning of Deviation in Latent Representation for Co-speech Gesture Video Generation
Huan Yang, Jiahui Chen, Chaofan Ding +5
Gestures are pivotal in enhancing co-speech communication. While recent works have mostly focused on point-level motion transformation or fully supervised motion representations th…
Bailing-TTS: Chinese Dialectal Speech Synthesis Towards Human-like Spontaneous Representation
Xinhan Di, Zihao Chen, Yunming Liang +3
Large-scale text-to-speech (TTS) models have made significant progress recently.However, they still fall short in the generation of Chinese dialectal speech. Toaddress this, we pro…
LWA-HAND: Lightweight Attention Hand for Interacting Hand Reconstruction
Xinhan Di, Pengqian Yu
Recent years have witnessed great success for hand reconstruction in real-time applications such as visual reality and augmented reality while interacting with two-hand reconstruct…
3D Reconstruction of Simple Objects from A Single View Silhouette Image
Xinhan Di, Pengqian Yu
While recent deep neural networks have achieved promising results for 3D reconstruction from a single-view image, these rely on the availability of RGB textures in images and extra…