most citedThe Multimodal Information based Speech Processing (MISP) 2022 Challenge: Audio-Visual Diarization and Recognition

1 citations · 2 across the 7 of their papers we have counts for

collaborators

7 papers

cs.CV2024

StyleTokenizer: Defining Image Style by a Single Instance for Controlling Diffusion Models

Wen Li, Muyuan Fang, Cheng Zou +5

Despite the burst of innovative methods for controlling the diffusion process, effectively controlling image styles in text-to-image generation remains a challenging task. Many ada…

cs.CV2024

Learning Dynamic Tetrahedra for High-Quality Talking Head Synthesis

Zicheng Zhang, Ruobing Zheng, Ziwen Liu +7

Recent works in implicit representations, such as Neural Radiance Fields (NeRF), have advanced the generation of realistic and animatable head avatars from video sequences. These i…

cs.CV2024

M2-Encoder: Advancing Bilingual Image-Text Understanding by Large-scale Efficient Pretraining

Qingpei Guo, Furong Xu, Hanxiao Zhang +6

Vision-language foundation models like CLIP have revolutionized the field of artificial intelligence. Nevertheless, VLM models supporting multi-language, e.g., in both Chinese and…

cs.SD2024

Independent low-rank matrix analysis based on the Sinkhorn divergence source model for blind source separation

Jianyu Wang, Shanzheng Guan, Jingdong Chen +1

The so-called independent low-rank matrix analysis (ILRMA) has demonstrated a great potential for dealing with the problem of determined blind source separation (BSS) for audio and…

eess.AS20231 cited

The Multimodal Information Based Speech Processing (MISP) 2023 Challenge: Audio-Visual Target Speaker Extraction

Shilong Wu, Chenxi Wang, Hang Chen +13

Previous Multimodal Information based Speech Processing (MISP) challenges mainly focused on audio-visual speech recognition (AVSR) with commendable success. However, the most advan…

cs.CV2023

Mapping EEG Signals to Visual Stimuli: A Deep Learning Approach to Match vs. Mismatch Classification

Yiqian Yang, Zhengqiao Zhao, Qian Wang +2

Existing approaches to modeling associations between visual stimuli and brain responses are facing difficulties in handling between-subject variance and model generalization. Inspi…