6 papers
DMP-TTS: Disentangled multi-modal Prompting for Controllable Text-to-Speech with Chained Guidance
Kang Yin, Chunyu Qiang, Sirui Zhao +6
Controllable text-to-speech (TTS) systems face significant challenges in achieving independent manipulation of speaker timbre and speaking style, often suffering from entanglement…
LIHE: Linguistic Instance-Split Hyperbolic-Euclidean Framework for Generalized Weakly-Supervised Referring Expression Comprehension
Xianglong Shi, Silin Cheng, Sirui Zhao +4
Existing Weakly-Supervised Referring Expression Comprehension (WREC) methods, while effective, are fundamentally limited by a one-to-one mapping assumption, hindering their ability…
Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router
Yubo Huang, Weiqiang Wang, Sirui Zhao +3
Recent years have witnessed remarkable advances in audio-driven talking head generation. However, existing approaches predominantly focus on single-character scenarios. While some…
MERba: Multi-Receptive Field MambaVision for Micro-Expression Recognition
Xinglong Mao, Shifeng Liu, Sirui Zhao +4
Micro-expressions (MEs) are brief, involuntary facial movements that reveal genuine emotions, offering valuable insights for psychological assessment and criminal investigations. D…
MER-CLIP: AU-Guided Vision-Language Alignment for Micro-Expression Recognition
Shifeng Liu, Xinglong Mao, Sirui Zhao +3
As a critical psychological stress response, micro-expressions (MEs) are fleeting and subtle facial movements revealing genuine emotions. Automatic ME recognition (MER) holds valua…
MELLM: A Flow-Guided Large Language Model for Micro-Expression Understanding
Sirui Zhao, Zhengye Zhang, Shifeng Liu +5
Micro-expressions (MEs), brief and low-intensity facial movements revealing concealed emotions, are crucial for affective computing. Despite notable progress in ME recognition, exi…