2 papers
cs.CV2025
Versatile Multimodal Controls for Expressive Talking Human Animation
Zheng Qin, Ruobing Zheng, Yabing Wang +5
In filmmaking, directors typically allow actors to perform freely based on the script before providing specific guidance on how to present key actions. AI-generated content faces s…
cs.CV2024
Multimodal LLM Enhanced Cross-lingual Cross-modal Retrieval
Yabing Wang, Le Wang, Qiang Zhou +4
Cross-lingual cross-modal retrieval (CCR) aims to retrieve visually relevant content based on non-English queries, without relying on human-labeled cross-modal data pairs during tr…