7 citations · 8 across the 7 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2025
Referring Expression Instance Retrieval and A Strong End-to-End Baseline
Xiangzhao Hao, Kuan Zhu, Hongyu Guo +5
Using natural language to query visual information is a fundamental need in real-world applications. Text-Image Retrieval (TIR) retrieves a target image from a gallery based on an…
cs.CV2023
PMMTalk: Speech-Driven 3D Facial Animation from Complementary Pseudo Multi-modal Features
Tianshun Han, Shengnan Gui, Yiqing Huang +9
Speech-driven 3D facial animation has improved a lot recently while most related works only utilize acoustic modality and neglect the influence of visual and textual cues, leading…