41 citations · 114 across the 7 of their papers we have counts for
Showing cs.CVShow all
3 papers · 1 filter
cs.CV2025
Audio-Driven Talking Face Video Generation with Joint Uncertainty Learning
Yifan Xie, Fei Ma, Yi Bin +2
Talking face video generation with arbitrary speech audio is a significant challenge within the realm of digital human technology. The previous studies have emphasized the signific…
cs.CV2023★ 41 cited
Your Negative May not Be True Negative: Boosting Image-Text Matching with False Negative Elimination
Haoxuan Li, Yi Bin, Junrong Liao +2
Most existing image-text matching methods adopt triplet loss as the optimization objective, and choosing a proper negative sample for the triplet of <anchor, positive, negative> is…
cs.CV2023★ 33 cited
Unifying Two-Stream Encoders with Transformers for Cross-Modal Retrieval
Yi Bin, Haoxuan Li, Yahui Xu +3
Most existing cross-modal retrieval methods employ two-stream encoders with different architectures for images and texts, \textit{e.g.}, CNN for images and RNN/Transformer for text…