5 citations · 7 across the 3 of their papers we have counts for
3 papers
cs.CV2022
Dilated Context Integrated Network with Cross-Modal Consensus for Temporal Emotion Localization in Videos
Juncheng Li, Junlin Xie, Linchao Zhu +8
Understanding human emotions is a crucial ability for intelligent robots to provide better human-robot interactions. The existing works are limited to trimmed video-level emotion c…
cs.AI2022★ 5 cited
BOSS: Bottom-up Cross-modal Semantic Composition with Hybrid Counterfactual Training for Robust Content-based Image Retrieval
Wenqiao Zhang, Jiannan Guo, Mengze Li +5
Content-Based Image Retrieval (CIR) aims to search for a target image by concurrently comprehending the composition of an example image and a complementary text, which potentially…
cs.CV2021★ 2 cited
Consensus Graph Representation Learning for Better Grounded Image Captioning
Wenqiao Zhang, Haochen Shi, Siliang Tang +3
The contemporary visual captioning models frequently hallucinate objects that are not actually in a scene, due to the visual misclassification or over-reliance on priors that resul…