activity
20222025
most citedEmbedded Heterogeneous Attention Transformer for Cross-lingual Image Captioning

24 citations · 31 across the 6 of their papers we have counts for

collaborators

7 papers

cs.CV2025

Rebalancing Contrastive Alignment with Bottlenecked Semantic Increments in Text-Video Retrieval

Jian Xiao, Zijie Song, Jialong Hu +4

Recent progress in text-video retrieval has been largely driven by contrastive learning. However, existing methods often overlook the effect of the modality gap, which causes ancho…

cs.MM2025★ 1 cited

Concept Drift Guided LayerNorm Tuning for Efficient Multimodal Metaphor Identification

Wenhao Qian, Zhenzhen Hu, Zijie Song +1

Metaphorical imagination, the ability to connect seemingly unrelated concepts, is fundamental to human cognition and communication. While understanding linguistic metaphors has adv…

cs.CV2025★ 2 cited

Video Flow as Time Series: Discovering Temporal Consistency and Variability for VideoQA

Zijie Song, Zhenzhen Hu, Yixiao Ma +2

Video Question Answering (VideoQA) is a complex video-language task that demands a sophisticated understanding of both visual content and temporal dynamics. Traditional Transformer…

cs.CV2023★ 4 cited

Grid Jigsaw Representation with CLIP: A New Perspective on Image Clustering

Zijie Song, Zhenzhen Hu, Richang Hong

Unsupervised representation learning for image clustering is essential in computer vision. Although the advancement of visual models has improved image clustering with efficient vi…

cs.IR2023

CDR: Conservative Doubly Robust Learning for Debiased Recommendation

ZiJie Song, JiaWei Chen, Sheng Zhou +4

In recommendation systems (RS), user behavior data is observational rather than experimental, resulting in widespread bias in the data. Consequently, tackling bias has emerged as a…

cs.CV2023★ 24 cited

Embedded Heterogeneous Attention Transformer for Cross-lingual Image Captioning

Zijie Song, Zhenzhen Hu, Yuanen Zhou +3

Cross-lingual image captioning is a challenging task that requires addressing both cross-lingual and cross-modal obstacles in multimedia analysis. The crucial issue in this task is…