2 citations · 3 across the 3 of their papers we have counts for
4 papers
Sparse Shortcuts: Facilitating Efficient Fusion in Multimodal Large Language Models
Jingrui Zhang, Feng Liang, Yong Zhang +3
With the remarkable success of large language models (LLMs) in natural language understanding and generation, multimodal large language models (MLLMs) have rapidly advanced in thei…
OVG-HQ: Online Video Grounding with Hybrid-modal Queries
Runhao Zeng, Jiaqi Mao, Minghao Lai +5
Video grounding (VG) task focuses on locating specific moments in a video based on a query, usually in text form. However, traditional VG struggles with some scenarios like streami…
COHESION: Composite Graph Convolutional Network with Dual-Stage Fusion for Multimodal Recommendation
Jinfeng Xu, Zheyu Chen, Wei Wang +3
Recent works in multimodal recommendations, which leverage diverse modal information to address data sparsity and enhance recommendation accuracy, have garnered considerable intere…
A Survey on Multimodal Recommender Systems: Recent Advances and Future Directions
Jinfeng Xu, Zheyu Chen, Shuo Yang +5
Acquiring valuable data from the rapidly expanding information on the internet has become a significant concern, and recommender systems have emerged as a widely used and effective…