4 papers
S2FGL: Spatial Spectral Federated Graph Learning
Zihan Tan, Suyuan Huang, Guancheng Wan +3
Federated Graph Learning (FGL) combines the privacy-preserving capabilities of federated learning (FL) with the strong graph modeling capability of Graph Neural Networks (GNNs). Cu…
From Image to Video, what do we need in multimodal LLMs?
Suyuan Huang, Haoxin Zhang, Linqing Zhong +4
Covering from Image LLMs to the more complex Video LLMs, the Multimodal Large Language Models (MLLMs) have demonstrated profound capabilities in comprehending cross-modal informati…
ScalingNote: Scaling up Retrievers with Large Language Models for Real-World Dense Retrieval
Suyuan Huang, Chao Zhang, Yuanyuan Wu +12
Dense retrieval in most industries employs dual-tower architectures to retrieve query-relevant documents. Due to online deployment requirements, existing real-world dense retrieval…
Vript: A Video Is Worth Thousands of Words
Dongjie Yang, Suyuan Huang, Chengqiang Lu +5
Advancements in multimodal learning, particularly in video understanding and generation, require high-quality video-text datasets for improved model performance. Vript addresses th…