6 papers
MOON3.0: Reasoning-aware Multimodal Representation Learning for E-commerce Product Understanding
Junxian Wu, Chenghan Fu, Zhanheng Nie +6
With the rapid growth of e-commerce, exploring general representations rather than task-specific ones has attracted increasing attention. Although recent multimodal large language…
MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding
Zhanheng Nie, Chenghan Fu, Daoze Zhang +5
Recent Multimodal Large Language Models (MLLMs) have significantly advanced e-commerce product understanding. However, they still face three challenges: (i) the modality imbalance…
Multi-Modal Style Transfer-based Prompt Tuning for Efficient Federated Domain Generalization
Yuliang Chen, Xi Lin, Jun Wu +5
Federated Domain Generalization (FDG) aims to collaboratively train a global model across distributed clients that can generalize well on unseen domains. However, existing FDG meth…
Learning to Hear by Seeing: It's Time for Vision Language Models to Understand Artistic Emotion from Sight and Sound
Dengming Zhang, Weitao You, Jingxiong Li +6
Emotion understanding is critical for making Large Language Models (LLMs) more general, reliable, and aligned with humans. Art conveys emotion through the joint design of visual an…
Controllable Video-to-Music Generation with Multiple Time-Varying Conditions
Junxian Wu, Weitao You, Heda Zuo +3
Music enhances video narratives and emotions, driving demand for automatic video-to-music (V2M) generation. However, existing V2M methods relying solely on visual features or suppl…
GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions
Heda Zuo, Weitao You, Junxian Wu +5
Composing music for video is essential yet challenging, leading to a growing interest in automating music generation for video applications. Existing approaches often struggle to a…