most citedExtending Visual Dynamics for Video-to-Music Generation

1 citations · 1 across the 4 of their papers we have counts for

collaborators

7 papers

cs.IR2025

AUV-Fusion: Cross-Modal Adversarial Fusion of User Interactions and Visual Perturbations Against VARS

Hai Ling, Tianchi Wang, Xiaohao Liu +3

Modern Visual-Aware Recommender Systems (VARS) exploit the integration of user interaction data and visual features to deliver personalized recommendations with high precision. How…

cs.IR2025

LLM2Rec: Large Language Models Are Powerful Embedding Models for Sequential Recommendation

Yingzhi He, Xiaohao Liu, An Zhang +2

Sequential recommendation aims to predict users' future interactions by modeling collaborative filtering (CF) signals from historical behaviors of similar users or items. Tradition…

cs.CL2025

L-MTP: Leap Multi-Token Prediction Beyond Adjacent Context for Large Language Models

Xiaohao Liu, Xiaobo Xia, Weixiang Zhao +6

Large language models (LLMs) have achieved notable progress. Despite their success, next-token prediction (NTP), the dominant method for LLM training and inference, is constrained…

cs.MM20251 cited

Extending Visual Dynamics for Video-to-Music Generation

Xiaohao Liu, Teng Tu, Yunshan Ma +1

Music profoundly enhances video production by improving quality, engagement, and emotional resonance, sparking growing interest in video-to-music generation. Despite recent advance…

cs.LG2025

Continual Multimodal Contrastive Learning

Xiaohao Liu, Xiaobo Xia, See-Kiong Ng +1

Multimodal Contrastive Learning (MCL) advances in aligning different modalities and generating multimodal representations in a joint space. By leveraging contrastive learning acros…

cs.CV2024

Towards Modality Generalization: A Benchmark and Prospective Analysis

Xiaohao Liu, Xiaobo Xia, Zhuo Huang +2

Multi-modal learning has achieved remarkable success by integrating information from various modalities, achieving superior performance in tasks like recognition and retrieval comp…