Publications (25)
Video-Bench: Human-Aligned Video Generation Benchmark
Hui Han, Siyuan Li, Jiaqi Chen +10
Video generation assessment is essential for ensuring that generative models produce visually realistic, high-quality videos while aligning with human expectations. Current video g…
The 2nd EReL@MIR Workshop on Efficient Representation Learning for Multimodal Information Retrieval
Junchen Fu, Xuri Ge, Xin Xin +6
Multimodal representation learning has attracted increasing attention in AI, driven by the strong performance of large, pretrained multimodal foundation models such as Qwen, LLaVA,…
NineRec: A Benchmark Dataset Suite for Evaluating Transferable Recommendation
Jiaqi Zhang, Yu Cheng, Yongxin Ni +6
Large foundational models, through upstream pre-training and downstream fine-tuning, have achieved immense success in the broad AI community due to improved model performance and s…
Stream-aware Side Adaptation for Large Pre-trained Multimodal Embedding Models in Sequential Recommendation
Junchen Fu, Kaiwen Zheng, Ioannis Arapakis +4
Recently, large pretrained multimodal embedding models such as Qwen3-VL Embedding have shown strong promise for sequential recommendation, as they provide reusable semantic item re…
Benchmarking Multimodal Large Language Models for Missing Modality Completion in Product Catalogues
Junchen Fu, Wenhao Deng, Kaiwen Zheng +5
Missing-modality information on e-commerce platforms, such as absent product images or textual descriptions, often arises from annotation errors or incomplete metadata, impairing b…
LLMPopcorn: Exploring LLMs as Assistants for Popular Micro-video Generation
Junchen Fu, Xuri Ge, Kaiwen Zheng +5
In an era where micro-videos dominate platforms like TikTok and YouTube, AI-generated content is nearing cinematic quality. The next frontier is using large language models (LLMs)…