papers

Publications (25)

cs.CV2025

Video-Bench: Human-Aligned Video Generation Benchmark

Hui Han, Siyuan Li, Jiaqi Chen +10

Video generation assessment is essential for ensuring that generative models produce visually realistic, high-quality videos while aligning with human expectations. Current video g…

cs.IR2026

The 2nd EReL@MIR Workshop on Efficient Representation Learning for Multimodal Information Retrieval

Junchen Fu, Xuri Ge, Xin Xin +6

Multimodal representation learning has attracted increasing attention in AI, driven by the strong performance of large, pretrained multimodal foundation models such as Qwen, LLaVA,…

cs.IR2024

NineRec: A Benchmark Dataset Suite for Evaluating Transferable Recommendation

Jiaqi Zhang, Yu Cheng, Yongxin Ni +6

Large foundational models, through upstream pre-training and downstream fine-tuning, have achieved immense success in the broad AI community due to improved model performance and s…

cs.IR2026

Stream-aware Side Adaptation for Large Pre-trained Multimodal Embedding Models in Sequential Recommendation

Junchen Fu, Kaiwen Zheng, Ioannis Arapakis +4

Recently, large pretrained multimodal embedding models such as Qwen3-VL Embedding have shown strong promise for sequential recommendation, as they provide reusable semantic item re…

cs.MM2026

Benchmarking Multimodal Large Language Models for Missing Modality Completion in Product Catalogues

Junchen Fu, Wenhao Deng, Kaiwen Zheng +5

Missing-modality information on e-commerce platforms, such as absent product images or textual descriptions, often arises from annotation errors or incomplete metadata, impairing b…

cs.CL2026

LLMPopcorn: Exploring LLMs as Assistants for Popular Micro-video Generation

Junchen Fu, Xuri Ge, Kaiwen Zheng +5

In an era where micro-videos dominate platforms like TikTok and YouTube, AI-generated content is nearing cinematic quality. The next frontier is using large language models (LLMs)…