4 papers
Frozen LVLMs for Micro-Video Recommendation: A Systematic Study of Feature Extraction and Fusion
Huatuan Sun, Yunshan Ma, Changguang Wu +3
Frozen Large Video Language Models (LVLMs) are increasingly employed in micro-video recommendation due to their strong multimodal understanding. However, their integration lacks sy…
MAB-DQA: Addressing Query Aspect Importance in Document Question Answering with Multi-Armed Bandits
Yixin Xiang, Yunshan Ma, Xiaoyu Du +3
Document Question Answering (DQA) involves generating answers from a document based on a user's query, representing a key task in document understanding. This task requires interpr…
DMAP: Human-Aligned Structural Document Map for Multimodal Document Understanding
ShunLiang Fu, Yanxin Zhang, Yixin Xiang +2
Existing multimodal document question-answering (QA) systems predominantly rely on flat semantic retrieval, representing documents as a set of disconnected text chunks and largely…
SpotDiff: Spotting and Disentangling Interference in Feature Space for Subject-Preserving Image Generation
Yongzhi Li, Saining Zhang, Yibing Chen +3
Personalized image generation aims to faithfully preserve a reference subject's identity while adapting to diverse text prompts. Existing optimization-based methods ensure high fid…