Showing cs.IRShow all
2 papers · 1 filter
cs.IR2026
DMAP: Human-Aligned Structural Document Map for Multimodal Document Understanding
ShunLiang Fu, Yanxin Zhang, Yixin Xiang +2
Existing multimodal document question-answering (QA) systems predominantly rely on flat semantic retrieval, representing documents as a set of disconnected text chunks and largely…
cs.IR2025
Frozen LVLMs for Micro-Video Recommendation: A Systematic Study of Feature Extraction and Fusion
Huatuan Sun, Yunshan Ma, Changguang Wu +3
Frozen Large Video Language Models (LVLMs) are increasingly employed in micro-video recommendation due to their strong multimodal understanding. However, their integration lacks sy…