2 papers
cs.CV2026
Enhancing Localized Reasoning for Long Video Understanding via Efficient Segment-to-Video Supervision
Beibei Zhang, Chao Xu, Jun Lan +4
Though Multimodal Large Language Models (MLLMs) have shown impressive potential in video understanding, long video understanding (LVU) remains challenging since distracting noise i…
cs.MM2025
Harnessing Multimodal Large Language Models for Personalized Product Search with Query-aware Refinement
Beibei Zhang, Yanan Lu, Ruobing Xie +4
Personalized product search (PPS) aims to retrieve products relevant to the given query considering user preferences within their purchase histories. Since large language models (L…