8 citations · 9 across the 10 of their papers we have counts for
15 papers · 1 filter
Video-IFBench: Evaluating Instruction Following of Multimodal LLMs in Video Understanding Scenarios
Hongbo Liu, Peixian Chen, Sihan Liu +12
Multimodal Large Language Models (MLLMs) have shown strong performance in video understanding. However, their ability to follow instructions in this domain remains under-explored.…
STBridge: Shared-Target Alignment for Bridging Understanding and Generation in UMMs
Ye Wang, Hongjun Wang, Hao Fang +7
Unified multimodal models (UMMs) aim to integrate visual understanding and generation within a single architecture, but architectural unification alone does not ensure semantic con…
OmniView-Space: Reinforcing Spatial Reasoning via Multi-Perspective Spatial Mapping
Xudong Li, Mengdan Zhang, Peixian Chen +7
Spatial intelligence remains a persistent challenge for Multimodal Large Language Models (MLLMs), as it requires coherent spatial scene representations beyond basic object recognit…
Streaming Video Instruction Tuning
Jiaer Xia, Peixian Chen, Mengdan Zhang +2
We present Streamo, a real-time streaming video LLM that serves as a general-purpose interactive assistant. Unlike existing online video models that focus narrowly on question answ…
PromptMoE: Generalizable Zero-Shot Anomaly Detection via Visually-Guided Prompt Mixtures
Yuheng Shao, Lizhang Wang, Changhao Li +2
Zero-Shot Anomaly Detection (ZSAD) aims to identify and localize anomalous regions in images of unseen object classes. While recent methods based on vision-language models like CLI…
Zooming from Context to Cue: Hierarchical Preference Optimization for Multi-Image MLLMs
Xudong Li, Mengdan Zhang, Peixian Chen +8
Multi-modal Large Language Models (MLLMs) excel at single-image tasks but struggle with multi-image understanding due to cross-modal misalignment, leading to hallucinations (contex…