collaborators

10 papers

cs.CV2026

STBridge: Shared-Target Alignment for Bridging Understanding and Generation in UMMs

Ye Wang, Hongjun Wang, Hao Fang +7

Unified multimodal models (UMMs) aim to integrate visual understanding and generation within a single architecture, but architectural unification alone does not ensure semantic con…

cs.CV2026

OmniView-Space: Reinforcing Spatial Reasoning via Multi-Perspective Spatial Mapping

Xudong Li, Mengdan Zhang, Peixian Chen +7

Spatial intelligence remains a persistent challenge for Multimodal Large Language Models (MLLMs), as it requires coherent spatial scene representations beyond basic object recognit…

cs.CV2026

Streaming Video Instruction Tuning

Jiaer Xia, Peixian Chen, Mengdan Zhang +2

We present Streamo, a real-time streaming video LLM that serves as a general-purpose interactive assistant. Unlike existing online video models that focus narrowly on question answ…

cs.CV2026

Towards Artwork Explanation in Large-scale Vision Language Models

Kazuki Hayashi, Yusuke Sakai, Hidetaka Kamigaito +2

Large-scale Vision-Language Models (LVLMs) output text from images and instructions, demonstrating capabilities in text generation and comprehension. However, it has not been clari…

cs.CV2025

PromptMoE: Generalizable Zero-Shot Anomaly Detection via Visually-Guided Prompt Mixtures

Yuheng Shao, Lizhang Wang, Changhao Li +2

Zero-Shot Anomaly Detection (ZSAD) aims to identify and localize anomalous regions in images of unseen object classes. While recent methods based on vision-language models like CLI…

cs.CV2025

Long-VITA: Scaling Large Multi-modal Models to 1 Million Tokens with Leading Short-Context Accuracy

Yunhang Shen, Chaoyou Fu, Shaoqi Dong +14

We introduce Long-VITA, a simple yet effective large multi-modal model for long-context visual-language understanding tasks. It is adept at concurrently processing and analyzing mo…