activity
20242026
collaborators

5 papers

cs.CV2026

When Relations Break: Analyzing Relation Hallucination in Vision-Language Model Under Rotation and Noise

Philip Wootaek Shin, Ajay Narayanan Sridhar, Sivani Devarapalli +3

Vision-language models (VLMs) achieve strong multimodal performance but remain prone to relation hallucination, which requires accurate reasoning over inter-object interactions. We…

cs.CV2025

MammothModa2: A Unified AR-Diffusion Framework for Multimodal Understanding and Generation

Tao Shen, Xin Wan, Taicai Chen +10

Unified multimodal models aim to integrate understanding and generation within a single framework, yet bridging the gap between discrete semantic reasoning and high-fidelity visual…

cs.CV2025

TimeSearch-R: Adaptive Temporal Search for Long-Form Video Understanding via Self-Verification Reinforcement Learning

Junwen Pan, Qizhe Zhang, Rui Zhang +5

Temporal search aims to identify a minimal set of relevant frames from tens of thousands based on a given query, serving as a foundation for accurate long-form video understanding.…

cs.CV2025

ZoomV: Temporal Zoom-in for Efficient Long Video Understanding

Yuan Zhang, Junwen Pan, Rui Zhang +5

Long video understanding poses a fundamental challenge for large video-language models (LVLMs) due to the overwhelming number of frames and the risk of losing essential context thr…

cs.CV2024

MammothModa: Multi-Modal Large Language Model

Qi She, Junwen Pan, Xin Wan +3

In this report, we introduce MammothModa, yet another multi-modal large language model (MLLM) designed to achieve state-of-the-art performance starting from an elementary baseline.…