activity
20242026
collaborators

11 papers

cs.CV2026

AdaThinkV: Adaptive Thinking for Token-Efficient Video Reasoning

Jingqi Tian, Haoji Zhang, Lin Chen +7

Chain-of-thought (CoT) reasoning can improve performance on difficult video questions but often wastes decoding tokens on simple ones. We study whether a video multimodal large lan…

cs.CV2026

Delayed Bidirectional Alignment via Disentangled Audio Semantics for Audio-Visual Segmentation

Jingqi Tian, Yiheng Du, Haoji Zhang +6

Audio-Visual Segmentation (AVS) aims to localize sound-producing objects at the pixel level by integrating auditory and visual cues. However, existing methods often struggle with m…

cs.CV2026

ChatUMM: Robust Context Tracking for Conversational Interleaved Generation

Wenxun Dai, Zhiyuan Zhao, Yule Zhong +12

Unified multimodal models (UMMs) have achieved remarkable progress yet remain constrained by a single-turn interaction paradigm, effectively functioning as solvers for independent…

cs.CV2026

FoodMonitor: Benchmarking MLLMs for Explainable Compliance Analysis

Ruihao Xu, Xingming Shui, Jingxuan Niu +4

As AI-powered compliance monitoring becomes increasingly important in public governance and industrial safety, the ability to provide verifiable evidence and traceable accountabili…

cs.CV2026

Self-Calibrated CLIP for Training-Free Open-Vocabulary Segmentation

Sule Bai, Yong Liu, Yifei Han +4

Recent advancements in pre-trained vision-language models like CLIP have enabled the task of open-vocabulary segmentation. CLIP demonstrates impressive zero-shot capabilities in va…

cs.CV2025

VG-Refiner: Towards Tool-Refined Referring Grounded Reasoning via Agentic Reinforcement Learning

Yuji Wang, Wenlong Liu, Jingxuan Niu +2

Tool-integrated visual reasoning (TiVR) has demonstrated great potential in enhancing multimodal problem-solving. However, existing TiVR paradigms mainly focus on integrating vario…