7 papers
Multi-Agent Self-Improving Reinforcement Learning for Video Reasoning
Mingwen Zhang, Jisheng Dang, Minqiang Yang +3
Video reasoning tasks such as grounded video question answering and temporal grounding require selecting temporal evidence that supports the query. In many current training setups,…
PhysMLLMs: Spatial Priors for Unified Referring Segmentation and Grounded Reasoning of Images and Videos
Siyao Yan, Bo Han, Jisheng Dang +7
Video multimodal large language models support language guided video segmentation, but they often show spatio temporal inconsistencies, e.g., jitter, drift, and identity switches.…
AffectOmni: RL-Verifiable People-Centric Grounded Affective Reasoning for Social and Art-Related Scenes
Yibo Wang, Rui Yang, Jisheng Dang +7
Multimodal large language models (MLLMs) achieve strong performance on VQA and scene understanding, yet affective reasoning remains vulnerable to shortcut behavior. Models may pred…
A Robust Point Cloud Analysis Framework Inspired By Primary Visual Cortex
Jisheng Dang, Dengyue Pan, Delin Deng +6
Despite significant advancements in point cloud analysis, reducing energy consumption and improving robustness remain understudied, largely due to the inherent limitations of Convo…
MODF-SIR: A Multi-agent Omni-modal Distilled Framework for Social Intelligence Reasoning
Shang Ma, Jisheng Dang, Wencan Zhang +6
We propose a multi-agent collaborative framework built upon a lightweight Multimodal Large Language Model (MLLM), specifically designed for social intelligence reasoning. A key fea…
Decoupled Seg Tokens Make Stronger Reasoning Video Segmenter and Grounder
Dang Jisheng, Wu Xudong, Wang Bimei +7
Existing video segmenter and grounder approaches, exemplified by Sa2VA, directly fuse features within segmentation models. This often results in an undesirable entanglement of dyna…