collaborators

14 papers

cs.AI2026

V-DEAL: Diagnosing Video Safety De-Calibration as an Understanding-Refusal Coupling Failure

Zhetong Zhang, Honghao Fu, Miao Xu +2

As Video Large Language Models are increasingly deployed in real-world applications, ensuring their safety alignment has become critical. Counterintuitively, we find that harmful v…

cs.AI2026

LenGuard-GPC: Length Guarding with Guided-Prompt Consistency for Spatial Reasoning Reinforce Learning

Xingjian Tao, Yiwei Wang, Yujun Cai +1

Multi-view spatial reasoning requires vision-language models to compare visual evidence across images, align object correspondences, and infer spatial relations over long visual co…

cs.CV2026

What Should a Streaming Video Model Remember?

Haonan Ge, Yiwei Wang, Hang Wu +1

Streaming video understanding models must answer queries at any moment during an ongoing stream, using only what they have observed so far and under fixed memory and computation bu…

cs.CL2026

GeoSVG-RL: Geometry-Aware Reinforcement Learning for Layout-Constrained Text-to-SVG Diagram Generation

Sifan Li, Yujun Cai, Hongkai Chen +1

Generating structured, editable diagrams remains a significant challenge for contemporary large language models, despite their proficiency in general-purpose vector code generation…

cs.CV2026

Mitigating Coordinate Prediction Bias from Positional Encoding Failures

Xingjian Tao, Yiwei Wang, Yujun Cai +3

While Multimodal Large Language Models (MLLMs) excel at general vision-language tasks, precise coordinate prediction remains a significant challenge, particularly as high-resolutio…

cs.CV2026

Video-STAR: Reinforcing Open-Vocabulary Action Recognition with Tools

Zhenlong Yuan, Xiangyan Qu, Chengxuan Qian +8

Multimodal large language models (MLLMs) have demonstrated remarkable potential in bridging visual and textual reasoning, yet their reliance on text-centric priors often limits the…