activity
20242026
collaborators
Showing cs.CVShow all

8 papers · 1 filter

cs.CV2026

Efficient Spatio-Temporal Grounding with Multimodal Large Models via Second-Level Tracking and RL Verification

Tianshu Zhang, Yan Wang, Ji Qi +1

Spatio-temporal grounding in long videos requires precise temporal localization and robust object tracking conditioned on natural-language queries. While recent vision-language mod…

cs.CV2026

Yuvion VL: A Multimodal Foundation Model for Adversarial Content and AI Safety

Shikai Qiu, Xiaowen Xu, Benlei Cui +55

General-purpose models often struggle to reliably identify and understand real-world multimodal risks, largely due to the inherent multimodal adversarial nature of content and AI s…

cs.CV2026

HG-Bench: A Benchmark for Multi-Page Handwritten Answer-Region Grounding in Automated Homework Assessment

Chuangxin Zhao, Boyan Shi, Yanling Wang +7

Automated homework assessment depends not only on recognizing student answers, but also on accurately locating where each answer and each intermediate reasoning step appears in noi…

cs.CV2026

Video2Code: Generating Interactive Webpages from UI Videos via Action-Aware Revisit

Mingde Xu, Zhen Yang, Yan Wang +7

UI videos provide a natural input for generating interactive webpages, as they capture both webpage appearance and action-triggered state transitions. However, directly applying vi…

cs.CV2026

GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents

V Team, Wenyi Hong, Xiaotao Gu +94

We present GLM-5V-Turbo, a step toward native foundation models for multimodal agents. As foundation models are increasingly deployed in real environments, agentic capability depen…

cs.CV2026

ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents

Fanqing Meng, Lingxiao Du, Zijian Wu +46

Language-model agents are increasingly used as persistent coworkers that assist users across multiple working days. During such workflows, the surrounding environment may change in…