activity
20242026
collaborators

11 papers

cs.CV2026

Efficient Spatio-Temporal Grounding with Multimodal Large Models via Second-Level Tracking and RL Verification

Tianshu Zhang, Yan Wang, Ji Qi +1

Spatio-temporal grounding in long videos requires precise temporal localization and robust object tracking conditioned on natural-language queries. While recent vision-language mod…

cs.CL2026

Yuvion LLM: An Adversarially-Aware Large Language Model for Content And AI Safety

Ting Ma, Xiufeng Huang, Benlei Cui +43

As large language models are increasingly deployed in real-world systems, safety failures can still lead to harmful outputs and dangerous misuse. We argue that the essence of safet…

cs.CV2026

Yuvion VL: A Multimodal Foundation Model for Adversarial Content and AI Safety

Shikai Qiu, Xiaowen Xu, Benlei Cui +55

General-purpose models often struggle to reliably identify and understand real-world multimodal risks, largely due to the inherent multimodal adversarial nature of content and AI s…

cs.RO2026

ASSCG: Just-Right Gating over Chattering for Fast-Slow LLM Planning in Autonomous Driving

Sining Ang, Yuan Chen, Liu Haiyan +5

Large language models (LLMs) can improve autonomous driving planning but are costly to query online, and existing fast-slow planners often rely on hand-designed triggering rules th…

cs.CV2026

HG-Bench: A Benchmark for Multi-Page Handwritten Answer-Region Grounding in Automated Homework Assessment

Chuangxin Zhao, Boyan Shi, Yanling Wang +7

Automated homework assessment depends not only on recognizing student answers, but also on accurately locating where each answer and each intermediate reasoning step appears in noi…

cs.CV2026

Video2Code: Generating Interactive Webpages from UI Videos via Action-Aware Revisit

Mingde Xu, Zhen Yang, Yan Wang +7

UI videos provide a natural input for generating interactive webpages, as they capture both webpage appearance and action-triggered state transitions. However, directly applying vi…