collaborators

5 papers

cs.CV2026

HarmVideoBench: Benchmarking Harmful Video Understanding in Large Multimodal Models

Jiajun Wu, Haoyu Kang, Yining Sun +13

Large vision-language models (LVLMs) have recently shown immense potential in automated content moderation, sparking growing interest in developing harmful-video benchmarks. Howeve…

cs.CV2026

VPA-Guard: Defending and Benchmarking Image-to-Video Generation Against Visual Prompt Attacks

Yining Sun, Haoyu Kang, Jiajun Wu +7

Recent advancements in Image-to-Video (I2V) generation have transformed input images from simple appearance references into interactive control interfaces where visual cues such as…

cs.RO2026

RoboRouter: Training-Free Policy Routing for Robotic Manipulation

Yiteng Chen, Zhe Cao, Hongjia Ren +9

Research on robotic manipulation has developed a diverse set of policy paradigms, including vision-language-action (VLA) models, vision-action (VA) policies, and code-based composi…

cs.CV2026

Bridging Creative Intent and Visual Quality: Creator-Driven Recurrent Video Generation with Agentic Feedback Loops

Denis Savytski, Aiden Lei, Heding Liu +4

Generative AI has made content creation increasingly accessible, but many AI-generated videos lack narrative coherence and creative direction, issues that become more substantial a…

cs.CV2024

Towards Small Object Editing: A Benchmark Dataset and A Training-Free Approach

Qihe Pan, Zhen Zhao, Zicheng Wang +5

A plethora of text-guided image editing methods has recently been developed by leveraging the impressive capabilities of large-scale diffusion-based generative models especially St…