11 papers
MIRA: Multimodal Iterative Reasoning Agent for Image Editing
Ziyun Zeng, Hang Hua, Jiebo Luo
Instruction-guided image editing offers an intuitive way for users to edit images with natural language. However, diffusion-based editing models often struggle to accurately interp…
StreamSense: Streaming Social Task Detection with Selective Vision-Language Model Routing
Han Wang, Deyi Ji, Lanyun Zhu +2
Live streaming platforms require real-time monitoring and reaction to social signals, utilizing partial and asynchronous evidence from video, text, and audio. We propose StreamSens…
Video-LMM Post-Training: A Deep Dive into Video Reasoning with Large Multimodal Models
Yolo Y. Tang, Jing Bi, Pinxin Liu +24
Video understanding represents the most challenging frontier in computer vision, requiring models to reason about complex spatiotemporal relationships, long-term dependencies, and…
Latent Chain-of-Thought for Visual Reasoning
Guohao Sun, Hang Hua, Jian Wang +5
Chain-of-thought (CoT) reasoning is critical for improving the interpretability and reliability of Large Vision-Language Models (LVLMs). However, existing training algorithms such…
Exploring the Adversarial Vulnerabilities of Vision-Language-Action Models in Robotics
Taowen Wang, Cheng Han, James Chenhao Liang +6
Recently in robotics, Vision-Language-Action (VLA) models have emerged as a transformative approach, enabling robots to execute complex tasks by integrating visual and linguistic i…
SafeCFG: Controlling Harmful Features with Dynamic Safe Guidance for Safe Generation
Jiadong Pan, Liang Li, Hongcheng Gao +3
Diffusion models (DMs) have demonstrated exceptional performance in text-to-image tasks, leading to their widespread use. With the introduction of classifier-free guidance (CFG), t…