activity
20242026
collaborators

6 papers

cs.CV2026

Chain of World: World Model Thinking in Latent Motion

Fuxiang Yang, Donglin Di, Lulu Tang +6

Vision-Language-Action (VLA) models are a promising path toward embodied intelligence, yet they often overlook the predictive and temporal-causal structure underlying visual dynami…

cs.CV2025

Driving in Spikes: An Entropy-Guided Object Detector for Spike Cameras

Ziyan Liu, Qi Su, Lulu Tang +2

Object detection in autonomous driving suffers from motion blur and saturation under fast motion and extreme lighting. Spike cameras, offer microsecond latency and ultra high dynam…

cs.CR2025

VisualDAN: Exposing Vulnerabilities in VLMs with Visual-Driven DAN Commands

Aofan Liu, Lulu Tang

Vision-Language Models (VLMs) have garnered significant attention for their remarkable ability to interpret and generate multimodal content. However, securing these models against…

cs.CR2025

PiCo: Jailbreaking Multimodal Large Language Models via Pictorial Code Contextualization

Aofan Liu, Lulu Tang, Ting Pan +3

Multimodal Large Language Models (MLLMs), which integrate vision and other modalities into Large Language Models (LLMs), significantly enhance AI capabilities but also introduce ne…

cs.CV2025

You See it, You Got it: Learning 3D Creation on Pose-Free Videos at Scale

Baorui Ma, Huachen Gao, Haoge Deng +4

Recent 3D generation models typically rely on limited-scale 3D `gold-labels' or 2D diffusion priors for 3D content creation. However, their performance is upper-bounded by constrai…

cs.CV2024

Tokenize Anything via Prompting

Ting Pan, Lulu Tang, Xinlong Wang +1

We present a unified, promptable model capable of simultaneously segmenting, recognizing, and captioning anything. Unlike SAM, we aim to build a versatile region representation in…