collaborators

9 papers

cs.CV2026

LaTtE-Flow: Layerwise Timestep-Expert Flow-based Transformer

Ying Shen, Zhiyang Xu, Jiuhai Chen +6

Recent advances in multimodal foundation models unifying image understanding and generation have opened exciting avenues for tackling a wide range of vision-language tasks within a…

cs.CL2026

Multi-Turn Reflective Masking Elicits Reasoning in Mask Diffusion Models

Yanming Zhang, Yihan Bian, Jingyuan Qi +3

While reasoning on autoregressive (AR) models is often performed by chain-of-thought reasoning and reflection, their refinement of previous outputs still relies on fully sequential…

cs.RO2026

ControlMap: Controllable High-Definition Map Generation for Traffic Scenario Simulation

Marwan Farag, Steffen Wäldele, Yu Yao

Simulation is central to validating autonomous driving systems, yet current pipelines are limited by insufficient scenario diversity due to costly High Definition (HD) map creation…

cs.CV2026

SmartCLIP: Modular Vision-language Alignment with Identification Guarantees

Shaoan Xie, Lingjing Kong, Yujia Zheng +5

Contrastive Language-Image Pre-training (CLIP)~\citep{radford2021learning} has emerged as a pivotal model in computer vision and multimodal learning, achieving state-of-the-art per…

cs.CV2026

Jailbreaks on Vision Language Model via Multimodal Reasoning

Aarush Noheria, Yuguang Yao

Vision-language models (VLMs) have become central to tasks such as visual question answering, image captioning, and text-to-image generation. However, their outputs are highly sens…

cs.CV2026

SuperFlow: Training Flow Matching Models with RL on the Fly

Kaijie Chen, Zhiyang Xu, Ying Shen +3

Recent progress in flow-based generative models and reinforcement learning (RL) has improved text-image alignment and visual quality. However, current RL training for flow models s…