9 papers
LaTtE-Flow: Layerwise Timestep-Expert Flow-based Transformer
Ying Shen, Zhiyang Xu, Jiuhai Chen +6
Recent advances in multimodal foundation models unifying image understanding and generation have opened exciting avenues for tackling a wide range of vision-language tasks within a…
Multi-Turn Reflective Masking Elicits Reasoning in Mask Diffusion Models
Yanming Zhang, Yihan Bian, Jingyuan Qi +3
While reasoning on autoregressive (AR) models is often performed by chain-of-thought reasoning and reflection, their refinement of previous outputs still relies on fully sequential…
ControlMap: Controllable High-Definition Map Generation for Traffic Scenario Simulation
Marwan Farag, Steffen Wäldele, Yu Yao
Simulation is central to validating autonomous driving systems, yet current pipelines are limited by insufficient scenario diversity due to costly High Definition (HD) map creation…
SmartCLIP: Modular Vision-language Alignment with Identification Guarantees
Shaoan Xie, Lingjing Kong, Yujia Zheng +5
Contrastive Language-Image Pre-training (CLIP)~\citep{radford2021learning} has emerged as a pivotal model in computer vision and multimodal learning, achieving state-of-the-art per…
Jailbreaks on Vision Language Model via Multimodal Reasoning
Aarush Noheria, Yuguang Yao
Vision-language models (VLMs) have become central to tasks such as visual question answering, image captioning, and text-to-image generation. However, their outputs are highly sens…
SuperFlow: Training Flow Matching Models with RL on the Fly
Kaijie Chen, Zhiyang Xu, Ying Shen +3
Recent progress in flow-based generative models and reinforcement learning (RL) has improved text-image alignment and visual quality. However, current RL training for flow models s…