5 papers
Safe Autoregressive Image Generation with Iterative Self-Improving Codebooks
Yunqi Xue, Zhijiang Li, Philip Torr +1
Unlike diffusion-based models that operate in continuous latent spaces, autoregressive unified multimodal models produce images by sequentially predicting discretized visual tokens…
The Path Matters: Learning a Token-Commitment Policy for Diffusion Language Models
Bohang Sun, Max Zhu, Francesco Caso +5
Diffusion large language models promise faster generation by refining many token positions in parallel, but this parallelism introduces a hidden control problem: which proposed tok…
TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs
Andong Hua, Colton Bishop, Igor Mordatch +5
Multimodal large language models (MLLMs) should generate consistent responses given semantically equivalent inputs across modalities. However, we observe a systematic discrepancy i…
Flaw or Artifact? Rethinking Prompt Sensitivity in Evaluating LLMs
Andong Hua, Kenan Tang, Chenhe Gu +3
Prompt sensitivity, referring to the phenomenon where paraphrasing (i.e., repeating something written or spoken using different words) leads to significant changes in large languag…
Improving Adversarial Transferability in MLLMs via Dynamic Vision-Language Alignment Attack
Chenhe Gu, Jindong Gu, Andong Hua +1
Multimodal Large Language Models (MLLMs), built upon LLMs, have recently gained attention for their capabilities in image recognition and understanding. However, while MLLMs are vu…