collaborators

5 papers

cs.CV2026

Safe Autoregressive Image Generation with Iterative Self-Improving Codebooks

Yunqi Xue, Zhijiang Li, Philip Torr +1

Unlike diffusion-based models that operate in continuous latent spaces, autoregressive unified multimodal models produce images by sequentially predicting discretized visual tokens…

cs.CL2026

The Path Matters: Learning a Token-Commitment Policy for Diffusion Language Models

Bohang Sun, Max Zhu, Francesco Caso +5

Diffusion large language models promise faster generation by refining many token positions in parallel, but this parallelism introduces a hidden control problem: which proposed tok…

cs.CL2026

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs

Andong Hua, Colton Bishop, Igor Mordatch +5

Multimodal large language models (MLLMs) should generate consistent responses given semantically equivalent inputs across modalities. However, we observe a systematic discrepancy i…

cs.CL2025

Flaw or Artifact? Rethinking Prompt Sensitivity in Evaluating LLMs

Andong Hua, Kenan Tang, Chenhe Gu +3

Prompt sensitivity, referring to the phenomenon where paraphrasing (i.e., repeating something written or spoken using different words) leads to significant changes in large languag…

cs.CV2025

Improving Adversarial Transferability in MLLMs via Dynamic Vision-Language Alignment Attack

Chenhe Gu, Jindong Gu, Andong Hua +1

Multimodal Large Language Models (MLLMs), built upon LLMs, have recently gained attention for their capabilities in image recognition and understanding. However, while MLLMs are vu…