collaborators

5 papers

cs.CV2026

BabyVision: Visual Reasoning Beyond Language

Liang Chen, Weichu Xie, Yiyan Liang +27

While humans develop core visual skills long before acquiring language, contemporary Multimodal LLMs (MLLMs) still rely heavily on linguistic priors to compensate for their fragile…

cs.CV2026

WorldBench: A Challenging and Visually Diverse Multimodal Reasoning Benchmark

Yida Yin, Harish Krishnakumar, Chung Peng Lee +9

In real-world applications, models are expected to perform reliably across diverse settings. Yet, many existing multimodal benchmarks expand task types without capturing the visual…

cs.LG2026

Muddit: Liberating Generation Beyond Text-to-Image with a Unified Discrete Diffusion Model

Qingyu Shi, Jinbin Bai, Zhuoran Zhao +7

Unified generation models aim to handle diverse tasks across modalities -- such as text generation, image generation, and vision-language reasoning -- within a single architecture…

cs.CV2026

VisionFoundry: Teaching VLMs Visual Perception with Synthetic Images

Guanyu Zhou, Yida Yin, Wenhao Chai +3

Vision-language models (VLMs) still struggle with visual perception tasks such as spatial understanding and viewpoint recognition. One plausible contributing factor is that natural…

cs.CV2025

Multimodal Representation Alignment for Image Generation: Text-Image Interleaved Control Is Easier Than You Think

Liang Chen, Shuai Bai, Wenhao Chai +5

The field of advanced text-to-image generation is witnessing the emergence of unified frameworks that integrate powerful text encoders, such as CLIP and T5, with Diffusion Transfor…