3 papers
cs.CV2026
WorldBench: A Challenging and Visually Diverse Multimodal Reasoning Benchmark
Yida Yin, Harish Krishnakumar, Chung Peng Lee +9
In real-world applications, models are expected to perform reliably across diverse settings. Yet, many existing multimodal benchmarks expand task types without capturing the visual…
cs.CV2026
VisionFoundry: Teaching VLMs Visual Perception with Synthetic Images
Guanyu Zhou, Yida Yin, Wenhao Chai +3
Vision-language models (VLMs) still struggle with visual perception tasks such as spatial understanding and viewpoint recognition. One plausible contributing factor is that natural…
cs.CV2025
Multimodal Representation Alignment for Image Generation: Text-Image Interleaved Control Is Easier Than You Think
Liang Chen, Shuai Bai, Wenhao Chai +5
The field of advanced text-to-image generation is witnessing the emergence of unified frameworks that integrate powerful text encoders, such as CLIP and T5, with Diffusion Transfor…