3 papers
cs.CV2026
An Exam for Active Observers
Jiarui Zhang, Muzi Tao, Shangshang Wang +3
Human vision is a closed loop: gaze is continuously redirected by intermediate hypotheses rather than a single snapshot. Decades of psychophysics and cognitive science have argued…
cs.CV2026
Asymmetric Idiosyncrasies in Multimodal Models
Muzi Tao, Chufan Shi, Huijuan Wang +2
In this work, we study idiosyncrasies in the caption models and their downstream impact on text-to-image models. We design a systematic analysis: given either a generated caption o…
cs.CL2026
UReason: Benchmarking Reasoning-to-Generation Alignment in Unified Multimodal Models
Cheng Yang, Chufan Shi, Bo Shui +7
Unified multimodal models (UMMs) aim to integrate multimodal understanding and generation within a unified architecture, yet it remains unclear to what extent textual and visual mo…