4 papers
Can Text-to-Image Models Draw from the Right Frame of Reference?
Zheyuan Gu, Ruihang Li, Yong Huang +5
Spatial instruction following has become a crucial requirement for text-to-image (T2I) generation. A common challenge arises when directional expressions are interpreted under diff…
Test-Time Curriculum for Open-Set AIGC Detection
Yiqian Zhang, Zheyuan Gu, Xiangzhao Hao +8
AI-generated image detectors deployed in open-world environments inevitably face distribution shifts as new and stronger generative models continue to emerge. Although existing met…
OmniVTLA: Vision-Tactile-Language-Action Models with Semantic-Aligned Tactile Sensing
Zhengxue Cheng, Yiqian Zhang, Anni Tang +5
Recent vision-language-action (VLA) models build upon vision-language foundations, and have achieved promising results and exhibit the possibility of task generalization in robot m…
CLEAR: Unlocking Generative Potential for Degraded Image Understanding in Unified Multimodal Models
Xiangzhao Hao, Zefeng Zhang, Zhenyu Zhang +6
Image degradation from blur, noise, compression, and poor illumination severely undermines multimodal understanding in real-world settings. Unified multimodal models that combine u…