5 papers
AIDE: Agentically Improve Visual Language Model with Domain Experts
Ming-Chang Chiu, Fuxiao Liu, Karan Sapra +5
The enhancement of Visual Language Models (VLMs) has traditionally relied on knowledge distillation from larger, more capable models. This dependence creates a fundamental bottlene…
ColorSense: A Study on Color Vision in Machine Visual Recognition
Ming-Chang Chiu, Yingfei Wang, Derrick Eui Gyu Kim +2
Color vision is essential for human visual perception, but its impact on machine perception is still underexplored. There has been an intensified demand for understanding its role…
MegaCOIN: Enhancing Medium-Grained Color Perception for Vision-Language Models
Ming-Chang Chiu, Shicheng Wen, Pin-Yu Chen +1
In vision-language models (VLMs), the ability to perceive and interpret color and physical environment is crucial for achieving contextually accurate understanding and interaction.…
Behavioral Bias of Vision-Language Models: A Behavioral Finance View
Yuhang Xiao, Yudi Lin, Ming-Chang Chiu
Large Vision-Language Models (LVLMs) evolve rapidly as Large Language Models (LLMs) was equipped with vision modules to create more human-like models. However, we should carefully…
VideoPoet: A Large Language Model for Zero-Shot Video Generation
Dan Kondratyuk, Lijun Yu, Xiuye Gu +28
We present VideoPoet, a language model capable of synthesizing high-quality video, with matching audio, from a large variety of conditioning signals. VideoPoet employs a decoder-on…