collaborators

5 papers

cs.CV2025

AIDE: Agentically Improve Visual Language Model with Domain Experts

Ming-Chang Chiu, Fuxiao Liu, Karan Sapra +5

The enhancement of Visual Language Models (VLMs) has traditionally relied on knowledge distillation from larger, more capable models. This dependence creates a fundamental bottlene…

cs.CV2025

ColorSense: A Study on Color Vision in Machine Visual Recognition

Ming-Chang Chiu, Yingfei Wang, Derrick Eui Gyu Kim +2

Color vision is essential for human visual perception, but its impact on machine perception is still underexplored. There has been an intensified demand for understanding its role…

cs.CV2024

MegaCOIN: Enhancing Medium-Grained Color Perception for Vision-Language Models

Ming-Chang Chiu, Shicheng Wen, Pin-Yu Chen +1

In vision-language models (VLMs), the ability to perceive and interpret color and physical environment is crucial for achieving contextually accurate understanding and interaction.…

cs.CL2024

Behavioral Bias of Vision-Language Models: A Behavioral Finance View

Yuhang Xiao, Yudi Lin, Ming-Chang Chiu

Large Vision-Language Models (LVLMs) evolve rapidly as Large Language Models (LLMs) was equipped with vision modules to create more human-like models. However, we should carefully…

cs.CV2024

VideoPoet: A Large Language Model for Zero-Shot Video Generation

Dan Kondratyuk, Lijun Yu, Xiuye Gu +28

We present VideoPoet, a language model capable of synthesizing high-quality video, with matching audio, from a large variety of conditioning signals. VideoPoet employs a decoder-on…