2 papers
cs.CV2025
AIDE: Agentically Improve Visual Language Model with Domain Experts
Ming-Chang Chiu, Fuxiao Liu, Karan Sapra +5
The enhancement of Visual Language Models (VLMs) has traditionally relied on knowledge distillation from larger, more capable models. This dependence creates a fundamental bottlene…
cs.CV2024
MegaCOIN: Enhancing Medium-Grained Color Perception for Vision-Language Models
Ming-Chang Chiu, Shicheng Wen, Pin-Yu Chen +1
In vision-language models (VLMs), the ability to perceive and interpret color and physical environment is crucial for achieving contextually accurate understanding and interaction.…