7 papers
Information-Regularized Attention for Visual-Centric Reasoning
Guohao Sun, Xiaofang Wang, Yash Patel +3
Vision-language models (VLMs) have become a paradigm for multimodal learning, yet remain unstable due to object hallucination, weak visual grounding, and catastrophic forgetting af…
HyCTAS: Multi-Objective Hybrid Convolution-Transformer Architecture Search for Real-Time Image Segmentation
Hongyuan Yu, Cheng Wan, Xiyang Dai +6
Real-time image segmentation demands architectures that preserve fine spatial detail while capturing global context under tight latency and memory budgets. Image segmentation is on…
In-Depth and In-Breadth: Pre-training Multimodal Language Models Customized for Comprehensive Chart Understanding
Wan-Cyuan Fan, Yen-Chun Chen, Mengchen Liu +3
Recent methods for customizing Large Vision Language Models (LVLMs) for domain-specific tasks have shown promising results in scientific chart comprehension. However, existing appr…
On Pre-training of Multimodal Language Models Customized for Chart Understanding
Wan-Cyuan Fan, Yen-Chun Chen, Mengchen Liu +2
Recent studies customizing Multimodal Large Language Models (MLLMs) for domain-specific tasks have yielded promising results, especially in the field of scientific chart comprehens…
Benchmarking Large and Small MLLMs
Xuelu Feng, Yunsheng Li, Dongdong Chen +4
Large multimodal language models (MLLMs) such as GPT-4V and GPT-4o have achieved remarkable advancements in understanding and generating multimodal content, showcasing superior qua…
Exploring Invariance in Images through One-way Wave Equations
Yinpeng Chen, Dongdong Chen, Xiyang Dai +5
In this paper, we empirically reveal an invariance over images-images share a set of one-way wave equations with latent speeds. Each image is uniquely associated with a solution to…