activity
20242026
collaborators

7 papers

cs.CV2026

Information-Regularized Attention for Visual-Centric Reasoning

Guohao Sun, Xiaofang Wang, Yash Patel +3

Vision-language models (VLMs) have become a paradigm for multimodal learning, yet remain unstable due to object hallucination, weak visual grounding, and catastrophic forgetting af…

cs.CV2025

HyCTAS: Multi-Objective Hybrid Convolution-Transformer Architecture Search for Real-Time Image Segmentation

Hongyuan Yu, Cheng Wan, Xiyang Dai +6

Real-time image segmentation demands architectures that preserve fine spatial detail while capturing global context under tight latency and memory budgets. Image segmentation is on…

cs.CL2025

In-Depth and In-Breadth: Pre-training Multimodal Language Models Customized for Comprehensive Chart Understanding

Wan-Cyuan Fan, Yen-Chun Chen, Mengchen Liu +3

Recent methods for customizing Large Vision Language Models (LVLMs) for domain-specific tasks have shown promising results in scientific chart comprehension. However, existing appr…

cs.CV2025

On Pre-training of Multimodal Language Models Customized for Chart Understanding

Wan-Cyuan Fan, Yen-Chun Chen, Mengchen Liu +2

Recent studies customizing Multimodal Large Language Models (MLLMs) for domain-specific tasks have yielded promising results, especially in the field of scientific chart comprehens…

cs.CV2025

Benchmarking Large and Small MLLMs

Xuelu Feng, Yunsheng Li, Dongdong Chen +4

Large multimodal language models (MLLMs) such as GPT-4V and GPT-4o have achieved remarkable advancements in understanding and generating multimodal content, showcasing superior qua…

cs.CV2024

Exploring Invariance in Images through One-way Wave Equations

Yinpeng Chen, Dongdong Chen, Xiyang Dai +5

In this paper, we empirically reveal an invariance over images-images share a set of one-way wave equations with latent speeds. Each image is uniquely associated with a solution to…