4 citations · 9 across the 26 of their papers we have counts for
4 papers · 1 filter
Subspace Alignment for Vision-Language Model Test-time Adaptation
Zhichen Zeng, Wenxuan Bao, Xiao Lin +8
Vision-language models (VLMs), despite their extraordinary zero-shot capabilities, are vulnerable to distribution shifts. Test-time adaptation (TTA) emerges as a predominant strate…
DenseAnnotate: Enabling Scalable Dense Caption Collection for Images and 3D Scenes via Spoken Descriptions
Xiaoyu Lin, Aniket Ghorpade, Hansheng Zhu +9
With the rapid adoption of multimodal large language models (MLLMs) across diverse applications, there is a pressing need for task-centered, high-quality training data. A key limit…
UI-UG: A Unified MLLM for UI Understanding and Generation
Hao Yang, Weijie Qiu, Ru Zhang +8
Although Multimodal Large Language Models (MLLMs) have been widely applied across domains, they are still facing challenges in domain-specific tasks, such as User Interface (UI) un…
MORALISE: A Structured Benchmark for Moral Alignment in Visual Language Models
Xiao Lin, Zhining Liu, Ze Yang +10
Warning: This paper contains examples of harmful language and images. Reader discretion is advised. Recently, vision-language models have demonstrated increasing influence in moral…