3 papers
cs.CV2026
REF-VLM: Triplet-Based Referring Paradigm for Unified Visual Decoding
Yan Tai, Luhao Zhu, Yunan Ding +4
Multimodal Large Language Models (MLLMs) demonstrate robust zero-shot capabilities across diverse vision-language tasks after training on mega-scale datasets. However, dense predic…
cs.CV2025
Exploring Representation Invariance in Finetuning
Wenqiang Zu, Shenghao Xie, Hao Chen +9
Foundation models pretrained on large-scale natural images are widely adapted to various cross-domain low-resource downstream tasks, benefiting from generalizable and transferable…
cs.CV2024
Learning from Pattern Completion: Self-supervised Controllable Generation
Zhiqiang Chen, Guofan Fan, Jinying Gao +4
The human brain exhibits a strong ability to spontaneously associate different visual attributes of the same or similar visual scene, such as associating sketches and graffiti with…