2 papers
cs.CV2024
Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning
Jingru Yang, Huan Yu, Yang Jingxin +4
Multimodal Large Language Models (MLLMs) excel at descriptive tasks within images but often struggle with precise object localization, a critical element for reliable visual interp…
cs.CV2024
Morpho-Aware Global Attention for Image Matting
Jingru Yang, Chengzhi Cao, Chentianye Xu +4
Vision Transformers (ViTs) and Convolutional Neural Networks (CNNs) face inherent challenges in image matting, particularly in preserving fine structural details. ViTs, with their…