5 citations · 7 across the 5 of their papers we have counts for
5 papers
Link-Context Learning for Multimodal LLMs
Yan Tai, Weichen Fan, Zhao Zhang +3
The ability to learn from context with novel concepts, and deliver appropriate responses are essential in human conversations. Despite current Multimodal Large Language Models (MLL…
MM-DAG: Multi-task DAG Learning for Multi-modal Data -- with Application for Traffic Congestion Analysis
Tian Lan, Ziyue Li, Zhishuai Li +6
This paper proposes to learn Multi-task, Multi-modal Direct Acyclic Graphs (MM-DAGs), which are commonly observed in complex systems, e.g., traffic, manufacturing, and weather syst…
Balancing Logit Variation for Long-tailed Semantic Segmentation
Yuchao Wang, Jingjing Fei, Haochen Wang +5
Semantic segmentation usually suffers from a long-tail data distribution. Due to the imbalanced number of samples across categories, the features of those tail classes may get sque…
Advancing Referring Expression Segmentation Beyond Single Image
Yixuan Wu, Zhao Zhang, Xie Chi +2
Referring Expression Segmentation (RES) is a widely explored multi-modal task, which endeavors to segment the pre-existing object within a single image with a given linguistic expr…
Multi-Source Video Domain Adaptation with Temporal Attentive Moment Alignment
Yuecong Xu, Jianfei Yang, Haozhi Cao +4
Multi-Source Domain Adaptation (MSDA) is a more practical domain adaptation scenario in real-world scenarios. It relaxes the assumption in conventional Unsupervised Domain Adaptati…