most citedMulti-Source Video Domain Adaptation with Temporal Attentive Moment Alignment

5 citations · 7 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CV20231 cited

Link-Context Learning for Multimodal LLMs

Yan Tai, Weichen Fan, Zhao Zhang +3

The ability to learn from context with novel concepts, and deliver appropriate responses are essential in human conversations. Despite current Multimodal Large Language Models (MLL…

stat.ML2023

MM-DAG: Multi-task DAG Learning for Multi-modal Data -- with Application for Traffic Congestion Analysis

Tian Lan, Ziyue Li, Zhishuai Li +6

This paper proposes to learn Multi-task, Multi-modal Direct Acyclic Graphs (MM-DAGs), which are commonly observed in complex systems, e.g., traffic, manufacturing, and weather syst…

cs.CV20231 cited

Balancing Logit Variation for Long-tailed Semantic Segmentation

Yuchao Wang, Jingjing Fei, Haochen Wang +5

Semantic segmentation usually suffers from a long-tail data distribution. Due to the imbalanced number of samples across categories, the features of those tail classes may get sque…

cs.CV2023

Advancing Referring Expression Segmentation Beyond Single Image

Yixuan Wu, Zhao Zhang, Xie Chi +2

Referring Expression Segmentation (RES) is a widely explored multi-modal task, which endeavors to segment the pre-existing object within a single image with a given linguistic expr…

cs.CV20215 cited

Multi-Source Video Domain Adaptation with Temporal Attentive Moment Alignment

Yuecong Xu, Jianfei Yang, Haozhi Cao +4

Multi-Source Domain Adaptation (MSDA) is a more practical domain adaptation scenario in real-world scenarios. It relaxes the assumption in conventional Unsupervised Domain Adaptati…