93 citations · 136 across the 16 of their papers we have counts for
6 papers · 1 filter
Bi-discriminator Domain Adversarial Neural Networks with Class-Level Gradient Alignment
Chuang Zhao, Hongke Zhao, Hengshu Zhu +4
Unsupervised domain adaptation aims to transfer rich knowledge from the annotated source domain to the unlabeled target domain with the same label space. One prevalent solution is…
Woodpecker: Hallucination Correction for Multimodal Large Language Models
Shukang Yin, Chaoyou Fu, Sirui Zhao +7
Hallucination is a big shadow hanging over the rapidly evolving Multimodal Large Language Models (MLLMs), referring to the phenomenon that the generated text is inconsistent with t…
A Solution to CVPR'2023 AQTC Challenge: Video Alignment for Multi-Step Inference
Chao Zhang, Shiwei Wu, Sirui Zhao +2
Affordance-centric Question-driven Task Completion (AQTC) for Egocentric Assistant introduces a groundbreaking scenario. In this scenario, through learning instructional videos, AI…
Recognizing Unseen Objects via Multimodal Intensive Knowledge Graph Propagation
Likang Wu, Zhi Li, Hongke Zhao +5
Zero-Shot Learning (ZSL), which aims at automatically recognizing unseen objects, is a promising learning paradigm to understand new real-world knowledge for machines continuously.…
A Survey on Multimodal Large Language Models
Shukang Yin, Chaoyou Fu, Sirui Zhao +4
Recently, Multimodal Large Language Model (MLLM) represented by GPT-4V has been a new rising research hotspot, which uses powerful Large Language Models (LLMs) as a brain to perfor…
AU-aware graph convolutional network for Macro- and Micro-expression spotting
Shukang Yin, Shiwei Wu, Tong Xu +3
Automatic Micro-Expression (ME) spotting in long videos is a crucial step in ME analysis but also a challenging task due to the short duration and low intensity of MEs. When solvin…