6 citations · 12 across the 13 of their papers we have counts for
20 papers
Event-Aware Instructed Assistant for Referring Video Segmentation
Jinyu Liu, Henghui Ding, Shuting He +1
Existing referring video segmentation methods often treat a video as a single event consisting of multiple images, overlooking the fact that a video typically contains multiple dis…
Embodied AI: From LLMs to World Models
Tongtong Feng, Xin Wang, Yu-Gang Jiang +1
Embodied Artificial Intelligence (AI) is an intelligent system paradigm for achieving Artificial General Intelligence (AGI), serving as the cornerstone for various applications and…
FedAPT: Federated Adversarial Prompt Tuning for Vision-Language Models
Kun Zhai, Siheng Chen, Xingjun Ma +1
Federated Prompt Tuning (FPT) is an efficient method for cross-client collaborative fine-tuning of large Vision-Language Models (VLMs). However, models tuned using FPT are vulnerab…
BrokenVideos: A Benchmark Dataset for Fine-Grained Artifact Localization in AI-Generated Videos
Jiahao Lin, Weixuan Peng, Bojia Zi +4
Recent advances in deep generative models have led to significant progress in video generation, yet the fidelity of AI-generated videos remains limited. Synthesized content often e…
Domain Expansion and Boundary Growth for Open-Set Single-Source Domain Generalization
Pengkun Jiao, Na Zhao, Jingjing Chen +1
Open-set single-source domain generalization aims to use a single-source domain to learn a robust model that can be generalized to unknown target domains with both domain shifts an…
Navigating Weight Prediction with Diet Diary
Yinxuan Gui, Bin Zhu, Jingjing Chen +2
Current research in food analysis primarily concentrates on tasks such as food recognition, recipe retrieval and nutrition estimation from a single image. Nevertheless, there is a…