activity
20242026
collaborators

15 papers

cs.CV2026

Event-Aware Instructed Assistant for Referring Video Segmentation

Jinyu Liu, Henghui Ding, Shuting He +1

Existing referring video segmentation methods often treat a video as a single event consisting of multiple images, overlooking the fact that a video typically contains multiple dis…

cs.AI2025

Embodied AI: From LLMs to World Models

Tongtong Feng, Xin Wang, Yu-Gang Jiang +1

Embodied Artificial Intelligence (AI) is an intelligent system paradigm for achieving Artificial General Intelligence (AGI), serving as the cornerstone for various applications and…

cs.CV2025

FedAPT: Federated Adversarial Prompt Tuning for Vision-Language Models

Kun Zhai, Siheng Chen, Xingjun Ma +1

Federated Prompt Tuning (FPT) is an efficient method for cross-client collaborative fine-tuning of large Vision-Language Models (VLMs). However, models tuned using FPT are vulnerab…

cs.CV2025

BrokenVideos: A Benchmark Dataset for Fine-Grained Artifact Localization in AI-Generated Videos

Jiahao Lin, Weixuan Peng, Bojia Zi +4

Recent advances in deep generative models have led to significant progress in video generation, yet the fidelity of AI-generated videos remains limited. Synthesized content often e…

cs.CV2025

DiffusionAD: Norm-guided One-step Denoising Diffusion for Anomaly Detection

Hui Zhang, Zheng Wang, Dan Zeng +2

Anomaly detection has garnered extensive applications in real industrial manufacturing due to its remarkable effectiveness and efficiency. However, previous generative-based models…

cs.CV2025

Instruction-Guided Scene Text Recognition

Yongkun Du, Zhineng Chen, Yuchen Su +2

Multi-modal models have shown appealing performance in visual recognition tasks, as free-form text-guided training evokes the ability to understand fine-grained visual content. How…