From the 1 of 9 linked papers with an AI index.
9 papers
EgoProceVQA: A Novel Egocentric Procedural Understanding Task with Self-Skill-Exploration Agent
Junlong Li, Junxi Li, Yuxiang Yang +3
The paper introduces EgoProceVQA, a new egocentric video question answering benchmark focused on procedural reasoning, and proposes EgoProceAgent, a self-exploring agent that learn…
NTIRE 2026 Challenge on Bitstream-Corrupted Video Restoration: Methods and Results
Wenbin Zou, Tianyi Liu, Kejun Wu +37
This paper reports on the NTIRE 2026 Challenge on Bitstream-Corrupted Video Restoration (BSCVR). The challenge aims to advance research on recovering visually coherent videos from…
Mixture-of-Modality-Experts with Holistic Token Learning for Fine-Grained Multimodal Visual Analytics in Driver Action Recognition
Tianyi Liu, Yiming Li, Wenqian Wang +4
Robust multimodal visual analytics remains challenging when heterogeneous modalities provide complementary but input-dependent evidence for decision-making.Existing multimodal lear…
Accelerating Diffusion-based Video Editing via Heterogeneous Caching: Beyond Full Computing at Sampled Denoising Timestep
Tianyi Liu, Ye Lu, Linfeng Zhang +5
Diffusion-based video editing has emerged as an important paradigm for high-quality and flexible content generation. However, despite their generality and strong modeling capacity,…
Building Egocentric Procedural AI Assistant: Methods, Benchmarks, and Challenges
Junlong Li, Huaiyuan Xu, Sijie Cheng +4
Driven by recent advances in vision-language models (VLMs) and egocentric perception research, the emerging topic of an egocentric procedural AI assistant (EgoProceAssist) is intro…
EDVD-LLaMA: Explainable Deepfake Video Detection via Multimodal Large Language Model Reasoning
Haoran Sun, Chen Cai, Huiping Zhuang +3
The rapid development of deepfake video technology has not only facilitated artistic creation but also made it easier to spread misinformation. Traditional deepfake video detection…