5 papers · 1 filter
Query-Conditioned Evidential Keyframe Sampling for MLLM-Based Long-Form Video Understanding
Yiheng Wang, Lichen Zhu, Yueqian Lin +4
Multimodal Large Language Models (MLLMs) have shown strong performance on video question answering, but their application to long-form videos is constrained by limited context leng…
Dual Prompt-Driven Feature Encoding for Nighttime UAV Tracking
Yiheng Wang, Changhong Fu, Liangliang Yao +2
Robust feature encoding constitutes the foundation of UAV tracking by enabling the nuanced perception of target appearance and motion, thereby playing a pivotal role in ensuring re…
TeleEgo: Benchmarking Egocentric AI Assistants in the Wild
Jiaqi Yan, Ruilong Ren, Jingren Liu +13
Egocentric AI assistants in real-world settings must process multi-modal inputs (video, audio, text), respond in real time, and retain evolving long-term memory. However, existing…
Prompt-Driven Temporal Domain Adaptation for Nighttime UAV Tracking
Changhong Fu, Yiheng Wang, Liangliang Yao +3
Nighttime UAV tracking under low-illuminated scenarios has achieved great progress by domain adaptation (DA). However, previous DA training-based works are deficient in narrowing t…
Enhancing Nighttime UAV Tracking with Light Distribution Suppression
Liangliang Yao, Changhong Fu, Yiheng Wang +2
Visual object tracking has boosted extensive intelligent applications for unmanned aerial vehicles (UAVs). However, the state-of-the-art (SOTA) enhancers for nighttime UAV tracking…