36 papers · 1 filter
MOJITO: Modal Joint Learning for Unified End-to-End Autonomous Driving
Zhijing Cheng, Xuancheng Zhang, Donglin Di +4
End-to-end autonomous driving systems commonly follow a cascaded two-stage pipeline where a perception stage compresses multi-modal sensor inputs into a compact context and a downs…
TextGaze: Prompting Gaze Target Estimation with Textual Scene Cues
Junhui She, Fei Wang, Kun Li +4
Gaze target estimation aims to infer the position of a person's gaze within a scene. Within mainstream design logic, multi-branch methods require extra supervision and annotations,…
SafeGuard: A Multi-Agent Perception-Reasoning Framework for Social-Risk AI-Generated Video Detection
Wenlin Wu, Sheng Zhou, Peipei Song +3
As video generation paradigms evolve from localized manipulation to full-scene synthesis, AI-generated video detection becomes increasingly challenging, as forgeries exhibit cohere…
A New Multi-Domain Benchmark for Micro-Action Recognition and Detection
Yanbin Hao, Pengyu Liu, Xing Wei +3
Micro-actions are short-duration, low-amplitude subtle body movements at the whole-body level that can reveal latent intentions, involuntary reactions, and fine-grained affective c…
MuKV: Multi-Grained KV Cache Compression for Long Streaming Video Question-Answering
Junbin Xiao, Jiajun Chen, Tianxiang Sun +2
Long streaming video QA remains challenging due to growing visual tokens and limited reasoning length of large language models (LLMs). KV-caching stores the Key-Value (KV) of the h…
Memory-Augmented Query Intent Understanding for Efficient Chat-based Image Retrieval
Xianke Chen, Daizong Liu, Yushuo Lou +5
Different from traditional text-to-image retrieval tasks, chat-based image retrieval allows the human-interactive system to iteratively clarify and refine user intent through multi…