7 papers · 1 filter
TextGaze: Prompting Gaze Target Estimation with Textual Scene Cues
Junhui She, Fei Wang, Kun Li +4
Gaze target estimation aims to infer the position of a person's gaze within a scene. Within mainstream design logic, multi-branch methods require extra supervision and annotations,…
Evidence Packing for Cross-Domain Image Deepfake Detection with LVLMs
Yuxin Liu, Fei Wang, Kun Li +4
Image Deepfake Detection (IDD) separates manipulated images from authentic ones by spotting artifacts of synthesis or tampering. Although large vision-language models (LVLMs) offer…
Face-Guided Sentiment Boundary Enhancement for Weakly-Supervised Temporal Sentiment Localization
Cailing Han, Zhangbin Li, Jinxing Zhou +5
Point-level weakly-supervised temporal sentiment localization (P-WTSL) aims to detect sentiment-relevant segments in untrimmed multimodal videos using timestamp sentiment annotatio…
Training-Free Multimodal Deepfake Detection via Graph Reasoning
Yuxin Liu, Fei Wang, Kun Li +5
Multimodal deepfake detection (MDD) aims to uncover manipulations across visual, textual, and auditory modalities, thereby reinforcing the reliability of modern information systems…
CLASP: Cross-modal Salient Anchor-based Semantic Propagation for Weakly-supervised Dense Audio-Visual Event Localization
Jinxing Zhou, Ziheng Zhou, Yanghao Zhou +3
The Dense Audio-Visual Event Localization (DAVEL) task aims to temporally localize events in untrimmed videos that occur simultaneously in both the audio and visual modalities. Thi…
Exploiting Ensemble Learning for Cross-View Isolated Sign Language Recognition
Fei Wang, Kun Li, Yiqi Nie +5
In this paper, we present our solution to the Cross-View Isolated Sign Language Recognition (CV-ISLR) challenge held at WWW 2025. CV-ISLR addresses a critical issue in traditional…