most citedPatch-level Sounding Object Tracking for Audio-Visual Question Answering

2 citations · 4 across the 11 of their papers we have counts for

collaborators
Showing cs.CVShow all

8 papers · 1 filter

cs.CV2024

Linguistics-Vision Monotonic Consistent Network for Sign Language Production

Xu Wang, Shengeng Tang, Peipei Song +3

Sign Language Production (SLP) aims to generate sign videos corresponding to spoken language sentences, where the conversion of sign Glosses to Poses (G2P) is the key step. Due to…

cs.CV2024

Sign-IDD: Iconicity Disentangled Diffusion for Sign Language Production

Shengeng Tang, Jiayi He, Dan Guo +3

Sign Language Production (SLP) aims to generate semantically consistent sign videos from textual statements, where the conversion from textual glosses to sign poses (G2P) is a cruc…

cs.CV2024

Dense Audio-Visual Event Localization under Cross-Modal Consistency and Multi-Temporal Granularity Collaboration

Ziheng Zhou, Jinxing Zhou, Wei Qian +3

In the field of audio-visual learning, most research tasks focus exclusively on short videos. This paper focuses on the more practical Dense Audio-Visual Event Localization (DAVEL)…

cs.CV2024

Towards Pixel-Level Prediction for Gaze Following: Benchmark and Approach

Feiyang Liu, Dan Guo, Jingyuan Xu +4

Following the gaze of other people and analyzing the target they are looking at can help us understand what they are thinking, and doing, and predict the actions that may follow. E…

cs.CV2024

Modality Alignment Meets Federated Broadcasting

Yuting Ma, Shengeng Tang, Xiaohua Xu +1

Federated learning (FL) has emerged as a powerful approach to safeguard data privacy by training models across distributed edge devices without centralizing local data. Despite adv…

cs.CV2024

Shaping a Stabilized Video by Mitigating Unintended Changes for Concept-Augmented Video Editing

Mingce Guo, Jingxuan He, Shengeng Tang +2

Text-driven video editing utilizing generative diffusion models has garnered significant attention due to their potential applications. However, existing approaches are constrained…