4 papers
Generalizable Detection of AI Generated Images with Large Models and Fuzzy Decision Tree
Fei Wu, Guanghao Ding, Zijian Niu +4
The malicious use and widespread dissemination of AI-generated images pose a serious threat to the authenticity of digital content. Existing detection methods exploit low-level art…
Cascade-Free Mandarin Visual Speech Recognition via Semantic-Guided Cross-Representation Alignment
Lei Yang, Yi He, Fei Wu +1
Chinese mandarin visual speech recognition (VSR) is a task that has advanced in recent years, yet still lags behind the performance on non-tonal languages such as English. One prim…
Landmark Guided Visual Feature Extractor for Visual Speech Recognition with Limited Resource
Lei Yang, Junshan Jin, Mingyuan Zhang +3
Visual speech recognition is a technique to identify spoken content in silent speech videos, which has raised significant attention in recent years. Advancements in data-driven dee…
Enhancing Visual Forced Alignment with Local Context-Aware Feature Extraction and Multi-Task Learning
Yi He, Lei Yang, Shilin Wang
This paper introduces a novel approach to Visual Forced Alignment (VFA), aiming to accurately synchronize utterances with corresponding lip movements, without relying on audio cues…