From the 1 of 10 linked papers with an AI index.
10 papers
MSEditor: Toward Consistent Multi-Shot Video Editing
Kunyu Feng, Yue Ma, Bingyuan Wang +6
In this paper, we tackle the problem of performing consistent, unified modifications to a multi-shot video sequence. This task is particularly challenging because multi-shot videos…
EmoWorld: A Decoupled Affective Field for Controllable Emotional Video Generation
Bingyuan Wang, Baistan Zhyldyzbekov, Kunyu Feng +1
Emotion shapes how viewers interpret a scene, yet existing video generators entangle global atmosphere, affect-bearing semantic cues, and temporal progression within a single text…
Teffic-Audio: Tell Fact from Fiction
Wan Lin, Li Wang, Jindong Wang +2
The paper presents Teffic-Audio, a speech deepfake detection system that uses a Conformer-based encoder with attentive pooling and a training recipe focused on multi-source data an…
VoxSafeBench: Not Just What Is Said, but Who, How, and Where
Yuxiang Wang, Hongyu Liu, Yijiang Xu +9
As speech language models (SLMs) transition from personal devices into shared, multi-user environments, their responses must account for far more than the words alone. Who is speak…
Follow-Your-Motion: Video Motion Transfer via Efficient Spatial-Temporal Decoupled Finetuning
Yue Ma, Yulong Liu, Qiyuan Zhu +8
Recently, breakthroughs in the video diffusion transformer have shown remarkable capabilities in diverse motion generations. As for the motion-transfer task, current methods mainly…
InstanceAnimator: Multi-Instance Sketch Video Colorization
Yinhan Zhang, Yue Ma, Bingyuan Wang +5
We propose InstanceAnimator, a novel Diffusion Transformer framework for multi-instance sketch video colorization. Existing methods suffer from three core limitations: inflexible u…