1 citations · 1 across the 6 of their papers we have counts for
6 papers
SlowFast-VGen: Slow-Fast Learning for Action-Driven Long Video Generation
Yining Hong, Beide Liu, Maxine Wu +9
Human beings are endowed with a complementary learning system, which bridges the slow learning of general world dynamics with fast storage of episodic memory from a new experience.…
IDOL: Unified Dual-Modal Latent Diffusion for Human-Centric Joint Video-Depth Generation
Yuanhao Zhai, Kevin Lin, Linjie Li +7
Significant advances have been made in human-centric video generation, yet the joint video-depth generation problem remains underexplored. Most existing monocular depth estimation…
STAT: Towards Generalizable Temporal Action Localization
Yangcen Liu, Ziyi Liu, Yuanhao Zhai +3
Weakly-supervised temporal action localization (WTAL) aims to recognize and localize action instances with only video-level labels. Despite the significant progress, existing metho…
SOAR: Scene-debiasing Open-set Action Recognition
Yuanhao Zhai, Ziyi Liu, Zhenyu Wu +5
Deep learning models have a risk of utilizing spurious clues to make predictions, such as recognizing actions based on the background scene. This issue can severely degrade the ope…
Towards Generic Image Manipulation Detection with Weakly-Supervised Self-Consistency Learning
Yuanhao Zhai, Tianyu Luan, David Doermann +1
As advanced image manipulation techniques emerge, detecting the manipulation becomes increasingly important. Despite the success of recent learning-based approaches for image manip…
Language-guided Human Motion Synthesis with Atomic Actions
Yuanhao Zhai, Mingzhen Huang, Tianyu Luan +5
Language-guided human motion synthesis has been a challenging task due to the inherent complexity and diversity of human behaviors. Previous methods face limitations in generalizat…