19 citations · 37 across the 6 of their papers we have counts for
15 papers · 1 filter
World Simulation with Video Foundation Models for Physical AI
NVIDIA, :, Arslan Ali +87
We introduce [Cosmos-Predict2.5], the latest generation of the Cosmos World Foundation Models for Physical AI. Built on a flow-based architecture, [Cosmos-Predict2.5] unifies Text2…
UFO: A Unified Framework towards Omni-supervised Object Detection
Zhongzheng Ren, Zhiding Yu, Xiaodong Yang +3
Existing work on object detection often relies on a single form of annotation: the model is trained using either accurate yet costly bounding boxes or cheaper but less expressive i…
Joint Disentangling and Adaptation for Cross-Domain Person Re-Identification
Yang Zou, Xiaodong Yang, Zhiding Yu +2
Although a significant progress has been witnessed in supervised person re-identification (re-id), it remains challenging to generalize re-id models to new domains due to the huge…
Instance-aware, Context-focused, and Memory-efficient Weakly Supervised Object Detection
Zhongzheng Ren, Zhiding Yu, Xiaodong Yang +4
Weakly supervised learning has emerged as a compelling tool for object detection by reducing the need for strong supervision during training. However, major challenges remain: (1)…
Dancing to Music
Hsin-Ying Lee, Xiaodong Yang, Ming-Yu Liu +4
Dancing to music is an instinctive move by humans. Learning to model the music-to-dance generation process is, however, a challenging problem. It requires significant efforts to me…
Few-shot Video-to-Video Synthesis
Ting-Chun Wang, Ming-Yu Liu, Andrew Tao +3
Video-to-video synthesis (vid2vid) aims at converting an input semantic video, such as videos of human poses or segmentation masks, to an output photorealistic video. While the sta…