57 papers
Branch2Skill: Efficient Skill Evolution Through Reasoning Trees
Yanwei Ren, Haotian Zhang, Likang Xiao +5
Skill evolution improves agent skills through feedback over time, with failed trajectories often providing informative signals by revealing incomplete or misleading behaviors. Howe…
Recommendation as Generation: Unifying Personalized Video Generation and Recommendation at Industrial Scale
Yanhua Cheng, Bo Wang, Haotian Zhang +17
Traditional short-video recommendation systems match user interest to a fixed pool of pre-produced videos, which limits their ability to capture fine-grained and dynamic preference…
MambaRaw: Selective State Space Modeling for Efficient 4K Raw Image Reconstruction
Peize Li, Fanhu Zeng, Tongda Xu +5
In-camera JPEG previews are ubiquitous in raw image formats and provide an sRGB reference at negligible storage cost. Although existing metadata-based reconstruction frameworks can…
Cosmos 3: Omnimodal World Models for Physical AI
NVIDIA, :, Aditi +293
We introduce Cosmos 3, a family of omnimodal world models designed to jointly process and generate language, image, video, audio, and action sequences within a unified mixture-of-t…
HIL: Hybrid Imitation Learning of Diverse Parkour Skills from Videos
Jiashun Wang, Yifeng Jiang, Haotian Zhang +4
Data-driven methods leveraging deep reinforcement learning have become the dominant paradigm for developing controllers that enable physically simulated characters to produce natur…
MMAE: A Massive Multitask Audio Editing Benchmark
Ziyang Ma, Ruiqi Yan, Ruiyang Xu +35
We introduce MMAE, a Massive Multitask Audio Editing benchmark, serving as the first comprehensive evaluation testbed designed for general-purpose instruction-based audio editing.…