8 papers
Milestone-Guided Policy Learning for Long-Horizon Language Agents
Zixuan Wang, Yuchen Yan, Hongxing Li +7
While long-horizon agentic tasks require language agents to perform dozens of sequential decisions, training such agents with reinforcement learning remains challenging. We identif…
RealCam: Real-Time Novel-View Video Generation with Interactive Camera Control
Youcan Xu, Jiaxin Shi, Zhen Wang +5
Camera-controlled video-to-video (V2V) generation enables dynamic viewpoint synthesis from monocular footage, holding immense potential for interactive filmmaking and live broadcas…
Physically Plausible Human-Object Rendering from Sparse Views via 3D Gaussian Splatting
Weiquan Wang, Jun Xiao, Yi Yang +2
Rendering realistic human-object interactions (HOIs) from sparse-view inputs is a challenging yet crucial task for various real-world applications. Existing methods often struggle…
Rendering Multi-Human and Multi-Object with 3D Gaussian Splatting
Weiquan Wang, Jun Xiao, Feifei Shao +3
Reconstructing dynamic scenes with multiple interacting humans and objects from sparse-view inputs is a critical yet challenging task, essential for creating high-fidelity digital…
Uncertainty-Aware 4D Gaussian Splatting for Monocular Occluded Human Rendering
Weiquan Wang, Feifei Shao, Lin Li +4
High-fidelity rendering of dynamic humans from monocular videos typically degrades catastrophically under occlusions. Existing solutions incorporate external priors-either hallucin…
em: Learning Hierarchical Hyperbolic Embeddings for Compositional Zero-Shot Learning
Lin Li, Jiahui Li, Jiaming Lei +3
Compositional zero-shot learning (CZSL) aims to recognize unseen state-object compositions by generalizing from a training set of their primitives (state and object). Current metho…