7 papers · 1 filter
X-Humanoid: Robotize Human Videos to Generate Humanoid Videos at Scale
Pei Yang, Hai Ci, Yiren Song +1
The advancement of embodied AI has unlocked significant potential for intelligent humanoid robots. However, progress in both Vision-Language-Action (VLA) models and world models is…
B2N3D: Progressive Learning from Binary to N-ary Relationships for 3D Object Grounding
Feng Xiao, Hongbin Xu, Hai Ci +1
Localizing 3D objects using natural language is essential for robotic scene understanding. The descriptions often involve multiple spatial relationships to distinguish similar obje…
DiffSeg30k: A Multi-Turn Diffusion Editing Benchmark for Localized AIGC Detection
Hai Ci, Ziheng Peng, Pei Yang +2
Diffusion-based editing enables realistic modification of local image regions, making AI-generated content harder to detect. Existing AIGC detection benchmarks focus on classifying…
Cyc3D: Fine-grained Controllable 3D Generation via Cycle Consistency Regularization
Hongbin Xu, Chaohui Yu, Feng Xiao +5
Despite the remarkable progress of 3D generation, achieving controllability, i.e., ensuring consistency between generated 3D content and input conditions like edge and depth, remai…
Impossible Videos
Zechen Bai, Hai Ci, Mike Zheng Shou
Synthetic videos nowadays is widely used to complement data scarcity and diversity of real-world videos. Current synthetic datasets primarily replicate real-world scenarios, leavin…
Anti-Reference: Universal and Immediate Defense Against Reference-Based Generation
Yiren Song, Shengtao Lou, Xiaokang Liu +4
Diffusion models have revolutionized generative modeling with their exceptional ability to produce high-fidelity images. However, misuse of such potent tools can lead to the creati…