11 papers
DreamTraj: Generating 6-DoF Object Trajectories by Reading Unrendered Video Diffusion Latents
Tongsheng Ding, Zhen Luo, Yixuan Yang +4
Accurate prediction of object trajectories during manipulation is essential for closing the perception-action loop. Progress is limited on two fronts: available datasets lack fine-…
Trans2Occ: Voxel Occupancy Estimation and Grasp for Transparent Objects from Simulation to Reality
Yixuan Yang, Sha Zhang, Rui Li +12
Transparent objects remain challenging for robotic perception due to unreliable depth sensing caused by refraction and reflection. While prior approaches rely on multi-view reconst…
Convergence of Spectral Descent for Non-smooth Optimization
Yixuan Yang, Yuqing He, Song Li
The Muon optimizer has recently demonstrated remarkable empirical success in training large language models. However, the theoretical understanding of its mechanisms remains limite…
Toward Native Multimodal Modeling: A Roadmap
Siyu An, Junru Lu, Junnan Dong +18
Multimodal modeling represents a vital step from modality-agnostic reasoning toward world modeling. While early approaches predominantly rely on late-fusion that assembles encoders…
STABLE: Simulation-Ready Tabletop Layout Generation via a Semantics-Physics Dual System
Zhen Luo, Yixuan Yang, Xudong Xu +5
Generating simulation-ready tabletop scenes from task instructions is an intriguing and promising research direction in the field of Embodied AI. However, existing task-to-scene ge…
Code-as-Room: Generating 3D Rooms from Top-Down View Images via Agentic Code Synthesis
Yixuan Yang, Zhen Luo, Wanshui Gan +5
Designing realistic and functional 3D indoor rooms is essential for a wide range of applications, including interior design, virtual reality, gaming, and embodied AI. While recent…