activity
20242026
collaborators

11 papers

cs.CV2026

DreamTraj: Generating 6-DoF Object Trajectories by Reading Unrendered Video Diffusion Latents

Tongsheng Ding, Zhen Luo, Yixuan Yang +4

Accurate prediction of object trajectories during manipulation is essential for closing the perception-action loop. Progress is limited on two fronts: available datasets lack fine-…

cs.RO2026

Trans2Occ: Voxel Occupancy Estimation and Grasp for Transparent Objects from Simulation to Reality

Yixuan Yang, Sha Zhang, Rui Li +12

Transparent objects remain challenging for robotic perception due to unreliable depth sensing caused by refraction and reflection. While prior approaches rely on multi-view reconst…

cs.LG2026

Convergence of Spectral Descent for Non-smooth Optimization

Yixuan Yang, Yuqing He, Song Li

The Muon optimizer has recently demonstrated remarkable empirical success in training large language models. However, the theoretical understanding of its mechanisms remains limite…

cs.CV2026

Toward Native Multimodal Modeling: A Roadmap

Siyu An, Junru Lu, Junnan Dong +18

Multimodal modeling represents a vital step from modality-agnostic reasoning toward world modeling. While early approaches predominantly rely on late-fusion that assembles encoders…

cs.CV2026

STABLE: Simulation-Ready Tabletop Layout Generation via a Semantics-Physics Dual System

Zhen Luo, Yixuan Yang, Xudong Xu +5

Generating simulation-ready tabletop scenes from task instructions is an intriguing and promising research direction in the field of Embodied AI. However, existing task-to-scene ge…

cs.CV2026

Code-as-Room: Generating 3D Rooms from Top-Down View Images via Agentic Code Synthesis

Yixuan Yang, Zhen Luo, Wanshui Gan +5

Designing realistic and functional 3D indoor rooms is essential for a wide range of applications, including interior design, virtual reality, gaming, and embodied AI. While recent…