collaborators

18 papers

cs.AI2026

ToolLIFT: Lifting Tool-Specific Trajectories into Function-Level Graphs for Generalizable Tool Planning

Xiuhui You, Jiayi Luo, Zichao Shen +2

Historical tool-use trajectories provide valuable experience for large language model (LLM) agents to plan and coordinate tool usage. Existing approaches directly construct tool-le…

cs.RO2026

EmbodiedVAE: Disentangled Video VAE for Efficient and Controllable Embodied Manipulation

Jiayi Luo, Hanxin Zhu, Chen Gao +5

Latent diffusion models (LDMs) have recently significantly advanced embodied learning in constructing powerful embodied manipulation world models. However, despite the remarkable p…

cs.CV2026

Token Radius Attention for Efficient Video Generation

Jiayu Chen, Zhikun Jiang, Maoliang Li +6

Video Diffusion Transformers (VDiTs) enable high-fidelity generation but incur quadratic cost from dense 3D self-attention. Existing head- and block-level sparse methods share comp…

cs.CV2026

Embody4D: A Generalist Data Engine for Embodied 4D World Modeling

Peiyan Tu, Hanxin Zhu, Jingwen Sun +6

Embodied agents require robust and comprehensive 3D spatiotemporal representations to support spatial reasoning, manipulation understanding, and downstream decision making. However…

cs.CV2026

Physics-Informed Video Generation via Mixture-of-Experts Latent Alignment

Cong Wang, Hanxin Zhu, Jiayi Luo +6

Large-scale video generation models have made remarkable progress in semantic consistency and visual quality, producing videos that are increasingly coherent and visually convincin…

cs.CV2026

OrthoPhys: Physically Plausible Video Generation with Orthogonal-View Geometry Guidance

Cong Wang, Hanxin Zhu, Xiao Tang +4

Recent progress in video generation has led to substantial improvements in visual fidelity, yet ensuring physically consistent motion remains a fundamental challenge. Intuitively,…