16 papers
Disentangling 3D Modeling from Spatial Reasoning
Haoze Sun, Jiequan Cui, Qingshan Xu +1
In this work, we explore an alternative paradigm for spatial reasoning by explicitly disentangling 3D perception from reasoning, rather than jointly acquiring implicit 3D perceptio…
Class-frequency Guided Noise Schedule for Diffusion Models
Jiequan Cui, Beier Zhu, Qingshan Xu +3
In this paper, we are the first to examine the correlations between class frequency and the multi-scale noise schedule within diffusion models. For score-based generative models, l…
iFLYTEK-Embodied-Omni Technical Report
Yuan Zhang, Jingfei Ni, Guanchen Lu +12
General-purpose embodied agents must understand multimodal instructions, anticipate how their environment will evolve, and produce precise control actions over extended horizons. E…
From Uncertainty to Stability and Fidelity: Guiding Sparse-View 3D Gaussian Splatting with Fisher Information
Junbao Zhou, Qingshan Xu, Yuan Zhou +7
3D Gaussian Splatting (3DGS) has emerged as a promising technique for novel view synthesis. However, 3DGS requires dense input views to achieve high-quality rendering. In sparse-vi…
Generalized Kullback-Leibler Divergence Loss
Jiequan Cui, Beier Zhu, Qingshan Xu +5
In this paper, we delve deeper into the Kullback-Leibler (KL) Divergence loss and mathematically prove that it is equivalent to the Decoupled Kullback-Leibler (DKL) Divergence loss…
Streaming Drag-Oriented Interactive Video Manipulation: Drag Anything, Anytime!
Junbao Zhou, Yuan Zhou, Kesen Zhao +4
Achieving streaming, fine-grained control over the outputs of autoregressive video diffusion models remains challenging, making it difficult to ensure that they consistently align…