5 papers
SPARGen: Unifying Spatial Perception and Reasoning through Native Multimodal Generation
Jinsheng Quan, Jianhua Li, Siyi Xie +7
Spatial perception and reasoning from visual observations require recovering geometric structure, establishing correspondences, and understanding spatial relations. Existing approa…
Vision as Unified Multimodal Generation
Xiaoyang Han, Jianhua Li, Kewang Deng +14
We formulate computer vision as unified multimodal generation, where heterogeneous visual tasks are expressed in the native text and image generation spaces of a unified multimodal…
Same Evidence, Different Answers: Canonical-Context On-Policy Distillation for Multi-Turn Language Models
Zizhuo Lin, Quanling Liu, Jinsheng Quan +6
Large language models (LLMs) often solve a task when all instructions are given in a single prompt, but fail when the same information is revealed gradually across turns. When a cl…
ParticleGS: Learning Neural Gaussian Particle Dynamics from Videos for Prior-free Physical Motion Extrapolation
Jinsheng Quan, Qiaowei Miao, Yichao Xu +5
The ability to extrapolate dynamic 3D scenes beyond the observed timeframe is fundamental to advancing physical world understanding and predictive modeling. Existing dynamic 3D rec…
Advances in 4D Generation: A Survey
Qiaowei Miao, Kehan Li, Jinsheng Quan +6
Generative artificial intelligence has recently progressed from static image and video synthesis to 3D content generation, culminating in the emergence of 4D generation-the task of…