Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
Rethinking Point Clouds as Sequences: A Causal Next-Token Predictive Learning Framework
Yumeng Yao, Jingzhi Dong, Haowen Gu +4
With the rapid progress of multimodal foundation models and predictive pre-training, an important open question is how to equip 3D point clouds with a pre-training paradigm that is…
cs.CV2025
MotionGPT3: Human Motion as a Second Modality
Bingfan Zhu, Biao Jiang, Sunyi Wang +5
With the rapid progress of large language models (LLMs), multimodal frameworks that unify understanding and generation have become promising, yet they face increasing complexity as…