2 papers
cs.LG2026
Heterogeneous Parallelism for Multimodal Large Language Model Training
Yashaswi Karnati, Kamran Jafari, Akash Mehra +10
Foundation model training is becoming multimodal, from post-training pipelines to large-scale pretraining. As modality coverage broadens, context windows grow, and encoder LLM scal…
cs.CV2026
Rethinking Point Clouds as Sequences: A Causal Next-Token Predictive Learning Framework
Yumeng Yao, Jingzhi Dong, Haowen Gu +4
With the rapid progress of multimodal foundation models and predictive pre-training, an important open question is how to equip 3D point clouds with a pre-training paradigm that is…