From the 1 of 5 linked papers with an AI index.
5 papers
DistillAlign: Coordinating Mode Covering and Mode Seeking in Autoregressive Video Distillation
Jiaxing Li, Kai Zou, Cindy Zhou +7
The paper studies autoregressive video distillation, showing that aligning the student model’s mode coverage with the teacher’s distribution improves generation quality and diversi…
OctoPipe: Reducing Pipeline Bubbles for Heterogeneous Models via Co-Optimizing Partitioning, Placement, and Scheduling
Jihu Guo, Tenghui Ma, Wei Gao +6
Pipeline parallelism is widely used to train large language models (LLMs). However, increasing heterogeneity in model architectures exacerbates pipeline bubbles, thereby reducing t…
How Mobile World Model Guides GUI Agents?
Weikai Xu, Kun Huang, Yunren Feng +10
Recent advances in vision-language models have enabled mobile GUI agents to perceive visual interfaces and execute user instructions, but reliable prediction of action consequences…
Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale
Liujie Zhang, Benzhe Ning, Rui Yang +8
Reinforcement learning (RL) post-training has proven effective at unlocking reasoning, self-reflection, and tool-use capabilities in large language models. As models extend to omni…
CharacterShot: Controllable and Consistent 4D Character Animation
Junyao Gao, Jiaxing Li, Wenran Liu +5
In this paper, we propose \textbf{CharacterShot}, a controllable and consistent 4D character animation framework that enables any individual designer to create dynamic 3D character…