6 papers
Athena-Brain Technical Report: An Efficient Robot Brain for General Intelligence and Embodied Interaction
Jialian Li, Junhong Liu, Yuchen Cao +6
Large language models (LLMs) have demonstrated remarkable capabilities in language understanding, reasoning, and world knowledge. As embodied agents become increasingly capable, th…
Engagement Process: Rethinking the Temporal Interface of Action and Observation
Jialian Li, Yuchen Cao, Junhong Liu +5
Task completion in digital and physical environments increasingly involves complex temporal interaction, where actions and observations unfold over different time scales rather tha…
Ideas in Inference-time Scaling can Benefit Generative Pre-training Algorithms
Jiaming Song, Linqi Zhou
Generative pre-training is often framed through a false dichotomy between autoregressive models for discrete signals and diffusion models for continuous signals. We argue that the…
Terminal Velocity Matching
Linqi Zhou, Mathias Parger, Ayaan Haque +1
We propose Terminal Velocity Matching (TVM), a generalization of flow matching that enables high-fidelity one- and few-step generative modeling. TVM models the transition between a…
Inductive Moment Matching
Linqi Zhou, Stefano Ermon, Jiaming Song
Diffusion models and Flow Matching generate high-quality samples but are slow at inference, and distilling them into few-step models often leads to instability and extensive tuning…
Personalized Preference Fine-tuning of Diffusion Models
Meihua Dang, Anikait Singh, Linqi Zhou +2
RLHF techniques like DPO can significantly improve the generation quality of text-to-image diffusion models. However, these methods optimize for a single reward that aligns model g…