papers

Publications (30)

cs.AI2024

Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Chunting Zhou, Lili Yu, Arun Babu +7

We introduce Transfusion, a recipe for training a multi-modal model over discrete and continuous data. Transfusion combines the language modeling loss function (next token predicti…

cs.LG2026

: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities

Physical Intelligence, Bo Ai, Ali Amin +85

We present a new robotic foundation model, called , that can enable strong out-of-the-box performance in a wide range of scenarios. can follow diverse language…

cs.LG2024

Megalodon: Efficient LLM Pretraining and Inference with Unlimited Context Length

Xuezhe Ma, Xiaomeng Yang, Wenhan Xiong +7

The quadratic complexity and weak length extrapolation of Transformers limits their ability to scale to long sequences, and while sub-quadratic solutions like linear attention and…

cond-mat.mtrl-sci2010

Functionalized Graphene for High Performance Two-dimensional Spintronics Devices

Linze Li, Rui Qin, Hong Li +5

Using first-principles calculations, we explore the possibility of functionalized graphene as high performance two-dimensional spintronics device. Graphene functionalized with O on…

cs.CL2025

Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models

Weixin Liang, Lili Yu, Liang Luo +8

The development of large language models (LLMs) has expanded to multi-modal systems capable of processing text, images, and speech within a unified framework. Training these models…

cs.LG2023

Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Lili Yu, Bowen Shi, Ramakanth Pasunuru +24

We present CM3Leon (pronounced "Chameleon"), a retrieval-augmented, token-based, decoder-only multi-modal language model capable of generating and infilling both text and images. C…

cs.LG2023

MEGABYTE: Predicting Million-byte Sequences with Multiscale Transformers

Lili Yu, Dániel Simig, Colin Flaherty +3

Autoregressive transformers are spectacular models for short sequences but scale poorly to long sequences such as high-resolution images, podcasts, code, or books. We proposed Mega…

cs.CL2021

Nutribullets Hybrid: Multi-document Health Summarization

Darsh J Shah, Lili Yu, Tao Lei +1

We present a method for generating comparative summaries that highlights similarities and contradictions in input documents. The key challenge in creating such summaries is the lac…

cs.LG2023

Jointly Training Large Autoregressive Multimodal Models

Emanuele Aiello, Lili Yu, Yixin Nie +2

In recent years, advances in the large-scale pretraining of language and text-to-image models have revolutionized the field of machine learning. Yet, integrating these two modaliti…

cs.LG2025

Mixture-of-Mamba: Enhancing Multi-Modal State-Space Models with Modality-Aware Sparsity

Weixin Liang, Junhong Shen, Genghan Zhang +3

State Space Models (SSMs) have emerged as efficient alternatives to Transformers for sequential modeling, but their inability to leverage modality-specific features limits their pe…

cs.CV2023

VideoOFA: Two-Stage Pre-Training for Video-to-Text Generation

Xilun Chen, Lili Yu, Wenhan Xiong +3

We propose a new two-stage pre-training framework for video-to-text generation tasks such as video captioning and video question answering: A generative encoder-decoder model is fi…

cs.CL2021

Nutri-bullets: Summarizing Health Studies by Composing Segments

Darsh J Shah, Lili Yu, Tao Lei +1

We introduce \emph{Nutri-bullets}, a multi-document summarization task for health and nutrition. First, we present two datasets of food and health summaries from multiple scientifi…

cs.CL2025

LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Weijia Shi, Xiaochuang Han, Chunting Zhou +4

We present LMFusion, a framework for empowering pretrained text-only large language models (LLMs) with multimodal generative capabilities, enabling them to understand and generate…

physics.soc-ph2020

On the evolution of word usage of classical Chinese poetry

Liang Liu, Lili Yu

The hierarchy of classical Chinese poetry has been broadly acknowledged by a number of studies in Chinese literature. However, quantitative investigations about the evolutionary li…

cs.LG2025

Knowledge Insulating Vision-Language-Action Models: Train Fast, Run Fast, Generalize Better

Danny Driess, Jost Tobias Springenberg, Brian Ichter +8

Vision-language-action (VLA) models provide a powerful approach to training control policies for physical systems, such as robots, by combining end-to-end learning with transfer of…

cond-mat.mes-hall2012

Integrated Circuits Based on Bilayer MoS2 Transistors

Han Wang, Lili Yu, Yi-Hsien Lee +7

Two-dimensional (2D) materials, such as molybdenum disulfide (MoS2), have been shown to exhibit excellent electrical and optical properties. The semiconducting nature of MoS2 allow…

cs.LG2025

: a Vision-Language-Action Model with Open-World Generalization

Physical Intelligence, Kevin Black, Noah Brown +33

In order for robots to be useful, they must perform practically relevant tasks in the real world, outside of the lab. While vision-language-action (VLA) models have demonstrated im…

cs.CV2025

CAT: Content-Adaptive Image Tokenization

Junhong Shen, Kushal Tirumala, Michihiro Yasunaga +4

Most existing image tokenizers encode images into a fixed number of tokens or patches, overlooking the inherent variability in image complexity. To address this, we introduce Conte…

cs.LG2020

Rationalizing Text Matching: Learning Sparse Alignments via Optimal Transport

Kyle Swanson, Lili Yu, Tao Lei

Selecting input features of top relevance has become a popular method for building self-explaining models. In this work, we extend this selective rationalization approach to text m…

cs.LG2025

: a VLA That Learns From Experience

Physical Intelligence, Ali Amin, Raichelle Aniceto +53

We study how vision-language-action (VLA) models can improve through real-world deployments via reinforcement learning (RL). We present a general-purpose method, RL with Experience…

cs.CL2023

LIMA: Less Is More for Alignment

Chunting Zhou, Pengfei Liu, Puxin Xu +12

Large language models are trained in two stages: (1) unsupervised pretraining from raw text, to learn general-purpose representations, and (2) large scale instruction tuning and re…

cs.CL2019

Building a Production Model for Retrieval-Based Chatbots

Kyle Swanson, Lili Yu, Christopher Fox +2

Response suggestion is an important task for building human-computer conversation systems. Recent approaches to conversation modeling have introduced new model architectures with i…

cond-mat.mtrl-sci2013

Large-scale 2D Electronics based on Single-layer MoS2 Grown by Chemical Vapor Deposition

Han Wang, Lili Yu, Yi-Hsien Lee +8

2D nanoelectronics based on single-layer MoS2 offers great advantages for both conventional and ubiquitous applications. This paper discusses the large-scale CVD growth of single-l…

q-bio.QM2024

Bayesian estimation of transmission networks for infectious diseases

Jianing Xu, Huimin Hu, Gregory Ellison +3

Reconstructing transmission networks is essential for identifying key factors like superspreaders and high-risk locations, which are critical for developing effective pandemic prev…

cs.CL2024

Improving Faithfulness of Abstractive Summarization by Controlling Confounding Effect of Irrelevant Sentences

Asish Ghoshal, Arash Einolghozati, Ankit Arun +6

Lack of factual correctness is an issue that still plagues state-of-the-art summarization systems despite their impressive progress on generating seemingly fluent summaries. In thi…

cond-mat.mtrl-sci2015

Parallel Stitching of Two-Dimensional Materials

Xi Ling, Yuxuan Lin, Qiong Ma +16

Diverse parallel stitched two-dimensional heterostructures are synthesized, including metal-semiconductor (graphene-MoS2), semiconductor-semiconductor (WS2-MoS2), and insulator-sem…

cs.CL2023

Scaling Laws for Generative Mixed-Modal Language Models

Armen Aghajanyan, Lili Yu, Alexis Conneau +7

Generative language models define distributions over sequences of tokens that can represent essentially any combination of data modalities (e.g., any permutation of image tokens fr…

cs.CL2020

Interactive Classification by Asking Informative Questions

Lili Yu, Howard Chen, Sida Wang +2

We study the potential for interaction in natural language classification. We add a limited form of interaction for intent classification, where users provide an initial query usin…

cs.CV2025

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation

Zhiyang Xu, Jiuhai Chen, Zhaojiang Lin +10

Recent advances in large language models (LLMs) have enabled multimodal foundation models to tackle both image understanding and generation within a unified framework. Despite thes…

cs.CL2024

Byte Latent Transformer: Patches Scale Better Than Tokens

Artidoro Pagnoni, Ram Pasunuru, Pedro Rodriguez +11

We introduce the Byte Latent Transformer (BLT), a new byte-level LLM architecture that, for the first time, matches tokenization-based LLM performance at scale with significant imp…