papers

Publications (29)

cs.CV2024

MaskBit: Embedding-free Image Generation via Bit Tokens

Mark Weber, Lijun Yu, Qihang Yu +4

Masked transformer models for class-conditional image generation have become a compelling alternative to diffusion models. Typically comprising two stages - an initial VQGAN model…

physics.soc-ph2022

Comparing value of travel time and value of travel time saving with heterogeneity in travelers

Lijun Yu, Baojun He

In research on the value of past time, the value of travel time and the value of saving travel time are two different concepts that have been vaguely distinguished over an extended…

physics.optics2024

Polarization Purity and Dispersion Characteristics of Nested Antiresonant Nodeless Hollow-Core Optical Fiber at Near- and Short-wave-IR Wavelengths for Quantum Communications

Ivi Afxenti, Lijun Yu, Taylor Shields +5

Advancements in quantum communication and sensing require improved optical transmission that ensures excellent state purity and reduced losses. While free-space optical communicati…

eess.AS2025

YuE: Scaling Open Foundation Models for Long-Form Music Generation

Ruibin Yuan, Hanfeng Lin, Shuyue Guo +55

We tackle the task of long-form music generation--particularly the challenging \textbf{lyrics-to-song} problem--by introducing YuE, a family of open foundation models based on the…

cs.CV2024

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Lijun Yu, José Lezama, Nitesh B. Gundavarapu +13

While Large Language Models (LLMs) are the dominant models for generative tasks in language, they do not perform as well as diffusion models on image and video generation. To effec…

cs.CV2024

Towards Multi-Task Multi-Modal Models: A Video Generative Perspective

Lijun Yu

Advancements in language foundation models have primarily fueled the recent surge in artificial intelligence. In contrast, generative learning of non-textual modalities, especially…

cs.CV2024

VideoPoet: A Large Language Model for Zero-Shot Video Generation

Dan Kondratyuk, Lijun Yu, Xiuye Gu +28

We present VideoPoet, a language model capable of synthesizing high-quality video, with matching audio, from a large variety of conditioning signals. VideoPoet employs a decoder-on…

cs.LG2023

Score-based Continuous-time Discrete Diffusion Models

Haoran Sun, Lijun Yu, Bo Dai +2

Score-based modeling through stochastic differential equations (SDEs) has provided a new perspective on diffusion models, and demonstrated superior performance on continuous data.…

cs.CV2018

Traffic Danger Recognition With Surveillance Cameras Without Training Data

Lijun Yu, Dawei Zhang, Xiangqun Chen +1

We propose a traffic danger recognition model that works with arbitrary traffic surveillance cameras to identify and predict car crashes. There are too many cameras to monitor manu…

cs.CV2026

SSD: Spatially Speculative Decoding Accelerates Autoregressive Image Generation

Shilong Xiang, Zirui Zhang, Lijun Yu +1

Autoregressive models excel in visual generation by treating images as 1D sequences of discrete tokens, mirroring language modeling. However, this flattening discards the intrinsic…

cs.AI2025

Tracing LLM Reasoning Processes with Strategic Games: A Framework for Planning, Revision, and Resource-Constrained Decision Making

Xiaopeng Yuan, Xingjian Zhang, Ke Xu +5

Large language models (LLMs) are increasingly used for tasks that require complex reasoning. Most benchmarks focus on final outcomes but overlook the intermediate reasoning steps -…

cs.LG2024

Unified Discrete Diffusion for Categorical Data

Lingxiao Zhao, Xueying Ding, Lijun Yu +1

Discrete diffusion models have seen a surge of attention with applications on naturally discrete data such as language and graphs. Although discrete-time discrete diffusion has bee…

cs.CV2024

A Versatile Diffusion Transformer with Mixture of Noise Levels for Audiovisual Generation

Gwanghyun Kim, Alonso Martinez, Yu-Chuan Su +8

Training diffusion models for audiovisual sequences allows for a range of generation tasks by learning conditional distributions of various input-output combinations of the two mod…

cs.CV2025

Language-Guided Image Tokenization for Generation

Kaiwen Zha, Lijun Yu, Alireza Fathi +4

Image tokenization, the process of transforming raw image pixels into a compact low-dimensional latent representation, has proven crucial for scalable and efficient image generatio…

cs.CV2023

MAGVIT: Masked Generative Video Transformer

Lijun Yu, Yong Cheng, Kihyuk Sohn +8

We introduce the MAsked Generative VIdeo Transformer, MAGVIT, to tackle various video synthesis tasks with a single model. We introduce a 3D tokenizer to quantize a video into spat…

cs.CL2023

DocumentNet: Bridging the Data Gap in Document Pre-Training

Lijun Yu, Jin Miao, Xiaoyu Sun +4

Document understanding tasks, in particular, Visually-rich Document Entity Retrieval (VDER), have gained significant attention in recent years thanks to their broad applications in…

cs.CL2025

Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Gheorghe Comanici, Eric Bieber, Mike Schaekermann +3431

In this report, we introduce the Gemini 2.X model family: Gemini 2.5 Pro and Gemini 2.5 Flash, as well as our earlier Gemini 2.0 Flash and Flash-Lite models. Gemini 2.5 Pro is our…

cs.LG2025

Efficient LLM Jailbreak via Adaptive Dense-to-sparse Constrained Optimization

Kai Hu, Weichen Yu, Yining Li +7

Recent research indicates that large language models (LLMs) are susceptible to jailbreaking attacks that can generate harmful content. This paper introduces a novel token-level att…

cs.CV2022

Argus++: Robust Real-time Activity Detection for Unconstrained Video Streams with Overlapping Cube Proposals

Lijun Yu, Yijun Qian, Wenhe Liu +1

Activity detection is one of the attractive computer vision tasks to exploit the video streams captured by widely installed cameras. Although achieving impressive performance, conv…

physics.optics2021

Generation and dynamics of soliton and soliton molecules from a VSe2/GO-based fiber laser

Benhai Wang, Haobin Han, Lijun Yu +2

Recently, in addition to exploring the application of new saturable absorber devices in fiber lasers, soliton dynamics has also become a focus of current research. In this article,…

cs.CV2026

Cubic Discrete Diffusion: Discrete Visual Generation on High-Dimensional Representation Tokens

Yuqing Wang, Chuofan Ma, Zhijie Lin +7

Visual generation with discrete tokens has gained significant attention as it enables a unified token prediction paradigm shared with language models, promising seamless multimodal…

cs.CV2023

Photorealistic Video Generation with Diffusion Models

Agrim Gupta, Lijun Yu, Kihyuk Sohn +6

We present W.A.L.T, a transformer-based approach for photorealistic video generation via diffusion modeling. Our approach has two key design decisions. First, we use a causal encod…

cs.CV2020

Training-free Monocular 3D Event Detection System for Traffic Surveillance

Lijun Yu, Peng Chen, Wenhe Liu +2

We focus on the problem of detecting traffic events in a surveillance scenario, including the detection of both vehicle actions and traffic collisions. Existing event detection sys…

cs.CV2026

Rethinking Generative Image Pretraining: How Far Are We From Scaling Up Next-Pixel Prediction?

Xinchen Yan, Chen Liang, Lijun Yu +3

This paper investigates the scaling properties of autoregressive next-pixel prediction, a simple, end-to-end yet under-explored framework for unified vision models. Starting with i…

cs.LG2018

MOBA-Slice: A Time Slice Based Evaluation Framework of Relative Advantage between Teams in MOBA Games

Lijun Yu, Dawei Zhang, Xiangqun Chen +1

Multiplayer Online Battle Arena (MOBA) is currently one of the most popular genres of digital games around the world. The domain of knowledge contained in these complicated games i…

cs.AI2026

Closing the Loop on Latent Reasoning via Test-Time Reconstruction

Xiaopeng Yuan, Haibo Jin, Ye Yu +4

Recent work moves intermediate reasoning from natural-language traces into latent or cache-level representations to reduce token overhead and avoid a discrete communication bottlen…

cs.CV2026

Archon: A Unified Multimodal Model for Holistic Digital Human Generation

Chong Bao, Shichen Liu, Lijun Yu +9

Digital humans are fundamental to immersive interaction, yet creating a unified model for holistic modalities, including text, audio, motion, and visual content, remains an open ch…

cs.CV2023

SPAE: Semantic Pyramid AutoEncoder for Multimodal Generation with Frozen LLMs

Lijun Yu, Yong Cheng, Zhiruo Wang +10

In this work, we introduce Semantic Pyramid AutoEncoder (SPAE) for enabling frozen LLMs to perform both understanding and generation tasks involving non-linguistic modalities such…

cs.CV2025

UISim: An Interactive Image-Based UI Simulator for Dynamic Mobile Environments

Jiannan Xiang, Yun Zhu, Lei Shu +6

Developing and testing user interfaces (UIs) and training AI agents to interact with them are challenging due to the dynamic and diverse nature of real-world mobile environments. E…