activity
20242026
collaborators

10 papers

cs.LG2026

The Efficiency Gap in Byte Modeling

Celine Lee, Jing Nathan Yan, Chen Liang +9

Modern language models have historically relied on two dominant design choices: subword tokenization and autoregressive (AR) ordering. These design decisions bake in priors that di…

cs.CV2026

RealCam: Real-Time Novel-View Video Generation with Interactive Camera Control

Youcan Xu, Jiaxin Shi, Zhen Wang +5

Camera-controlled video-to-video (V2V) generation enables dynamic viewpoint synthesis from monocular footage, holding immense potential for interactive filmmaking and live broadcas…

cs.CV2026

DiT as Real-Time Rerenderer: Streaming Video Stylization with Autoregressive Diffusion Transformer

Hengye Lyu, Zisu Li, Yue Hong +4

Recent advances in video generation models has significantly accelerated video generation and related downstream tasks. Among these, video stylization holds important research valu…

cs.LG2026

Generative Frontiers: Why Evaluation Matters for Diffusion Language Models

Patrick Pynadath, Jiaxin Shi, Ruqi Zhang

Diffusion language models have seen exciting recent progress, offering far more flexibility in generative trajectories than autoregressive models. This flexibility has motivated a…

cs.CV2026

Real-Time Motion-Controllable Autoregressive Video Diffusion

Kesen Zhao, Jiaxin Shi, Beier Zhu +5

Real-time motion-controllable video generation remains challenging due to the inherent latency of bidirectional diffusion models and the lack of effective autoregressive (AR) appro…

cs.CV2025

SpriteHand: Real-Time Versatile Hand-Object Interaction with Autoregressive Video Generation

Zisu Li, Hengye Lyu, Jiaxin Shi +4

Modeling and synthesizing complex hand-object interactions remains a significant challenge, even for state-of-the-art physics engines. Conventional simulation-based approaches rely…