collaborators

9 papers

cs.LG2026

Small LLMs: Pruning vs. Training from Scratch

Yufeng Xu, Taiming Lu, Kunjun Li +3

Pruning promises a shortcut to strong small language models. In this work, we examine this promise by pruning Llama-3.1-8B at pruning ratios of 0.5--0.8 with six methods spanning d…

cs.CV2026

i1: A Simple and Fully Open Recipe for Strong Text-to-Image Models

Boya Zeng, Tianze Luo, Shu Pu +4

Diffusion models have consistently driven progress in text-to-image generation. However, it is challenging to attribute recent progress to specific modeling and data choices: state…

cs.LG2026

Strong Teacher Not Needed? On Distillation in LLM Pretraining

Taiming Lu, Zhuang Liu

Knowledge distillation generally assumes a strong-to-weak relationship where stronger teachers yield better students. In this work, we examine this assumption about distillation in…

cs.LG2026

Stronger Normalization-Free Transformers

Mingzhi Chen, Taiming Lu, Jiachen Zhu +2

Although normalization layers have long been viewed as indispensable components of deep learning architectures, the recent introduction of Dynamic Tanh (DyT) has demonstrated that…

cs.CV2025

World-in-World: World Models in a Closed-Loop World

Jiahan Zhang, Muqing Jiang, Nanru Dai +14

Generative world models (WMs) can now simulate worlds with striking visual realism, which naturally raises the question of whether they can endow embodied agents with predictive pe…

cs.CV2025

EvoWorld: Evolving Panoramic World Generation with Explicit 3D Memory

Jiahao Wang, Luoxin Ye, TaiMing Lu +8

Humans possess a remarkable ability to mentally explore and replay 3D environments they have previously experienced. Inspired by this mental process, we present EvoWorld: a world m…