most citedTimeStep Master: Asymmetrical Mixture of Timestep LoRA Experts for Versatile and Efficient Diffusion Models in Vision

1 citations · 1 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CV2026

UniWeTok: An Unified Binary Tokenizer with Codebook Size for Unified Multimodal Large Language Model

Shaobin Zhuang, Yuang Ai, Jiaming Han +12

Unified Multimodal Large Language Models (MLLMs) require a visual representation that simultaneously supports high-fidelity reconstruction, complex semantic extraction, and generat…

cs.CV2025

WeTok: Powerful Discrete Tokenization for High-Fidelity Visual Reconstruction

Shaobin Zhuang, Yiwei Guo, Canmiao Fu +7

Visual tokenizer is a critical component for vision generation. However, the existing tokenizers often face unsatisfactory trade-off between compression ratios and reconstruction f…

cs.CV2025

Video-GPT via Next Clip Diffusion

Shaobin Zhuang, Zhipeng Huang, Ying Zhang +6

GPT has shown its remarkable success in natural language processing. However, the language sequence is not sufficient to describe spatial-temporal details in the visual world. Alte…

cs.CV2025★ 1 cited

TimeStep Master: Asymmetrical Mixture of Timestep LoRA Experts for Versatile and Efficient Diffusion Models in Vision

Shaobin Zhuang, Yiwei Guo, Yanbo Ding +7

Diffusion models have driven the advancement of vision generation over the past years. However, it is often difficult to apply these large models in downstream tasks, due to massiv…

cs.CV2025

Get In Video: Add Anything You Want to the Video

Shaobin Zhuang, Zhipeng Huang, Binxin Yang +7

Video editing increasingly demands the ability to incorporate specific real-world instances into existing footage, yet current approaches fundamentally fail to capture the unique v…