activity
20242026
collaborators

6 papers

cs.CV2026

FlowCoMotion: Text-to-Motion Generation via Token-Latent Flow Modeling

Dawei Guan, Di Yang, Chengjie Jin +1

Text-to-motion generation is driven by learning motion representations for semantic alignment with language. Existing methods rely on either continuous or discrete motion represent…

cs.LG2026

Optimal Scaling Needs Optimal Norm

Oleg Filatov, Jiangtao Wang, Jan Ebert +1

Despite recent progress in optimal hyperparameter transfer under model and dataset scaling, no unifying explanatory principle has been established. For Adam and Scion optimizers, w…

cs.LG2025

Data Pruning in Generative Diffusion Models

Rania Briq, Jiangtao Wang, Stefan Kesselheim

Data pruning is the problem of identifying a core subset that is most beneficial to training and discarding the remainder. While pruning strategies are well studied for discriminat…

cs.DC2025

Memory and Bandwidth are All You Need for Fully Sharded Data Parallel

Jiangtao Wang, Jan Ebert, Oleg Filatov +1

Transformer models have revolutionized a wide spectrum of disciplines, especially in language processing. The recent success has proven that model size scalability is crucial for a…

cs.LG2025

Time Transfer: On Optimal Learning Rate and Batch Size In The Infinite Data Limit

Oleg Filatov, Jan Ebert, Jiangtao Wang +1

One of the main challenges in optimal scaling of large language models (LLMs) is the prohibitive cost of hyperparameter tuning, particularly learning rate and batch size .…

cs.CV2024

Scaling Image Tokenizers with Grouped Spherical Quantization

Jiangtao Wang, Zhen Qin, Yifan Zhang +4

Vision tokenizers have gained a lot of attraction due to their scalability and compactness; previous works depend on old-school GAN-based hyperparameters, biased comparisons, and a…