From the 1 of 11 linked papers with an AI index.
4 papers · 1 filter
Qwen-CUA: Native Computer Use for (almost) Everything
Dunjie Lu, Shuai Bai, Tianyi Bai +42
Native computer use offers a general interface for agents to operate almost any software available to people, but requires long-horizon state tracking, large-scale interactive expe…
Decompose Sparsely Where You Should, Absorb Densely Where You Should No
Ruixuan Deng, Zehao Jin, Zekun Wang +1
Sparse autoencoders (SAEs) are typically trained to reconstruct the \textbf{entire} residual stream through a sparse dictionary, implicitly assuming that all activation content is…
SlimQwen: Exploring the Pruning and Distillation in Large MoE Model Pre-training
Shengkun Tang, Zekun Wang, Bo Zheng +7
Structured pruning and knowledge distillation (KD) are typical techniques for compressing large language models, but it remains unclear how they should be applied at pretraining sc…
Test-Time Compositional Generalization in Diffusion Models via Concept Discovery
Zekun Wang, Anant Gupta, Tianyi Zhu +1
Compositional generalization requires models to produce novel configurations from familiar parts. In diffusion models, prior compositional generation methods typically assume that…