7 papers
Covering Human Action Space for Computer Use: Data Synthesis and Benchmark
Miaosen Zhang, Xiaohan Zhao, Zhihong Tan +14
Computer-use agents (CUAs) automate on-screen work, as illustrated by GPT-5.4 and Claude. Yet their reliability on complex, low-frequency interactions is still poor, limiting user…
XPERT: Expert Knowledge Transfer for Effective Training of Language Models
Chang Liu, Boyu Shi, Xu Yang +1
Mixture-of-Experts (MoE) language models organize knowledge into explicitly routed expert modules, making expert-level representations traceable and analyzable. By analyzing expert…
Learngene Search Across Multiple Datasets for Building Variable-Sized Models
Boyu Shi, Junbo Zhou, Chang Liu +3
Deep learning methods are widely used under diverse resource constraints, resulting in models of varying sizes, such as the Vision Transformer (ViT) series. Deploying these models…
Towards On-Policy SFT: Distribution Discriminant Theory and its Applications in LLM Training
Miaosen Zhang, Yishan Liu, Shuxia Lin +8
Supervised fine-tuning (SFT) is computationally efficient but often yields inferior generalization compared to reinforcement learning (RL). This gap is primarily driven by RL's use…
Enhancing Multimodal In-Context Learning for Image Classification through Coreset Optimization
Huiyi Chen, Jiawei Peng, Kaihua Tang +2
In-context learning (ICL) enables Large Vision-Language Models (LVLMs) to adapt to new tasks without parameter updates, using a few demonstrations from a large support set. However…
Distribution-Conditional Generation: From Class Distribution to Creative Generation
Fu Feng, Yucheng Xie, Xu Yang +2
Text-to-image (T2I) diffusion models are effective at producing semantically aligned images, but their reliance on training data distributions limits their ability to synthesize tr…