collaborators

5 papers

cs.CV2026

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction

Zhongbin Guo, Jiahao Xie, Dongling Xiao +5

While Multimodal Large Language Models (MLLMs) have achieved remarkable progress, visual understanding and generation are typically treated as divergent objectives. Existing unifie…

cs.LG2026

PRISM: Priority-aware Rubric Internalization via Structured Multimodal Data Synthesis

Xiaomin He, Dongling Xiao, Jiahao Xie +4

Real-world multimodal instructions often bundle multiple requirements with unequal importance, yet most multimodal training data still reduce instruction following to answering one…

cs.CV2026

DecoupleMix: Decoupled Ratio Search and Convex Allocation for Scalable VLM Data Recipes

Jiahao Xie, Zhongbin Guo, Qianle Wang +4

While data curation for Vision Language Models (VLMs) is increasingly active, public practice for constructing pretraining mixtures remains largely heuristic: practitioners stack d…

cs.CL2024

Cherry on Top: Parameter Heterogeneity and Quantization in Large Language Models

Wanyun Cui, Qianle Wang

This paper reveals the phenomenon of parameter heterogeneity in large language models (LLMs). We find that a small subset of "cherry" parameters exhibit a disproportionately large…

cs.CL2024

Ada-Instruct: Adapting Instruction Generators for Complex Reasoning

Wanyun Cui, Qianle Wang

Instructions augmentation is a crucial step for unleashing the full potential of large language models (LLMs) in downstream tasks. Existing Self-Instruct methods primarily simulate…