4 papers
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction
Zhongbin Guo, Jiahao Xie, Dongling Xiao +5
While Multimodal Large Language Models (MLLMs) have achieved remarkable progress, visual understanding and generation are typically treated as divergent objectives. Existing unifie…
PRISM: Priority-aware Rubric Internalization via Structured Multimodal Data Synthesis
Xiaomin He, Dongling Xiao, Jiahao Xie +4
Real-world multimodal instructions often bundle multiple requirements with unequal importance, yet most multimodal training data still reduce instruction following to answering one…
DecoupleMix: Decoupled Ratio Search and Convex Allocation for Scalable VLM Data Recipes
Jiahao Xie, Zhongbin Guo, Qianle Wang +4
While data curation for Vision Language Models (VLMs) is increasingly active, public practice for constructing pretraining mixtures remains largely heuristic: practitioners stack d…
GuideBench: Benchmarking Domain-Oriented Guideline Following for LLM Agents
Lingxiao Diao, Xinyue Xu, Wanxuan Sun +2
Large language models (LLMs) have been widely deployed as autonomous agents capable of following user instructions and making decisions in real-world applications. Previous studies…