3 papers
cs.CV2026
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction
Zhongbin Guo, Jiahao Xie, Dongling Xiao +5
While Multimodal Large Language Models (MLLMs) have achieved remarkable progress, visual understanding and generation are typically treated as divergent objectives. Existing unifie…
cs.LG2026
PRISM: Priority-aware Rubric Internalization via Structured Multimodal Data Synthesis
Xiaomin He, Dongling Xiao, Jiahao Xie +4
Real-world multimodal instructions often bundle multiple requirements with unequal importance, yet most multimodal training data still reduce instruction following to answering one…
cs.CL2026
How Order-Sensitive Are LLMs? OrderProbe for Deterministic Structural Reconstruction
Zhaolu Kang, Yingjie He, Kehan Jiang +18
Large language models (LLMs) excel at semantic understanding, yet their ability to reconstruct internal structure from scrambled inputs remains underexplored. Sentence-level restor…