collaborators

10 papers

cs.CV2026

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction

Zhongbin Guo, Jiahao Xie, Dongling Xiao +5

While Multimodal Large Language Models (MLLMs) have achieved remarkable progress, visual understanding and generation are typically treated as divergent objectives. Existing unifie…

cs.LG2026

PRISM: Priority-aware Rubric Internalization via Structured Multimodal Data Synthesis

Xiaomin He, Dongling Xiao, Jiahao Xie +4

Real-world multimodal instructions often bundle multiple requirements with unequal importance, yet most multimodal training data still reduce instruction following to answering one…

cs.CV2026

DecoupleMix: Decoupled Ratio Search and Convex Allocation for Scalable VLM Data Recipes

Jiahao Xie, Zhongbin Guo, Qianle Wang +4

While data curation for Vision Language Models (VLMs) is increasingly active, public practice for constructing pretraining mixtures remains largely heuristic: practitioners stack d…

cs.CV2026

Video-MME-Logical: A Controlled Diagnostic Benchmark for Video Temporal-Logical Reasoning

Hohin Kwan, Hongyu Li, Ray Zhang +5

Recent interest in multimodal large language models (MLLMs) raises a central question: can they reason over dynamic visual evidence rather than merely recognize objects or events i…

cs.CV2026

TEVI: Text-Conditioned Editing of Visual Representations via Sparse Autoencoders for Improved Vision-Language Alignment

Sweta Mahajan, Sukrut Rao, Jiahao Xie +2

Vision-language models such as CLIP are highly useful for diverse tasks due to their shared image-text embedding space. Despite this, the image and text embeddings are often poorly…

cs.CV2026

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models

Jiahao Xie, Alessio Tonioni, Nathalie Rauschmayr +2

Reinforcement learning (RL) with verifiable rewards (RLVR) has demonstrated the great potential of enhancing the reasoning abilities in multimodal large language models (MLLMs). Ho…