4 papers
Grounding the Score: Explicit Visual Premise Verification for Reliable Vision-Language Process Reward Models
Junxin Wang, Dai Guan, Weijie Qiu +7
Vision-language process reward models (VL-PRMs) are increasingly used to score intermediate reasoning steps and rerank candidates under test-time scaling. However, they often funct…
Rationale Matters: Learning Transferable Rubrics via Proxy-Guided Critique for VLM Reward Models
Weijie Qiu, Dai Guan, Junxin Wang +6
Generative reward models (GRMs) for vision-language models (VLMs) often evaluate outputs via a three-stage pipeline: rubric generation, criterion-based scoring, and a final verdict…
Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models
Yongyu Mu, Hengyu Li, Junxin Wang +7
Previous work on augmenting large multimodal models (LMMs) for text-to-image (T2I) generation has focused on enriching the input space of in-context learning (ICL). This includes p…
Early Exit Is a Natural Capability in Transformer-based Models: An Empirical Study on Early Exit without Joint Optimization
Weiqiao Shan, Long Meng, Tong Zheng +5
Large language models (LLMs) exhibit exceptional performance across various downstream tasks. However, they encounter limitations due to slow inference speeds stemming from their e…