1 citations · 1 across the 20 of their papers we have counts for
Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
Fill the GAP: A Granular Alignment Paradigm for Visual Reasoning in Multimodal Large Language Models
Yanting Miao, Yutao Sun, Dexin Wang +8
Visual latent reasoning lets a multimodal large language model (MLLM) create intermediate visual evidence as continuous tokens, avoiding external tools or image generators. However…
cs.CV2026
Grounding the Score: Explicit Visual Premise Verification for Reliable Vision-Language Process Reward Models
Junxin Wang, Dai Guan, Weijie Qiu +7
Vision-language process reward models (VL-PRMs) are increasingly used to score intermediate reasoning steps and rerank candidates under test-time scaling. However, they often funct…
cs.CV2026
Rationale Matters: Learning Transferable Rubrics via Proxy-Guided Critique for VLM Reward Models
Weijie Qiu, Dai Guan, Junxin Wang +6
Generative reward models (GRMs) for vision-language models (VLMs) often evaluate outputs via a three-stage pipeline: rubric generation, criterion-based scoring, and a final verdict…