benchmark evaluation 1cross-image comparative reasoning 1dataset construction 1structured reward learning 1vision-language models 1
From the 1 of 3 linked papers with an AI index.
3 papers
cs.CV2026
CoRe: A Comprehensive Framework for Cross-Image Comparative Reasoning in Vision-Language Models
Lin Peng, Cong Wan, Zeyu Guo +2
The paper introduces CoRe, a framework that improves vision-language models' ability to perform fine-grained cross‑image comparative reasoning by providing a large triplet‑based da…
cs.LG2026
DataClaw0: Agentic Tailoring Multimodal Data from Raw Streams
Cong Wan, Zeyu Guo, Zijian Cai +6
Raw multimodal streams are abundant but noisy, redundant, and unaligned with any particular training objective. Turning them into supervision today means either brittle heuristics…
cs.CV2026
ReMoT: Reinforcement Learning with Motion Contrast Triplets
Cong Wan, Zeyu Guo, Jiangyang Li +5
We present ReMoT, a unified training paradigm to systematically address the fundamental shortcomings of VLMs in spatio-temporal consistency -- a critical failure point in navigatio…