From the 1 of 4 linked papers with an AI index.
4 papers
CoRe: A Comprehensive Framework for Cross-Image Comparative Reasoning in Vision-Language Models
Lin Peng, Cong Wan, Zeyu Guo +2
The paper introduces CoRe, a framework that improves vision-language models' ability to perform fine-grained cross‑image comparative reasoning by providing a large triplet‑based da…
DataClaw0: Agentic Tailoring Multimodal Data from Raw Streams
Cong Wan, Zeyu Guo, Zijian Cai +6
Raw multimodal streams are abundant but noisy, redundant, and unaligned with any particular training objective. Turning them into supervision today means either brittle heuristics…
ReMoT: Reinforcement Learning with Motion Contrast Triplets
Cong Wan, Zeyu Guo, Jiangyang Li +5
We present ReMoT, a unified training paradigm to systematically address the fundamental shortcomings of VLMs in spatio-temporal consistency -- a critical failure point in navigatio…
Cluster-Aware Neural Collapse Prompt Tuning for Long-Tailed Generalization of Vision-Language Models
Boyang Guo, Liang Li, Lin Peng +3
Prompt learning has emerged as an efficient alternative to fine-tuning pre-trained vision-language models (VLMs). Despite its promise, current methods still struggle to maintain ta…