2 papers
cs.CV2026
Learning to Rank Caption Chains for Video-Text Alignment
Ansel Blume, Burak Uzkent, Shalini Chaudhuri +1
Direct preference optimization (DPO) is an effective technique to train language models to generate preferred over dispreferred responses. However, this binary "winner-takes-all" a…
cs.CV2026
Perceptio: Perception Enhanced Vision Language Models via Spatial Token Generation
Yuchen Li, Amanmeet Garg, Shalini Chaudhuri +2
Large Vision Language Models (LVLMs) excel at semantic understanding but struggle with fine grained spatial grounding, as the model must implicitly infer complex geometry without e…