6 papers
Do Models Share Safety Representations? Cross-Model Steering for Safe Visual Generation
Tobia Poppi, Silvia Cappelletti, Sara Sarto +5
Recent progress in generative modeling has made safety control a central challenge, yet existing approaches remain largely model-specific, requiring retraining or tailored interven…
Learning to Rank Caption Chains for Video-Text Alignment
Ansel Blume, Burak Uzkent, Shalini Chaudhuri +1
Direct preference optimization (DPO) is an effective technique to train language models to generate preferred over dispreferred responses. However, this binary "winner-takes-all" a…
Narrative Aligned Long Form Video Question Answering
Rahul Jain, Keval Doshi, Burak Uzkent +1
Recent progress in multimodal large language models (MLLMs) has led to a surge of benchmarks for long-video reasoning. However, most existing benchmarks rely on localized cues and…
Perceptio: Perception Enhanced Vision Language Models via Spatial Token Generation
Yuchen Li, Amanmeet Garg, Shalini Chaudhuri +2
Large Vision Language Models (LVLMs) excel at semantic understanding but struggle with fine grained spatial grounding, as the model must implicitly infer complex geometry without e…
CounterVid: Counterfactual Video Generation for Mitigating Action and Temporal Hallucinations in Video-Language Models
Tobia Poppi, Burak Uzkent, Amanmeet Garg +7
Video-language models (VLMs) achieve strong multimodal understanding but remain prone to hallucinations, especially when reasoning about actions and temporal order. Existing mitiga…
From Frames to Clips: Training-free Adaptive Key Clip Selection for Long-Form Video Understanding
Guangyu Sun, Archit Singhal, Burak Uzkent +3
Video Large Language Models (VLMs) have achieved strong performance on various vision-language tasks, yet their practical use is limited by the massive number of visual tokens prod…