6 papers · 1 filter
LENS: Adaptive Spatio-Temporal Zooming for Keyframe Sampling in Long-Form Videos
Ce Zhang, Jinxi He, Katia Sycara +1
Despite rapid progress in Multi-modal Large Language Models (MLLMs), understanding long-form videos is still bottlenecked by limited context windows. While recent keyframe sampling…
Self-Correcting Decoding with Generative Feedback for Mitigating Hallucinations in Large Vision-Language Models
Ce Zhang, Zifu Wan, Zhehan Kan +7
While recent Large Vision-Language Models (LVLMs) have shown remarkable performance in multi-modal tasks, they are prone to generating hallucinatory text responses that do not alig…
Enhancing Vision-Language Few-Shot Adaptation with Negative Learning
Ce Zhang, Simon Stepputtis, Katia Sycara +1
Large-scale pre-trained Vision-Language Models (VLMs) have exhibited impressive zero-shot performance and transferability, allowing them to adapt to downstream tasks in a data-effi…
Dual Prototype Evolving for Test-Time Generalization of Vision-Language Models
Ce Zhang, Simon Stepputtis, Katia Sycara +1
Test-time adaptation, which enables models to generalize to diverse data with unlabeled test samples, holds significant value in real-world scenarios. Recently, researchers have ap…
HiKER-SGG: Hierarchical Knowledge Enhanced Robust Scene Graph Generation
Ce Zhang, Simon Stepputtis, Joseph Campbell +2
Being able to understand visual scenes is a precursor for many downstream tasks, including autonomous driving, robotics, and other vision-based approaches. A common approach enabli…
Test-time Distribution Learning Adapter for Cross-modal Visual Reasoning
Yi Zhang, Ce Zhang
Vision-Language Pre-Trained (VLP) models, such as CLIP, have demonstrated remarkable effectiveness in learning generic visual representations. Several approaches aim to efficiently…