77 citations · 341 across the 51 of their papers we have counts for
12 papers · 1 filter
InstructVideo: Instructing Video Diffusion Models with Human Feedback
Hangjie Yuan, Shiwei Zhang, Xiang Wang +7
Diffusion models have emerged as the de facto paradigm for video generation. However, their reliance on web-scale data of varied quality often yields results that are visually unap…
Charting New Territories: Exploring the Geographic and Geospatial Capabilities of Multimodal LLMs
Jonathan Roberts, Timo Lüddecke, Rehan Sheikh +2
Multimodal large language models (MLLMs) have shown remarkable capabilities across a broad range of tasks but their knowledge and abilities in the geographic and geospatial domains…
Visual Data-Type Understanding does not emerge from Scaling Vision-Language Models
Vishaal Udandarao, Max F. Burg, Samuel Albanie +1
Recent advances in the development of vision-language models (VLMs) are yielding remarkable success in recognizing visual semantic content, including impressive instances of compos…
Simple Baselines for Interactive Video Retrieval with Questions and Answers
Kaiqu Liang, Samuel Albanie
To date, the majority of video retrieval systems have been optimized for a "single-shot" scenario in which the user submits a query in isolation, ignoring previous interactions wit…
RLIPv2: Fast Scaling of Relational Language-Image Pre-training
Hangjie Yuan, Shiwei Zhang, Xiang Wang +7
Relational Language-Image Pre-training (RLIP) aims to align vision representations with relational texts, thereby advancing the capability of relational reasoning in computer visio…
arXiVeri: Automatic table verification with GPT
Gyungin Shin, Weidi Xie, Samuel Albanie
Without accurate transcription of numerical data in scientific documents, a scientist cannot draw accurate conclusions. Unfortunately, the process of copying numerical data from on…