activity
20172024
most citedRLIP: Relational Language-Image Pre-training for Human-Object Interaction Detection

29 citations · 86 across the 14 of their papers we have counts for

collaborators

14 papers

cs.LG20246 cited

On scalable oversight with weak LLMs judging strong LLMs

Zachary Kenton, Noah Y. Siegel, János Kramár +8

Scalable oversight protocols aim to enable humans to accurately supervise superhuman AI. In this paper we study debate, where two AI's compete to convince a judge; consultancy, whe…

cs.LG2024

HelloFresh: LLM Evaluations on Streams of Real-World Human Editorial Actions across X Community Notes and Wikipedia edits

Tim Franzmeyer, Aleksandar Shtedritski, Samuel Albanie +3

Benchmarks have been essential for driving progress in machine learning. A better understanding of LLM capabilities on real world tasks is vital for safe development. Designing ade…

eess.AS2024

A SOUND APPROACH: Using Large Language Models to generate audio descriptions for egocentric text-audio retrieval

Andreea-Maria Oncescu, João F. Henriques, Andrew Zisserman +2

Video databases from the internet are a valuable source of text-audio retrieval datasets. However, given that sound and vision streams represent different "views" of the data, trea…

cs.CV20231 cited

Simple Baselines for Interactive Video Retrieval with Questions and Answers

Kaiqu Liang, Samuel Albanie

To date, the majority of video retrieval systems have been optimized for a "single-shot" scenario in which the user submits a query in isolation, ignoring previous interactions wit…

cs.CV20232 cited

RLIPv2: Fast Scaling of Relational Language-Image Pre-training

Hangjie Yuan, Shiwei Zhang, Xiang Wang +7

Relational Language-Image Pre-training (RLIP) aims to align vision representations with relational texts, thereby advancing the capability of relational reasoning in computer visio…

cs.CL202319 cited

GPT4GEO: How a Language Model Sees the World's Geography

Jonathan Roberts, Timo Lüddecke, Sowmen Das +2

Large language models (LLMs) have shown remarkable capabilities across a broad range of tasks involving question answering and the generation of coherent text and code. Comprehensi…