3 papers
cs.CV2024
CatFree3D: Category-agnostic 3D Object Detection with Diffusion
Wenjing Bian, Zirui Wang, Andrea Vedaldi
Image-based 3D object detection is widely employed in applications such as autonomous vehicles and robotics, yet current systems struggle with generalisation due to complex problem…
cs.CL2024
Improving Language Understanding from Screenshots
Tianyu Gao, Zirui Wang, Adithya Bhaskar +1
An emerging family of language models (LMs), capable of processing both text and images within a single visual view, has the promise to unlock complex tasks such as chart understan…
cs.CV2023
Guiding Image Captioning Models Toward More Specific Captions
Simon Kornblith, Lala Li, Zirui Wang +1
Image captioning is conventionally formulated as the task of generating captions for images that match the distribution of reference image-caption pairs. However, reference caption…