55 citations · 64 across the 5 of their papers we have counts for
5 papers
Cubify Anything: Scaling Indoor 3D Object Detection
Justin Lazarow, David Griffiths, Gefen Kohavi +2
We consider indoor 3D object detection with respect to a single RGB(-D) frame acquired from a commodity handheld device. We seek to significantly advance the status quo with respec…
Understanding Alignment in Multimodal LLMs: A Comprehensive Study
Elmira Amirloo, Jean-Philippe Fauconnier, Christoph Roesmann +8
Preference alignment has become a crucial component in enhancing the performance of Large Language Models (LLMs), yet its impact in Multimodal Large Language Models (MLLMs) remains…
SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
Mingze Xu, Mingfei Gao, Zhe Gan +5
We propose SlowFast-LLaVA (or SF-LLaVA for short), a training-free video large language model (LLM) that can jointly capture detailed spatial semantics and long-range temporal cont…
4M-21: An Any-to-Any Vision Model for Tens of Tasks and Modalities
Roman Bachmann, Oğuzhan Fatih Kar, David Mizrahi +6
Current multimodal and multitask foundation models like 4M or UnifiedIO show promising results, but in practice their out-of-the-box abilities to accept diverse inputs and perform…
GAUDI: A Neural Architect for Immersive 3D Scene Generation
Miguel Angel Bautista, Pengsheng Guo, Samira Abnar +9
We introduce GAUDI, a generative model capable of capturing the distribution of complex and realistic 3D scenes that can be rendered immersively from a moving camera. We tackle thi…