Showing 2024Show all
2 papers · 1 filter
cs.CV2024
Cubify Anything: Scaling Indoor 3D Object Detection
Justin Lazarow, David Griffiths, Gefen Kohavi +2
We consider indoor 3D object detection with respect to a single RGB(-D) frame acquired from a commodity handheld device. We seek to significantly advance the status quo with respec…
cs.CV2024
4M-21: An Any-to-Any Vision Model for Tens of Tasks and Modalities
Roman Bachmann, OÄuzhan Fatih Kar, David Mizrahi +6
Current multimodal and multitask foundation models like 4M or UnifiedIO show promising results, but in practice their out-of-the-box abilities to accept diverse inputs and perform…