1 citations · 1 across the 1 of their papers we have counts for
7 papers
VueBuds: Visual Intelligence with Wireless Earbuds
Maruchi Kim, Rasya Fawwaz, Zhi Yang Lim +4
Despite their ubiquity, wireless earbuds remain audio-centric due to size and power constraints. We present VueBuds, the first camera-integrated wireless earbuds for egocentric vis…
Point-It-Out: Benchmarking Embodied Reasoning for Vision Language Models in Multi-Stage Visual Grounding
Haotian Xue, Yunhao Ge, Yu Zeng +4
Vision-Language Models (VLMs) have demonstrated impressive world knowledge across a wide range of tasks, making them promising candidates for embodied reasoning applications. Howev…
Cosmos World Foundation Model Platform for Physical AI
NVIDIA, :, Niket Agarwal +76
Physical AI needs to be trained digitally first. It needs a digital twin of itself, the policy model, and a digital twin of the world, the world model. In this paper, we present th…
Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation
Lu Ling, Chen-Hsuan Lin, Tsung-Yi Lin +7
Synthesizing interactive 3D scenes from text is essential for gaming, virtual reality, and embodied AI. However, existing methods face several challenges. Learning-based approaches…
Cosmos-Transfer1: Conditional World Generation with Adaptive Multimodal Control
NVIDIA, :, Hassan Abu Alhaija +38
We introduce Cosmos-Transfer, a conditional world generation model that can generate world simulations based on multiple spatial control inputs of various modalities such as segmen…
Edify Image: High-Quality Image Generation with Pixel Space Laplacian Diffusion Models
NVIDIA, :, Yuval Atzmon +29
We introduce Edify Image, a family of diffusion models capable of generating photorealistic image content with pixel-perfect accuracy. Edify Image utilizes cascaded pixel-space dif…