11 citations · 16 across the 16 of their papers we have counts for
4 papers · 2 filters
freePruner: A Training-free Approach for Large Multimodal Model Acceleration
Bingxin Xu, Yuzhang Shang, Yunhao Ge +2
Large Multimodal Models (LMMs) have demonstrated impressive capabilities in visual-language tasks but face significant deployment challenges due to their high computational demands…
Edify 3D: Scalable High-Quality 3D Asset Generation
NVIDIA, :, Maciej Bala +22
We introduce Edify 3D, an advanced solution designed for high-quality 3D asset generation. Our method first synthesizes RGB and surface normal images of the described object at mul…
Edify Image: High-Quality Image Generation with Pixel Space Laplacian Diffusion Models
NVIDIA, :, Yuval Atzmon +29
We introduce Edify Image, a family of diffusion models capable of generating photorealistic image content with pixel-perfect accuracy. Edify Image utilizes cascaded pixel-space dif…
Visual Fact Checker: Enabling High-Fidelity Detailed Caption Generation
Yunhao Ge, Xiaohui Zeng, Jacob Samuel Huffman +3
Existing automatic captioning methods for visual content face challenges such as lack of detail, content hallucination, and poor instruction following. In this work, we propose Vis…