5 citations · 5 across the 3 of their papers we have counts for
3 papers
Multi-Modal Hallucination Control by Visual Information Grounding
Alessandro Favero, Luca Zancato, Matthew Trager +5
Generative Vision-Language Models (VLMs) are prone to generate plausible-sounding textual answers that, however, are not always grounded in the input image. We investigate this phe…
SplatArmor: Articulated Gaussian splatting for animatable humans from monocular RGB videos
Rohit Jena, Ganesh Subramanian Iyer, Siddharth Choudhary +3
We propose SplatArmor, a novel approach for recovering detailed and animatable human models by `armoring' a parameterized body model with 3D Gaussians. Our approach represents the…
Mesh Strikes Back: Fast and Efficient Human Reconstruction from RGB videos
Rohit Jena, Pratik Chaudhari, James Gee +3
Human reconstruction and synthesis from monocular RGB videos is a challenging problem due to clothing, occlusion, texture discontinuities and sharpness, and framespecific pose chan…