76 citations · 76 across the 3 of their papers we have counts for
4 papers
CoVeR: Coverage-Based Token Pruning for Multi-View 3D Reasoning in VLMs
Nhat-Tan Bui, Varshini Elangovan, Arun Reddy Anugu +7
Representing a 3D scene as multi-view images allows 2D VLMs to reason in 3D by reusing priors from pre-training, sidestepping the scarcity of annotated 3D data. However, it produce…
Multimodal Fusion with LLMs for Engagement Prediction in Natural Conversation
Cheng Charles Ma, Kevin Hyekang Joo, Alexandria K. Vail +6
Over the past decade, wearable computing devices (``smart glasses'') have undergone remarkable advancements in sensor technology, design, and processing power, ushering in a new er…
Taming 3DGS: High-Quality Radiance Fields with Limited Resources
Saswat Subhajyoti Mallick, Rahul Goel, Bernhard Kerbl +3
3D Gaussian Splatting (3DGS) has transformed novel-view synthesis with its fast, interpretable, and high-fidelity rendering. However, its resource requirements limit its usability.…
SelfPose: 3D Egocentric Pose Estimation from a Headset Mounted Camera
Denis Tome, Thiemo Alldieck, Patrick Peluse +4
We present a solution to egocentric 3D body pose estimation from monocular images captured from downward looking fish-eye cameras installed on the rim of a head mounted VR device.…