2 citations · 2 across the 3 of their papers we have counts for
3 papers
EgoSocial: Benchmarking Proactive Intervention Ability of Omnimodal LLMs via Egocentric Social Interaction Perception
Xijun Wang, Tanay Sharma, Achin Kulshrestha +3
As AR/VR technologies become integral to daily life, there's a growing need for AI that understands human social dynamics from an egocentric perspective. However, current LLMs ofte…
Geometry Aware Passthrough Mitigates Cybersickness
Trishia El Chemaly, Mohit Goyal, Tinglin Duan +8
Virtual Reality headsets isolate users from the real-world by restricting their perception to the virtual-world. Video See-Through (VST) headsets address this by utilizing world-fa…
BRAVE: Broadening the visual encoding of vision-language models
Oğuzhan Fatih Kar, Alessio Tonioni, Petra Poklukar +3
Vision-language models (VLMs) are typically composed of a vision encoder, e.g. CLIP, and a language model (LM) that interprets the encoded features to solve downstream tasks. Despi…