collaborators

7 papers

cs.CV2025

MapTrace: Scalable Data Generation for Route Tracing on Maps

Artemis Panagopoulou, Aveek Purohit, Achin Kulshrestha +2

While Multimodal Large Language Models have achieved human-like performance on many visual and textual reasoning tasks, their proficiency in fine-grained spatial understanding, suc…

cs.HC2025

Sensible Agent: A Framework for Unobtrusive Interaction with Proactive AR Agents

Geonsun Lee, Min Xia, Nels Numan +9

Proactive AR agents promise context-aware assistance, but their interactions often rely on explicit voice prompts or responses, which can be disruptive or socially awkward. We intr…

cs.CV2025

EgoSocial: Benchmarking Proactive Intervention Ability of Omnimodal LLMs via Egocentric Social Interaction Perception

Xijun Wang, Tanay Sharma, Achin Kulshrestha +3

As AR/VR technologies become integral to daily life, there's a growing need for AI that understands human social dynamics from an egocentric perspective. However, current LLMs ofte…

cs.CV2025

EgoTrigger: Toward Audio-Driven Image Capture for Human Memory Enhancement in All-Day Energy-Efficient Smart Glasses

Akshay Paruchuri, Sinan Hersek, Lavisha Aggarwal +6

All-day smart glasses are likely to emerge as platforms capable of continuous contextual sensing, uniquely positioning them for unprecedented assistance in our daily lives. Integra…

cs.HC2025

Geometry Aware Passthrough Mitigates Cybersickness

Trishia El Chemaly, Mohit Goyal, Tinglin Duan +8

Virtual Reality headsets isolate users from the real-world by restricting their perception to the virtual-world. Video See-Through (VST) headsets address this by utilizing world-fa…

cs.CV2025

Omnia de EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric Videos

Chiara Plizzari, Alessio Tonioni, Yongqin Xian +2

Understanding fine-grained temporal dynamics is crucial in egocentric videos, where continuous streams capture frequent, close-up interactions with objects. In this work, we bring…