activity
20162026
most citedIMU2CLIP: Multimodal Contrastive Learning for IMU Motion Sensors from Egocentric Videos and Text

9 citations · 17 across the 16 of their papers we have counts for

collaborators
Showing cs.CVShow all

8 papers · 1 filter

cs.CV2026

MoE-ViE: Mixture of Experts Vision Encoder for Efficient Image and Video Understanding

Bonan Zhang, Shiyu Dong, Quan Hung Tran +9

Vision encoders are a critical component of vision-language models, and scaling their capacity effectively improves performance. However, dense scaling increases compute cost and i…

cs.CV2026

R3D: Quantitative 3D Spatial Reasoning for Egocentric Wearables

Maxwell Horton, Wei Lu, Quan Tran +6

Quantitative 3D spatial reasoning from egocentric RGB-D video is a critical capability for next-generation wearable assistants. Yet existing benchmarks do not reflect the challenge…

cs.CV2026

Plan, Watch, Recover: A Benchmark and Architectures for Proactive Procedural Assistance

Kaustav Kundu, Ritvik Shrivastava, Maxim Arap +13

We envision a proactive multi-modal assistant system which gives users real-time step-by-step guidance on a procedural task, autonomously deciding \textit{when} to interrupt, and \…

cs.CV2025

CRAG-MM: Multi-modal Multi-turn Comprehensive RAG Benchmark

Jiaqi Wang, Xiao Yang, Kai Sun +38

Wearable devices such as smart glasses are transforming the way people interact with their surroundings, enabling users to seek information regarding entities in their view. Multi-…

cs.CV2025

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding

Jang Hyun Cho, Andrea Madotto, Effrosyni Mavroudi +26

Vision-language models are integral to computer vision research, yet many high-performing models remain closed-source, obscuring their data, design and training recipe. The researc…

cs.CV2024

SnapNTell: Enhancing Entity-Centric Visual Question Answering with Retrieval Augmented Multimodal LLM

Jielin Qiu, Andrea Madotto, Zhaojiang Lin +7

Vision-extended LLMs have made significant strides in Visual Question Answering (VQA). Despite these advancements, VLLMs still encounter substantial difficulties in handling querie…