1 citations · 1 across the 2 of their papers we have counts for
3 papers · 1 filter
Are We There Yet? Exploring the Capabilities of MLLMs in Assistive AI Applications
Shayon Dasgupta, Avijit Dasgupta, C. V. Jawahar
Multimodal Large Language Models (MLLMs) have redefined visual understanding by combining vision encoders with large-scale language models. This unified architecture enables strong…
DriveSafe: A Framework for Risk Detection and Safety Suggestions in Driving Scenarios
Sainithin Artham, Shankar Gangisetty, Avijit Dasgupta +1
Comprehensive situational awareness is essential for autonomous vehicles operating in safety-critical environments, as it enables the identification and mitigation of potential ris…
Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives
Kristen Grauman, Andrew Westbury, Lorenzo Torresani +98
We present Ego-Exo4D, a diverse, large-scale multimodal multiview video dataset and benchmark challenge. Ego-Exo4D centers around simultaneously-captured egocentric and exocentric…