1 citations · 2 across the 7 of their papers we have counts for
8 papers · 1 filter
Can Vision-Language Models Answer Face to Face Questions in the Real-World?
Reza Pourreza, Rishit Dagli, Apratim Bhattacharyya +3
AI models have made significant strides in recent years in their ability to describe and answer questions about real-world images. They have also made progress in the ability to co…
AirLetters: An Open Video Dataset of Characters Drawn in the Air
Rishit Dagli, Guillaume Berger, Joanna Materzynska +2
We introduce AirLetters, a new video dataset consisting of real-world videos of human-generated, articulated motions. Specifically, our dataset requires a vision model to predict l…
What to Say and When to Say it: Live Fitness Coaching as a Testbed for Situated Interaction
Sunny Panchal, Apratim Bhattacharyya, Guillaume Berger +10
Vision-language models have shown impressive progress in recent years. However, existing models are largely limited to turn-based interactions, where each turn must be stepped (i.e…
HexaGen3D: StableDiffusion is just one step away from Fast and Diverse Text-to-3D Generation
Antoine Mercier, Ramin Nakhli, Mahesh Reddy +4
Despite the latest remarkable advances in generative modeling, efficient generation of high-quality 3D assets from textual prompts remains a difficult task. A key challenge lies in…
Efficient neural supersampling on a novel gaming dataset
Antoine Mercier, Ruan Erasmus, Yashesh Savani +3
Real-time rendering for video games has become increasingly challenging due to the need for higher resolutions, framerates and photorealism. Supersampling has emerged as an effecti…
Is end-to-end learning enough for fitness activity recognition?
Antoine Mercier, Guillaume Berger, Sunny Panchal +5
End-to-end learning has taken hold of many computer vision tasks, in particular, related to still images, with task-specific optimization yielding very strong performance. Neverthe…