activity
20162025
most citedQuickSRNet: Plain Single-Image Super-Resolution Architecture for Faster Inference on Mobile Platforms

1 citations · 2 across the 7 of their papers we have counts for

collaborators
Showing cs.CVShow all

8 papers · 1 filter

cs.CV2025

Can Vision-Language Models Answer Face to Face Questions in the Real-World?

Reza Pourreza, Rishit Dagli, Apratim Bhattacharyya +3

AI models have made significant strides in recent years in their ability to describe and answer questions about real-world images. They have also made progress in the ability to co…

cs.CV2024

AirLetters: An Open Video Dataset of Characters Drawn in the Air

Rishit Dagli, Guillaume Berger, Joanna Materzynska +2

We introduce AirLetters, a new video dataset consisting of real-world videos of human-generated, articulated motions. Specifically, our dataset requires a vision model to predict l…

cs.CV2024

What to Say and When to Say it: Live Fitness Coaching as a Testbed for Situated Interaction

Sunny Panchal, Apratim Bhattacharyya, Guillaume Berger +10

Vision-language models have shown impressive progress in recent years. However, existing models are largely limited to turn-based interactions, where each turn must be stepped (i.e…

cs.CV2024★ 1 cited

HexaGen3D: StableDiffusion is just one step away from Fast and Diverse Text-to-3D Generation

Antoine Mercier, Ramin Nakhli, Mahesh Reddy +4

Despite the latest remarkable advances in generative modeling, efficient generation of high-quality 3D assets from textual prompts remains a difficult task. A key challenge lies in…

cs.CV2023

Efficient neural supersampling on a novel gaming dataset

Antoine Mercier, Ruan Erasmus, Yashesh Savani +3

Real-time rendering for video games has become increasingly challenging due to the need for higher resolutions, framerates and photorealism. Supersampling has emerged as an effecti…

cs.CV2023

Is end-to-end learning enough for fitness activity recognition?

Antoine Mercier, Guillaume Berger, Sunny Panchal +5

End-to-end learning has taken hold of many computer vision tasks, in particular, related to still images, with task-specific optimization yielding very strong performance. Neverthe…