collaborators

5 papers

cs.CV2025

Temporal Object-Aware Vision Transformer for Few-Shot Video Object Detection

Yogesh Kumar, Anand Mishra

Few-shot Video Object Detection (FSVOD) addresses the challenge of detecting novel objects in videos with limited labeled examples, overcoming the constraints of traditional detect…

cs.CV2025

Aligning Moments in Time using Video Queries

Yogesh Kumar, Uday Agarwal, Manish Gupta +1

Video-to-video moment retrieval (Vid2VidMR) is the task of localizing unseen events or moments in a target video using a query video. This task poses several challenges, such as th…

cs.AI2025

The Amazon Nova Family of Models: Technical Report and Model Card

Amazon AGI, Aaron Langford, Aayush Shah +783

We present Amazon Nova, a new generation of state-of-the-art foundation models that deliver frontier intelligence and industry-leading price performance. Amazon Nova Pro is a highl…

cs.CV2025

Composite Sketch+Text Queries for Retrieving Objects with Elusive Names and Complex Interactions

Prajwal Gatti, Kshitij Parikh, Dhriti Prasanna Paul +2

Non-native speakers with limited vocabulary often struggle to name specific objects despite being able to visualize them, e.g., people outside Australia searching for numbats. Furt…

cs.CV2025

PatentLMM: Large Multimodal Model for Generating Descriptions for Patent Figures

Shreya Shukla, Nakul Sharma, Manish Gupta +1

Writing comprehensive and accurate descriptions of technical drawings in patent documents is crucial to effective knowledge sharing and enabling the replication and protection of i…