5 papers
Temporal Object-Aware Vision Transformer for Few-Shot Video Object Detection
Yogesh Kumar, Anand Mishra
Few-shot Video Object Detection (FSVOD) addresses the challenge of detecting novel objects in videos with limited labeled examples, overcoming the constraints of traditional detect…
Aligning Moments in Time using Video Queries
Yogesh Kumar, Uday Agarwal, Manish Gupta +1
Video-to-video moment retrieval (Vid2VidMR) is the task of localizing unseen events or moments in a target video using a query video. This task poses several challenges, such as th…
The Amazon Nova Family of Models: Technical Report and Model Card
Amazon AGI, Aaron Langford, Aayush Shah +783
We present Amazon Nova, a new generation of state-of-the-art foundation models that deliver frontier intelligence and industry-leading price performance. Amazon Nova Pro is a highl…
Composite Sketch+Text Queries for Retrieving Objects with Elusive Names and Complex Interactions
Prajwal Gatti, Kshitij Parikh, Dhriti Prasanna Paul +2
Non-native speakers with limited vocabulary often struggle to name specific objects despite being able to visualize them, e.g., people outside Australia searching for numbats. Furt…
PatentLMM: Large Multimodal Model for Generating Descriptions for Patent Figures
Shreya Shukla, Nakul Sharma, Manish Gupta +1
Writing comprehensive and accurate descriptions of technical drawings in patent documents is crucial to effective knowledge sharing and enabling the replication and protection of i…