3 papers
cs.CV2026
VideoWeave: A Data-Centric Approach for Efficient Video Understanding
Zane Durante, Silky Singh, Arpandeep Khatua +6
Training video-language models is often prohibitively expensive due to the high cost of processing long frame sequences and the limited availability of annotated long videos. We pr…
cs.CV2025
Towards Efficient Exemplar Based Image Editing with Multimodal VLMs
Avadhoot Jadhav, Ashutosh Srivastava, Abhinav Java +4
Text-to-Image Diffusion models have enabled a wide array of image editing applications. However, capturing all types of edits through text alone can be challenging and cumbersome.…
cs.CV2024
ReEdit: Multimodal Exemplar-Based Image Editing with Diffusion Models
Ashutosh Srivastava, Tarun Ram Menta, Abhinav Java +4
Modern Text-to-Image (T2I) Diffusion models have revolutionized image editing by enabling the generation of high-quality photorealistic images. While the de facto method for perfor…