7 papers
DesignSense: A Human Preference Dataset and Reward Modeling Framework for Graphic Layout Generation
Varun Gopal, Rishabh Jain, Aradhya Mathur +6
Graphic layouts serve as an important and engaging medium for visual communication across different channels. While recent layout generation models have demonstrated impressive cap…
AFRAgent : An Adaptive Feature Renormalization Based High Resolution Aware GUI agent
Neeraj Anand, Rishabh Jain, Sohan Patnaik +2
There is a growing demand for mobile user interface (UI) automation, driven by its broad applications across industries. With the advent of visual language models (VLMs), GUI autom…
Towards Efficient Exemplar Based Image Editing with Multimodal VLMs
Avadhoot Jadhav, Ashutosh Srivastava, Abhinav Java +4
Text-to-Image Diffusion models have enabled a wide array of image editing applications. However, capturing all types of edits through text alone can be challenging and cumbersome.…
Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning
Sohan Patnaik, Milan Aggarwal, Sumit Bhatia +1
LLMssuch as GPT-4 have shown a remarkable ability to solve complex questions by generating step-by-step rationales. Prior works have utilized this capability to improve smaller and…
EOPose : Exemplar-based object reposing using Generalized Pose Correspondences
Sarthak Mehrotra, Rishabh Jain, Mayur Hemani +2
Reposing objects in images has a myriad of applications, especially for e-commerce where several variants of product images need to be produced quickly. In this work, we leverage t…
It Helps to Take a Second Opinion: Teaching Smaller LLMs to Deliberate Mutually via Selective Rationale Optimisation
Sohan Patnaik, Milan Aggarwal, Sumit Bhatia +1
Very large language models (LLMs) such as GPT-4 have shown the ability to handle complex tasks by generating and self-refining step-by-step rationales. Smaller language models (SLM…