3 papers
cs.AI2024
V-Zen: Efficient GUI Understanding and Precise Grounding With A Novel Multimodal LLM
Abdur Rahman, Rajat Chawla, Muskaan Kumar +4
In the rapidly evolving landscape of AI research and application, Multimodal Large Language Models (MLLMs) have emerged as a transformative force, adept at interpreting and integra…
cs.AI2024
AUTONODE: A Neuro-Graphic Self-Learnable Engine for Cognitive GUI Automation
Arkajit Datta, Tushar Verma, Rajat Chawla +2
In recent advancements within the domain of Large Language Models (LLMs), there has been a notable emergence of agents capable of addressing Robotic Process Automation (RPA) challe…
cs.CV2024
Veagle: Advancements in Multimodal Representation Learning
Rajat Chawla, Arkajit Datta, Tushar Verma +6
Lately, researchers in artificial intelligence have been really interested in how language and vision come together, giving rise to the development of multimodal models that aim to…