3 papers
cs.CV2026
DeepInsert: Early Layer Bypass for Efficient and Performant Multimodal Understanding
Moulik Choraria, Xinbo Wu, Akhil Bhimaraju +5
Hyperscaling of data and parameter count in LLMs is yielding diminishing improvement when weighed against training costs, underlining a growing need for more efficient finetuning a…
cs.CV2025
PIXELS: Progressive Image Xemplar-based Editing with Latent Surgery
Shristi Das Biswas, Matthew Shreve, Xuelu Li +2
Recent advancements in language-guided diffusion models for image editing are often bottle-necked by cumbersome prompt engineering to precisely articulate desired changes. An intui…
cs.CV2024
Semantically Grounded QFormer for Efficient Vision Language Understanding
Moulik Choraria, Xinbo Wu, Sourya Basu +5
General purpose Vision Language Models (VLMs) have received tremendous interest in recent years, owing to their ability to learn rich vision-language correlations as well as their…