3 citations · 3 across the 3 of their papers we have counts for
5 papers
Pixel-Grounded Retrieval for Knowledgeable Large Multimodal Models
Jeonghwan Kim, Renjie Tao, Sanat Sharma +8
Visual Question Answering (VQA) often requires coupling fine-grained perception with factual knowledge beyond the input image. Prior multimodal Retrieval-Augmented Generation (MM-R…
ThematicPlane: Bridging Tacit User Intent and Latent Spaces for Image Generation
Daniel Lee, Nikhil Sharma, Donghoon Shin +4
Generative AI has made image creation more accessible, yet aligning outputs with nuanced creative intent remains challenging, particularly for non-experts. Existing tools often req…
Infogent: An Agent-Based Framework for Web Information Aggregation
Revanth Gangi Reddy, Sagnik Mukherjee, Jeonghwan Kim +3
Despite seemingly performant web agents on the task-completion benchmarks, most existing methods evaluate the agents based on a presupposition: the web navigation task consists of…
Aligning LLMs with Individual Preferences via Interaction
Shujin Wu, May Fung, Cheng Qian +3
As large language models (LLMs) demonstrate increasingly advanced capabilities, aligning their behaviors with human values and preferences becomes crucial for their wide adoption.…
ARMADA: Attribute-Based Multimodal Data Augmentation
Xiaomeng Jin, Jeonghwan Kim, Yu Zhou +4
In Multimodal Language Models (MLMs), the cost of manually annotating high-quality image-text pair data for fine-tuning and alignment is extremely high. While existing multimodal d…