5 papers
Pixel-Grounded Retrieval for Knowledgeable Large Multimodal Models
Jeonghwan Kim, Renjie Tao, Sanat Sharma +8
Visual Question Answering (VQA) often requires coupling fine-grained perception with factual knowledge beyond the input image. Prior multimodal Retrieval-Augmented Generation (MM-R…
ThematicPlane: Bridging Tacit User Intent and Latent Spaces for Image Generation
Daniel Lee, Nikhil Sharma, Donghoon Shin +4
Generative AI has made image creation more accessible, yet aligning outputs with nuanced creative intent remains challenging, particularly for non-experts. Existing tools often req…
Finer: Investigating and Enhancing Fine-Grained Visual Concept Recognition in Large Vision Language Models
Jeonghwan Kim, Heng Ji
Recent advances in instruction-tuned Large Vision-Language Models (LVLMs) have imbued the models with the ability to generate high-level, image-grounded explanations with ease. Whi…
Aligning LLMs with Individual Preferences via Interaction
Shujin Wu, May Fung, Cheng Qian +3
As large language models (LLMs) demonstrate increasingly advanced capabilities, aligning their behaviors with human values and preferences becomes crucial for their wide adoption.…
Infogent: An Agent-Based Framework for Web Information Aggregation
Revanth Gangi Reddy, Sagnik Mukherjee, Jeonghwan Kim +3
Despite seemingly performant web agents on the task-completion benchmarks, most existing methods evaluate the agents based on a presupposition: the web navigation task consists of…