activity
20242026
most citedThematicPlane: Bridging Tacit User Intent and Latent Spaces for Image Generation

3 citations · 3 across the 3 of their papers we have counts for

collaborators

5 papers

cs.CV2026

Pixel-Grounded Retrieval for Knowledgeable Large Multimodal Models

Jeonghwan Kim, Renjie Tao, Sanat Sharma +8

Visual Question Answering (VQA) often requires coupling fine-grained perception with factual knowledge beyond the input image. Prior multimodal Retrieval-Augmented Generation (MM-R…

cs.HC20253 cited

ThematicPlane: Bridging Tacit User Intent and Latent Spaces for Image Generation

Daniel Lee, Nikhil Sharma, Donghoon Shin +4

Generative AI has made image creation more accessible, yet aligning outputs with nuanced creative intent remains challenging, particularly for non-experts. Existing tools often req…

cs.AI2024

Infogent: An Agent-Based Framework for Web Information Aggregation

Revanth Gangi Reddy, Sagnik Mukherjee, Jeonghwan Kim +3

Despite seemingly performant web agents on the task-completion benchmarks, most existing methods evaluate the agents based on a presupposition: the web navigation task consists of…

cs.CL2024

Aligning LLMs with Individual Preferences via Interaction

Shujin Wu, May Fung, Cheng Qian +3

As large language models (LLMs) demonstrate increasingly advanced capabilities, aligning their behaviors with human values and preferences becomes crucial for their wide adoption.…

cs.AI2024

ARMADA: Attribute-Based Multimodal Data Augmentation

Xiaomeng Jin, Jeonghwan Kim, Yu Zhou +4

In Multimodal Language Models (MLMs), the cost of manually annotating high-quality image-text pair data for fine-tuning and alignment is extremely high. While existing multimodal d…