5 papers
From Plans to Pixels: Learning to Plan and Orchestrate for Open-Ended Image Editing
Anirudh Sundara Rajan, Krishna Kumar Singh, Yong Jae Lee
Modern image editing models produce realistic results but struggle with abstract, multi step instructions (e.g., ``make this advertisement more vegetarian-friendly''). Prior agent…
LoopUS: Recasting Pretrained LLMs into Looped Latent Refinement Models
Taekhyun Park, Yongjae Lee, Dohee Kim +1
Looped computation shows promise in improving the reasoning-oriented performance of LLMs by scaling test-time compute. However, existing approaches typically require either trainin…
Reasoning-Augmented Representations for Multimodal Retrieval
Jianrui Zhang, Anirudh Sundara Rajan, Brandon Han +3
Universal Multimodal Retrieval (UMR) seeks any-to-any search across text and vision, yet modern embedding models remain brittle when queries require latent reasoning (e.g., resolvi…
VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection
Zeyi Huang, Yuyang Ji, Anirudh Sundara Rajan +5
We introduce VisTA, a new reinforcement learning framework that empowers visual agents to dynamically explore, select, and combine tools from a diverse library based on empirical p…
Stay-Positive: A Case for Ignoring Real Image Features in Fake Image Detection
Anirudh Sundara Rajan, Yong Jae Lee
Detecting AI generated images is a challenging yet essential task. A primary difficulty arises from the detectors tendency to rely on spurious patterns, such as compression artifac…