5 papers
TRAP: Benchmark for Task-completion and Resistance to Active Privacy-extraction
Moon Ye-Bin, Nam Hyeon-Woo, Baek Seong-Eun +2
Agents are increasingly deployed in document-intensive workflows where sensitive private information is not an edge case but a routine input, e.g., an agent booking a flight needs…
Early Failure Detection and Intervention in Video Diffusion Models
Kwon Byung-Ki, Sohwi Lim, Nam Hyeon-Woo +2
Text-to-video (T2V) diffusion models have rapidly advanced, yet generations still occasionally fail in practice, such as low text-video alignment or low perceptual quality. Since d…
Patch-wise Retrieval: A Bag of Practical Techniques for Instance-level Matching
Wonseok Choi, Sohwi Lim, Nam Hyeon-Woo +4
Instance-level image retrieval aims to find images containing the same object as a given query, despite variations in size, position, or appearance. To address this challenging tas…
RetouchLLM: Training-free Code-based Image Retouching with Vision Language Models
Moon Ye-Bin, Roy Miles, Tae-Hyun Oh +2
Image retouching not only enhances visual quality but also serves as a means of expressing personal preferences and emotions. However, existing learning-based approaches require la…
Automated Model Discovery via Multi-modal & Multi-step Pipeline
Lee Jung-Mok, Nam Hyeon-Woo, Moon Ye-Bin +2
Automated model discovery is the process of automatically searching and identifying the most appropriate model for a given dataset over a large combinatorial search space. Existing…