4 papers
TRAP: Benchmark for Task-completion and Resistance to Active Privacy-extraction
Moon Ye-Bin, Nam Hyeon-Woo, Baek Seong-Eun +2
Agents are increasingly deployed in document-intensive workflows where sensitive private information is not an edge case but a routine input, e.g., an agent booking a flight needs…
SplatReasoner: Enhancing Embodied Reasoning and Grounding by Novel View Synthesis
Kim Yu-Ji, Dahye Lee, Kim Jun-Seong +6
Vision-Language Models (VLMs) have demonstrated strong reasoning capabilities over images and videos, yet their application to embodied scene understanding often constrained by the…
VSC: Visual Search Compositional Text-to-Image Diffusion Model
Do Huu Dat, Nam Hyeonu, Po-Yuan Mao +1
Text-to-image diffusion models have shown impressive capabilities in generating realistic visuals from natural-language prompts, yet they often struggle with accurately binding att…
BEAF: Observing BEfore-AFter Changes to Evaluate Hallucination in Vision-language Models
Moon Ye-Bin, Nam Hyeon-Woo, Wonseok Choi +1
Vision language models (VLMs) perceive the world through a combination of a visual encoder and a large language model (LLM). The visual encoder, pre-trained on large-scale vision-t…