4 papers
Show, Don't Ask: Generative Visual Disambiguation for Composed Image Retrieval with Turn-Valid Coverage
Amsisan Tran, Baogh Le, Tuan Kiet Pham +1
Composed image retrieval (CIR) uses a reference image and a text modification to search for a target image. However, such queries often describe several possible images rather than…
Resolving Ambiguity in Composed Image Retrieval via Calibrated Interaction
Amsisan Tran, Baogh Le, Tuan Kiet Pham +1
Composed image retrieval (CIR) searches a corpus with a reference image and a text describing how to modify it. Despite rapid progress from triplet-trained compositors to zero-shot…
RelWitness: Open-Vocabulary 3D Scene Graph Generation with Visual-Geometric Relation Witnesses
Minh Anh Nguyen, Quang Huy Tran, Bao Ngoc Le +2
Open-vocabulary 3D scene graph generation seeks to describe object instances and their relations with flexible natural-language predicates. The central difficulty is not only vocab…
ScriptHOI: Learning Scripted State Transitions for Open-Vocabulary Human-Object Interaction Detection
Minh Anh Nguyen, Quang Huy Tran, Bao Ngoc Le +3
Open-vocabulary human-object interaction (HOI) detection requires recognizing interaction phrases that may not appear as annotated categories during training. Recent vision-languag…