collaborators

6 papers

cs.CV2026

Show, Don't Ask: Generative Visual Disambiguation for Composed Image Retrieval with Turn-Valid Coverage

Amsisan Tran, Baogh Le, Tuan Kiet Pham +1

Composed image retrieval (CIR) uses a reference image and a text modification to search for a target image. However, such queries often describe several possible images rather than…

cs.CV2026

Resolving Ambiguity in Composed Image Retrieval via Calibrated Interaction

Amsisan Tran, Baogh Le, Tuan Kiet Pham +1

Composed image retrieval (CIR) searches a corpus with a reference image and a text describing how to modify it. Despite rapid progress from triplet-trained compositors to zero-shot…

cs.CV2026

RelWitness: Open-Vocabulary 3D Scene Graph Generation with Visual-Geometric Relation Witnesses

Minh Anh Nguyen, Quang Huy Tran, Bao Ngoc Le +2

Open-vocabulary 3D scene graph generation seeks to describe object instances and their relations with flexible natural-language predicates. The central difficulty is not only vocab…

cs.CV2026

ScriptHOI: Learning Scripted State Transitions for Open-Vocabulary Human-Object Interaction Detection

Minh Anh Nguyen, Quang Huy Tran, Bao Ngoc Le +3

Open-vocabulary human-object interaction (HOI) detection requires recognizing interaction phrases that may not appear as annotated categories during training. Recent vision-languag…

cs.CV2026

ReLIC-SGG: Relation Lattice Completion for Open-Vocabulary Scene Graph Generation

Amir Hosseini, Sara Farahani, Xinyi Li +1

Open-vocabulary scene graph generation (SGG) aims to describe visual scenes with flexible relation phrases beyond a fixed predicate set. Existing methods usually treat annotated tr…

cs.CV2026

CAGE-SGG: Counterfactual Active Graph Evidence for Open-Vocabulary Scene Graph Generation

Suiyang Guang, Chenyu Liu, Ruohan Zhang +1

Open-vocabulary scene graph generation (SGG) aims to describe visual scenes with flexible and fine-grained relation phrases beyond a fixed predicate vocabulary. While recent vision…