4 papers
Rethinking Progression of Memory State in Robotic Manipulation: An Object-Centric Perspective
Nhat Chung, Taisei Hanyu, Toan Nguyen +9
As embodied agents operate in increasingly complex environments, the ability to perceive, track, and reason about individual object instances over time becomes essential, especiall…
UNO: Unifying One-stage Video Scene Graph Generation via Object-Centric Visual Representation Learning
Huy Le, Nhat Chung, Tung Kieu +2
Video Scene Graph Generation (VidSGG) aims to represent dynamic visual content by detecting objects and modeling their temporal interactions as structured graphs. Prior studies typ…
BiMa: Towards Biases Mitigation for Text-Video Retrieval via Scene Element Guidance
Huy Le, Nhat Chung, Tung Kieu +2
Text-video retrieval (TVR) systems often suffer from visual-linguistic biases present in datasets, which cause pre-trained vision-language models to overlook key details. To addres…
With Limited Data for Multimodal Alignment, Let the STRUCTURE Guide You
Fabian Gröger, Shuo Wen, Huyen Le +1
Multimodal models have demonstrated powerful capabilities in complex tasks requiring multimodal alignment, including zero-shot classification and cross-modal retrieval. However, ex…