5 papers
Dynamic Context Adapters: Efficiently Infusing History into Vision-and-Language Models
Yuhang Song, Bor-Jiun Lin, Jiaxu Liu +3
Historical context integration presents a fundamental challenge for Vision-Language Models (VLMs) in sequential decision-making tasks. Current VLMs process visual inputs independen…
AffordMatcher: Affordance Learning in 3D Scenes from Visual Signifiers
Nghia Vu, Tuong Do, Khang Nguyen +8
Affordance learning is a complex challenge in many applications, where existing approaches primarily focus on the geometric structures, visual knowledge, and affordance labels of o…
Lightweight Temporal Transformer Decomposition for Federated Autonomous Driving
Tuong Do, Binh X. Nguyen, Quang D. Tran +3
Traditional vision-based autonomous driving systems often face difficulties in navigating complex environments when relying solely on single-image inputs. To overcome this limitati…
FedEFM: Federated Endovascular Foundation Model with Unseen Data
Tuong Do, Nghia Vu, Tudor Jianu +7
In endovascular surgery, the precise identification of catheters and guidewires in X-ray images is essential for reducing intervention risks. However, accurately segmenting cathete…
Fine-Grained Alignment in Vision-and-Language Navigation through Bayesian Optimization
Yuhang Song, Mario Gianni, Chenguang Yang +4
This paper addresses the challenge of fine-grained alignment in Vision-and-Language Navigation (VLN) tasks, where robots navigate realistic 3D environments based on natural languag…