activity
20242026
collaborators

5 papers

cs.CV2026

Dynamic Context Adapters: Efficiently Infusing History into Vision-and-Language Models

Yuhang Song, Bor-Jiun Lin, Jiaxu Liu +3

Historical context integration presents a fundamental challenge for Vision-Language Models (VLMs) in sequential decision-making tasks. Current VLMs process visual inputs independen…

cs.CV2026

AffordMatcher: Affordance Learning in 3D Scenes from Visual Signifiers

Nghia Vu, Tuong Do, Khang Nguyen +8

Affordance learning is a complex challenge in many applications, where existing approaches primarily focus on the geometric structures, visual knowledge, and affordance labels of o…

cs.CV2025

Lightweight Temporal Transformer Decomposition for Federated Autonomous Driving

Tuong Do, Binh X. Nguyen, Quang D. Tran +3

Traditional vision-based autonomous driving systems often face difficulties in navigating complex environments when relying solely on single-image inputs. To overcome this limitati…

cs.CV2025

FedEFM: Federated Endovascular Foundation Model with Unseen Data

Tuong Do, Nghia Vu, Tudor Jianu +7

In endovascular surgery, the precise identification of catheters and guidewires in X-ray images is essential for reducing intervention risks. However, accurately segmenting cathete…

cs.CV2024

Fine-Grained Alignment in Vision-and-Language Navigation through Bayesian Optimization

Yuhang Song, Mario Gianni, Chenguang Yang +4

This paper addresses the challenge of fine-grained alignment in Vision-and-Language Navigation (VLN) tasks, where robots navigate realistic 3D environments based on natural languag…