4 citations · 11 across the 23 of their papers we have counts for
10 papers · 1 filter
Focus Where It Counts: A Salience-Driven Vision-Language Model for Low Vision Assistance
Jiazhao Liang, Hao Huang, Shuaihang Yuan +8
Vision-language models (VLMs) are rapidly progressing and offer promising capabilities for assistive technologies supporting persons with blindness or low vision. However, existing…
VCS-SLAM: Geometry-Validated Semantic Evidence Fusion for 3D Gaussian SLAM
Raman Jha, Shuaihang Yuan, Yi Fang
Visual SLAM performance often deteriorates in complex real-world applications. Semantic 3D Gaussian SLAM commonly fuses 2D semantic priors into a persistent 3D map using uniform op…
Hierarchical Scoring with 3D Gaussian Splatting for Instance Image-Goal Navigation
Yijie Deng, Shuaihang Yuan, Geeta Chandra Raju Bethala +3
Instance Image-Goal Navigation (IIN) requires autonomous agents to identify and navigate to a target object or location depicted in a reference image captured from any viewpoint. W…
A Chain-of-Thought Subspace Meta-Learning for Few-shot Image Captioning with Large Vision and Language Models
Hao Huang, Shuaihang Yuan, Yu Hao +2
A large-scale vision and language model that has been pretrained on massive data encodes visual and linguistic prior, which makes it easier to generate images and language that are…
FairCLIP: Harnessing Fairness in Vision-Language Learning
Yan Luo, Min Shi, Muhammad Osama Khan +9
Fairness is a critical concern in deep learning, especially in healthcare, where these models influence diagnoses and treatment decisions. Although fairness has been investigated i…
A Multi-Modal Foundation Model to Assist People with Blindness and Low Vision in Environmental Interaction
Yu Hao, Fan Yang, Hao Huang +5
People with blindness and low vision (pBLV) encounter substantial challenges when it comes to comprehensive scene recognition and precise object identification in unfamiliar enviro…